blog: GPU Kubernetes infrastructure for AI workloads - #20631
blog: GPU Kubernetes infrastructure for AI workloads#20631workprentice[bot] wants to merge 1 commit into
Conversation
Covers the GPU/accelerator compute substrate for AI and agentic workloads on Kubernetes with Pulumi: GPU node pools across EKS/GKE/AKS, GPU operator install, Kueue quotas and ResourceFlavors, taints/tolerations, time-slicing vs MIG, and cost guardrails. Cross-links to the agent-runtime post (ai-agents-on-kubernetes) rather than duplicating it, since that post already covers agent CRDs and orchestration. This post closes the GPU/accelerator gap in our AI Infrastructure and Kubernetes GEO topics instead.
Social Media Reviewcontent/blog/gpu-kubernetes-ai-workloads/index.mdX — PASSLinkedIn — PASSBluesky — FAILReasons:
Suggested copyBluesky (176/300 chars) — minimum-change repair; split into two paragraphs at the existing sentence boundary, no wording changed:
Suggestions (advisory)These are stylistic notes — they don't block the post. X
Updated for commit |
Pre-merge Review — Last updated 2026-08-01T12:23:56ZTip Summary: This PR adds one new blog post — a single-subject how-to on provisioning and governing GPU capacity on Kubernetes with Pulumi (node pools, the NVIDIA GPU Operator, Kueue quotas, taints, time-slicing vs. MIG), positioned as the substrate companion to the existing How to Run AI Agents on Kubernetes with Pulumi and structured like it (FAQ section, "Where to go next" link list, Review confidence:
Investigation log
🔍 Verification trail58 claims extracted · 43 verified · 1 unverifiable · 3 contradicted · 1 framing-drift · 2 detector findings
📊 Editorial balanceSection depth, mention distribution, recommendation steering
🚨 Outstanding in this PRThese must be resolved or refuted before merging.
|
📋 Triaged verifier findingsI double-checked these and realized they weren't real findings — click to expand
💡 Pre-existing issues in touched files (optional)No pre-existing issues in touched files. ✅ Resolved since last reviewNo items resolved since the last review. 📜 Review history
Important Please don't hide, resolve, or delete this comment! It breaks things! 📖 How pre-merge review works — the full lifecycle, short-circuits, and escape hatches. |
|
Your site preview for commit 9ee6c6c is ready! 🎉 http://www-testing-pulumi-docs-origin-pr-20631-9ee6c6ce.s3-website.us-west-2.amazonaws.com Changed pages: |
What
Adds a new blog post: "GPU Kubernetes Infrastructure for AI Workloads with Pulumi" at
/blog/gpu-kubernetes-ai-workloads/.Why the angle changed from the original card
The Marketing Content Calendar card behind this PR asked for a post on "Pulumi + Kubernetes for Agentic AI Workloads." That exact angle was already shipped in PR #20550 (merged 2026-07-29): "How to Run AI Agents on Kubernetes with Pulumi" at
/blog/ai-agents-on-kubernetes/. Writing the card literally would have cannibalized a 3-day-old flagship post on the same keywords.Fresh Profound visibility data (2026-07-02 to 2026-08-01) shows Pulumi's two weakest GEO topics are AI Infrastructure (10.8% visibility, versus NVIDIA 64.1%, Google 61.8%, Databricks 53.6%, CoreWeave 47.8%, AWS 46.4%) and Kubernetes (no measured Pulumi visibility at all, versus Docker 61.1%, Kubernetes 56.2%, Azure 35.9%). Both gaps are shaped around compute/accelerator vendors, not agent frameworks.
So this post keeps the card's strategic intent (close the Kubernetes and AI Infrastructure gaps, ship Kubernetes-native code, hit every AEO requirement, Alex Leventer byline) but shifts the angle to the adjacent, unowned lane: the GPU/accelerator compute substrate underneath agentic workloads (node pools, GPU operator install, Kueue quotas, taints/tolerations, time-slicing vs. MIG, cost guardrails), rather than the agent runtime itself. The two posts cross-link and reinforce each other instead of competing for the same query space.
AEO / schema choices
faq_schema: trueandhowto_schema: true, matching the repo's schema-gating rules.howto-entity.htmluses a single running counter over every^\d+\.line in the document — any other numbered list would bleed into the HowTo steps.faq-entity.htmlhas no section scoping, so all?-ending H2s/H3s land inFAQPage.mainEntity), and a dedicated "Frequently asked questions" section adds five### ...?H3s targeting real query phrasings./blog/the-agentic-infrastructure-era/, quoted verbatim, including the proprietary "20% of infrastructure deployments" stat)./blog/ai-agents-on-kubernetes/,/blog/the-agentic-infrastructure-era/,/blog/ai-infrastructure-tools/,/docs/esc/, plus/registry/packages/kubernetes/(a generated route used the same way elsewhere in the repo).Verification
ai-agents-on-kubernetes,ai-infrastructure-tools,fully-automated-ai-inference-aws-azure-gcp-pulumi,self-host-gemma4-llama-cpp-k8s-tailscale-pulumi,ai-ml-on-kubernetes-google-cloud-llm-rag,kubernetes-chart-v4) to confirm no keyword or angle overlap and to harvest verified resource shapes (Helm v4Chart,apiextensions.CustomResourcepatterns, ESC-sourced credentials).alex-leventerauthor slug exists atdata/team/team/alex-leventer.toml.categoryscalar from the closed set, trailing newline present); a Node.js toolchain was not available in this session to runscripts/lint/lint-markdown.jsdirectly, so please treat CI's lint run as authoritative.🧠 This PR was created by workprentice on behalf of the Pulumi SEO/AEO content team — no
get_me-equivalent tool was available in this session to resolve a specific requester's username.