Blueprint · part 2 of 6

The same agentic stack with no platform vendor: CNCF, NVIDIA and Microsoft projects only

Part 1 mapped the Red Hat blueprint for AI agents on OpenShift. This part replaces every Red Hat component with the open project it packages, or with its NVIDIA or Microsoft equivalent, and looks at what you actually give up. Spoiler: not the architecture.

HokonokenSeptember 2026Reading time: 7 minViews are my own, not my employer's

Three things to take away

  1. Nothing in the diagram is proprietary. Every project is Apache 2.0 or MIT: I checked the licence of each repository on GitHub on 21 September 2026. The stack runs on any Kubernetes.
  2. What a platform vendor sells is not the architecture, it is the assembly. Certified images, tested combinations, an operator that installs the whole thing, and someone to call. Remove the vendor and that work moves to your team.
  3. Two things are not free, and it is worth knowing which. NVIDIA's NIM containers and CUDA-X libraries in production, and Microsoft's hosted control planes. Neither is needed: vLLM serves the models, and the Agent Governance Toolkit needs no Microsoft account.

Why build it this way

Three situations call for a vendor-free stack. A sovereign or air-gapped environment where the subscription model does not fit. A team that already runs Kubernetes and wants to understand every moving part before paying for a bundle. Or simply the exercise itself: if you can name the open project behind each box, you understand what the product adds, and you negotiate better.

The rules of the exercise: same flows as part 1, same two execution patterns, same two exits for the agent. Only the boxes change. Where a replacement is a CNCF project I give its maturity level; where it belongs to another foundation I say which.

Users · business applications · notebooks (Kubeflow) Envoy GatewayGateway API · request rate limitingCNCFCNCF graduated KeycloakOIDC, JWTCNCFCNCF incubating Open Policy Agentext_authz at ingressCNCFCNCF graduated SPIFFE / SPIREworkload identityCNCFCNCF graduated CONTROL BAND Argo CD · HelmGitOps, blue/green RolloutsCNCFCNCF graduated kagentagent lifecycle, HITLCNCFCNCF sandbox NemoClawagent-in-pod blueprintNVIDIApre-1.0 OpenShellpolicies, drafts, proverNVIDIAApache 2.0 Agent SandboxSIG Apps controllerK8Salpha Kyverno · cosignadmission, signed imagesCNCFCNCF incubating Kubeflownotebooks, model registryCNCFCNCF graduated PATTERN A · AGENT-AS-A-WORKLOAD · ONE POD PER AGENT OpenShell sandbox (NemoClaw): Landlock, seccomp, netns Harness + MCP clientsMicrosoft Agent Framework · NeMo Agent ToolkitMS · NVIDIA OpenShell supervisorRego policy · Privacy Router: where inference goesNVIDIA AGT · SDK or sidecarper-action policy, declared intentMS PATTERN B · AGENT-AS-A-SERVICE · SHARED RUNTIME OGX is open source and vendor-neutral; runs without Red Hat OGX serverex-Llama Stack · Responses API · tools · Vector_IOOAuth, ABAC, multi-tenantOSS Containers API → OpenShellearly validation, not shipping (Red Hat, 05.2026)NVIDIA AGT · server sidecarsame policies as pattern AMS Envoy AI GatewayMCP mode · per-tool CEL decidesENVOYv1.0 MCP serversDMS, CRM, ERP Business systemsAPIs, data AGT · ACTION GOVERNANCE (Microsoft, MIT) · proposed insertion, not in the blueprint PolicyCedar · Rego Intentmission, drift Approvalwebhook, elicitation AuditMerkle tree Envoy AI GatewayLLM mode · keys, tokens, costsENVOYv1.0 NeMo Guardrailsinput / outputNVIDIAApache 2.0 llm-d routingno model runs hereCNCFCNCF sandbox Vector storeMilvus · pgvector · QdrantOSS Model RegistryKubeflow · candidatesCNCFCNCF graduated vLLM · the engineKServe; NIM optionalPYTORCHCNCF incubating OIDC or API key direct request to the agent Responses API JWT JWT checked tool name + arguments RAG OTLP traces OBSERVABILITYOpenTelemetry Collector → Jaeger (traces) · Prometheus + DCGM exporter (GPU, vLLM) · OpenShell OCSF events · AGT Merkle audit PLATFORMKubernetes · containerd · Cilium (NetworkPolicy) · GPU Operator and Network Operator (NVIDIA) · Node Feature Discovery · Kueue · Rook-Ceph or Longhorn HARDWARENVIDIA GPUs; also AMD and Intel Gaudi · RoCE, InfiniBand · NVMe · on-premises, sovereign or cloud Maturity at the bottom of each box, as published by the vendor in September 2026; the inference chain is linearized: KServe drives llm-d, the guardrails are a called service.
Purple is a CNCF project, green NVIDIA, dashed blue Microsoft, grey another foundation. The two execution patterns are the ones from part 1: the agent in its own OpenShell-confined pod, or the shared OGX server behind a Responses API. OGX is open source and runs without Red Hat. The bottom label of every box is its maturity as published by its foundation or vendor.

What replaces what

RoleIn the Red Hat stackHereStatusWhat you lose
KubernetesOpenShiftKubernetesCNCF graduatedSupport, certified images, the OpenShift console
Ingress gatewayGateway API on Envoy (Kuadrant)Envoy GatewayCNCF graduated (Envoy)Nothing: same engine
External authorizationAuthorinoOpen Policy Agent as ext_authzCNCF graduatedThe AuthPolicy CRDs; you write Rego
MCP gateway and token quotasMCP gateway (Tech Preview), MaaS, LimitadorEnvoy AI Gateway: MCPRoute, per-tool CEL authorization, OAuth, token-based limiting. v1.0 since June 2026Built on Envoy Gateway; the project is moving to the Agentic AI FoundationThe MCP catalog in the dashboard
Agent lifecycleOGX operator (3.5 early access); Rossoctl, ex-Kagenti (research)kagentCNCF sandboxThe PatternFly agent catalog; kagent has its own UI
Harness and shared runtimeLangGraph, Deep Agents; OGX serverMicrosoft Agent Framework, NeMo Agent Toolkit; the same OGX server, open sourceMIT, Apache 2.0; OGX open sourceNothing: OGX runs without Red Hat
Content guardrailsTrustyAI + NeMo GuardrailsNeMo Guardrails aloneNVIDIA, Apache 2.0The TrustyAI detectors; NeMo covers them
Inference routingllm-d, semantic routingllm-d + Gateway API inference extensionCNCF sandbox; Kubernetes SIGRed Hat's semantic routing (roadmap)
Model servingKServe, vLLM, NIMKServe, vLLM; NIM optionalCNCF incubating; PyTorch FoundationNothing
Model registry, notebooksOpenShift AI dashboardKubeflow Model Registry and NotebooksCNCF graduatedThe unified dashboard
TracesTempo, MLflowJaegerCNCF graduatedMLflow experiment tracking (Linux Foundation, can be added)
Network, admission, signaturesOVN, Compliance Operator, ACSCilium, Kyverno, cosignCNCF graduated, incubating; OpenSSFThe shipped compliance profiles

Two decisions the variant forces you to make

Who decides a tool call. In the Red Hat stack, Authorino answers at the server-and-tool grain. Here, Envoy AI Gateway's MCPRoute carries its own per-tool authorization in CEL, and Open Policy Agent sits at the ingress. If you let both decide, you get two places where a call can be refused and nobody knows which one did it. In the diagram, OPA decides who reaches what at the door; MCPRoute decides each call. One point of decision per question.

Requests versus tokens. Envoy Gateway limits requests at the ingress. Envoy AI Gateway limits tokens on the LLM route, which is the quantity that costs money. Keeping the two apart avoids a common mistake: a request limit that lets a single prompt burn a whole budget.

What "free" means exactly

What you take on

Count the boxes in the diagram: about twenty projects from six foundations and two vendors, each with its own release cadence and its own security advisories. A platform vendor tests one combination and tells you when to upgrade. Without one, that calendar is yours, and so is the first integration of each new version. This is not an argument against the variant. It is the price, stated plainly, and for a sovereign environment it is often the right price to pay.

Next in the series
  1. Part 3What the regulations actually ask for (EU AI Act, NIS2, ISO 42001, Swiss FADP and FINMA), and which layer of the stack has to answer.
  2. Part 4Action governance with Microsoft's Agent Governance Toolkit: how those requirements are enforced, and where it plugs into this stack.
  3. Part 5Observability: how it is proven. What OpenTelemetry, OCSF and MLflow record, and what they cannot tell you.
  4. Part 6Rolling out a model blue/green: how the stack changes without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.

Sources
  1. Architect an open blueprint for cloud-native AI agents, Red Hat Developer, 20 July 2026, for the flows
  2. CNCF project pages: kagent, KServe; KServe becomes a CNCF incubating project
  3. Envoy AI Gateway, Model Context Protocol gateway; release notes
  4. llm-d documentation
  5. Working with OGX, Red Hat OpenShift AI 3.5 documentation
  6. Licences read on GitHub on 21 September 2026 for each repository named in the diagram

Independent work, not affiliated with the CNCF, Red Hat, NVIDIA or Microsoft. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.