Blueprint · part 2 of 6
The same agentic stack with no platform vendor: CNCF, NVIDIA and Microsoft projects only
Part 1 mapped the Red Hat blueprint for AI agents on OpenShift. This part replaces every Red Hat component with the open project it packages, or with its NVIDIA or Microsoft equivalent, and looks at what you actually give up. Spoiler: not the architecture.
HokonokenSeptember 2026Reading time: 7 minViews are my own, not my employer's
Three things to take away
- Nothing in the diagram is proprietary. Every project is Apache 2.0 or MIT: I checked the licence of each repository on GitHub on 21 September 2026. The stack runs on any Kubernetes.
- What a platform vendor sells is not the architecture, it is the assembly. Certified images, tested combinations, an operator that installs the whole thing, and someone to call. Remove the vendor and that work moves to your team.
- Two things are not free, and it is worth knowing which. NVIDIA's NIM containers and CUDA-X libraries in production, and Microsoft's hosted control planes. Neither is needed: vLLM serves the models, and the Agent Governance Toolkit needs no Microsoft account.
Why build it this way
Three situations call for a vendor-free stack. A sovereign or air-gapped environment where the subscription model does not fit. A team that already runs Kubernetes and wants to understand every moving part before paying for a bundle. Or simply the exercise itself: if you can name the open project behind each box, you understand what the product adds, and you negotiate better.
The rules of the exercise: same flows as part 1, same two execution patterns, same two exits for the agent. Only the boxes change. Where a replacement is a CNCF project I give its maturity level; where it belongs to another foundation I say which.
Two decisions the variant forces you to make
Who decides a tool call. In the Red Hat stack, Authorino answers at the server-and-tool grain. Here, Envoy AI Gateway's MCPRoute carries its own per-tool authorization in CEL, and Open Policy Agent sits at the ingress. If you let both decide, you get two places where a call can be refused and nobody knows which one did it. In the diagram, OPA decides who reaches what at the door; MCPRoute decides each call. One point of decision per question.
Requests versus tokens. Envoy Gateway limits requests at the ingress. Envoy AI Gateway limits tokens on the LLM route, which is the quantity that costs money. Keeping the two apart avoids a common mistake: a request limit that lets a single prompt burn a whole budget.
What "free" means exactly
- NVIDIA. OpenShell, NemoClaw, NeMo Guardrails and NeMo Agent Toolkit are Apache 2.0, without restriction. NIM and the CUDA-X libraries are not: free to develop with, licensed under NVIDIA AI Enterprise in production. The diagram keeps them optional; vLLM does the serving job.
- Microsoft. Agent Governance Toolkit and Agent Framework are MIT, without restriction. Agent 365 and Foundry Control Plane are paid, closed services and are not in the diagram. An agent governed by the toolkit needs no Microsoft account.
- CNCF and the other foundations. Everything is Apache 2.0. The cost moves to operations: each component is upgraded, patched and integrated by you, with no vendor assembling them.
- The models. gpt-oss is Apache 2.0. Nemotron is under NVIDIA's open model licence, Llama under Meta's community licence: free to use, with conditions on redistribution. Read them before serving a model to a customer.
What you take on
Count the boxes in the diagram: about twenty projects from six foundations and two vendors, each with its own release cadence and its own security advisories. A platform vendor tests one combination and tells you when to upgrade. Without one, that calendar is yours, and so is the first integration of each new version. This is not an argument against the variant. It is the price, stated plainly, and for a sovereign environment it is often the right price to pay.
Next in the series
- Part 3What the regulations actually ask for (EU AI Act, NIS2, ISO 42001, Swiss FADP and FINMA), and which layer of the stack has to answer.
- Part 4Action governance with Microsoft's Agent Governance Toolkit: how those requirements are enforced, and where it plugs into this stack.
- Part 5Observability: how it is proven. What OpenTelemetry, OCSF and MLflow record, and what they cannot tell you.
- Part 6Rolling out a model blue/green: how the stack changes without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.