Blueprint · part 1 of 6

AI agents on OpenShift: the open blueprint, layer by layer

Red Hat published its reference architecture for cloud-native AI agents in July 2026. I read it, then read the code of the projects it names. Here is the whole stack on one page, from the GPU to the user, and the two gaps Red Hat itself says are still open.

HokonokenSeptember 2026Reading time: 6 minViews are my own, not my employer's

Three things to take away

  1. An agent has two exits, and each has its own gateway: tool calls go through an MCP gateway, inference goes through a model gateway. Governance lives at those two doors, not inside the model.
  2. Red Hat governs by identity and tool authorization, never by reading the prompt. A prompt injection that makes the model call a forbidden tool fails at the infrastructure layer, because the model was never the one deciding.
  3. The blueprint states what it does not solve yet: human approval of high-risk actions and policy authoring at fleet scale. Those two sentences are the most useful part of the document.

Why a map

Every vendor now ships "agentic AI". What is hard to find is a picture of how the pieces fit when you actually have to run them: which component holds the agent's identity, which one decides whether a tool call goes through, where the model lives, what emits the audit trail. Red Hat's blueprint is the first one I have seen that answers those questions with named, open projects. So I drew it.

Everything below comes from public sources: the blueprint article, Red Hat's article on the OpenShift AI dashboard, and the source code of the projects on GitHub. Where a component's behaviour is only described in documentation and not visible in code, I say so.

Users · business applications · notebooks and IDEs Ingress gatewayGateway API · Envoy (Kuadrant)RED HATGA KeycloakOIDC, JWTCNCF Authorinoext_authz · server + toolRED HATGA Limitadortoken quotasRED HATGA CONTROL BAND GitOps · Helmdeclarative deploymentRHGA OGX Operatoragents as a shared serviceRH3.5 early access NemoClawagent-in-pod blueprintNVIDIApre-1.0 OpenShellpolicies, drafts, proverNVIDIAdev preview Agent SandboxSIG Apps controllerK8Salpha SPIFFE / SPIREworkload identityCNCFCNCF graduated Rossoctl (ex-Kagenti)research, not in blueprintRH PATTERN A · AGENT-AS-A-WORKLOAD · ONE POD PER AGENT OpenShell sandbox (NemoClaw): Landlock, seccomp, netns Harness + MCP clientsLangGraph · Deep Agents · NeMo Agent Toolkit OpenShell supervisorRego policy · Privacy Router: where inference goesNVIDIA AGT · SDK or sidecarper-action policy, declared intentMS PATTERN B · AGENT-AS-A-SERVICE · SHARED RUNTIME what OpenShift AI 3.5 ships OGX serverex-Llama Stack · Responses API · tools · Vector_IOOAuth, ABAC, multi-tenantRED HAT3.5 early access Containers API → OpenShellearly validation, not shipping (Red Hat, 05.2026)NVIDIA AGT · server sidecarsame policies as pattern AMS MCP gatewayKuadrant · Envoy ext_proc + brokerRED HATTech Preview MCP serversMCP catalogdev preview Business systemsAPIs, data AGT · ACTION GOVERNANCE (Microsoft, MIT) · proposed insertion, not in the blueprint PolicyCedar · Rego Intentmission, drift Approvalwebhook, elicitation AuditMerkle tree MaaSkeys, quotas · AuthorinoRED HATGA GuardrailsTrustyAI · NeMo GuardrailsRH + NVIDIA llm-d routingno model runs hereRH3.5 Vector storeMilvus · pgvector · QdrantRHGA Registry, catalogof modelsRHGA vLLM · the engineKServe; NIM optionalRH + NVIDIAGA OIDC or API key direct request to the agent Responses API JWT allowed? (server + tool) tool name + arguments RAG OTLP traces OBSERVABILITYOpenTelemetry Collector → MLflow Tracing · Tempo · Prometheus + DCGM · OpenShell OCSF events · MCP gateway "tool call" log PLATFORMRed Hat OpenShift · RHCOS · OVN-Kubernetes · GPU Operator and Network Operator (NVIDIA) · Node Feature Discovery · Kueue · ODF storage HARDWARENVIDIA GPUs; also AMD and Intel Gaudi · RoCE, InfiniBand · NVMe · on-premises, sovereign or cloud Maturity at the bottom of each box, as published by the vendor in September 2026; the inference chain is linearized: KServe drives llm-d, the guardrails are a called service.
The stack, top to bottom: users and the ingress gateway; the control band that deploys and confines agents; the two execution patterns, an agent in its own OpenShell-confined pod (the NemoClaw path) and the shared OGX server behind a Responses API (what OpenShift AI 3.5 ships); the two exits, tools on the right top, inference on the right bottom; observability, platform and hardware. Every box carries its maturity as the vendor publishes it. Red is Red Hat, green is NVIDIA, grey is community. The dashed blue block is not in the blueprint: it is where I would plug an action-governance toolkit, and part 4 of this series is about it.

The path of one request

  1. A user or an application authenticates with Keycloak and reaches the ingress gateway with a token or an API key.
  2. The request reaches the agent through one of two patterns. Pattern A, Agent-as-a-Workload: an agent in its own pod, confined by OpenShell, NVIDIA's sandbox (Landlock, seccomp, network namespaces, a Rego policy the agent may propose to change but cannot change alone); this is the NemoClaw path. Pattern B, Agent-as-a-Service: the shared OGX server (formerly Llama Stack) receives the request on its Responses API and runs the agent loop; this is what OpenShift AI 3.5 ships, and its OpenShell isolation is still an early validation, in Red Hat's own words. In both cases the control band gave the runtime a SPIFFE identity.
  3. To act, the agent calls a tool. The MCP gateway (Kuadrant, on Envoy) asks Authorino whether this identity may use this server and this tool, then routes to the MCP server. Authorino never sees the arguments of the call; the gateway can hand the tool name and arguments to an external guardrail service if you configure one.
  4. To reason, the agent calls a model. MaaS checks the key and the token quota, the content guardrails (TrustyAI, NeMo Guardrails) filter input and output, llm-d picks the right pod, and KServe serves the model with vLLM on GPU. NIM is optional.
  5. Everything emits traces to OpenTelemetry. OpenShell adds OCSF security events; the MCP gateway logs every tool call with the caller, the tool and the server.

Two design choices worth copying

The model is never consulted for authorization. Identity says who the agent is; token claims say which tools that identity may call. The gateway does not read the prompt. That is what makes prompt injection a contained problem instead of an existential one: the attacker can make the model want to call a tool, not make the gateway allow it.

Two engines, two jobs. vLLM is the inference engine: it loads the weights and computes the tokens. llm-d runs no model at all: it routes each request to the vLLM pod where the prefix cache is warm, and it is what lets you run two versions of a model side by side behind one name. Keep the two apart in your head and the whole inference plane becomes easy to reason about.

What Red Hat says it does not solve yet

Human approval workflows for high-risk actions still need product-level patterns.
Sandbox policy authoring is operationally complex at fleet scale.
Red Hat Developer, "Architect an open blueprint for cloud-native AI agents", 20 July 2026, section "what it does not solve yet". The same section adds: "no single vendor closes the list".

I checked both statements against the code. The MCP gateway's design document lists the protocol mechanism that would let a tool call pause for a human decision as a non-goal, deferred to a future iteration. The agent platform that grew out of Kagenti has a "Gateway Policies" screen that still reads "Coming Soon". None of this is a criticism: the blueprint is honest about it, which is rare. It simply tells you where the work is.

The words console, dashboard and UI do not appear in the blueprint either. The operator's experience today is YAML, the OpenShift console's event feed, and traces in MLflow. That is the third gap, and the one I find most interesting.

What this costs, and what is free

Every project in the diagram is open source: Apache 2.0 for the Red Hat and NVIDIA pieces, MIT for Llama Stack. What you pay Red Hat for is the assembly, the certified images and the support: OpenShift AI is the product, Open Data Hub is the same code without the subscription. What you may pay NVIDIA for is NIM and the CUDA-X libraries in production; vLLM does the serving job without them. Part 2 of this series rebuilds the same architecture from CNCF projects only, with no platform vendor at all.

Next in the series
  1. Part 2The same architecture in CNCF, NVIDIA and Microsoft projects, without a subscription.
  2. Part 3What the regulations actually ask for (EU AI Act, NIS2, ISO 42001, Swiss FADP and FINMA), and which layer of the stack has to answer.
  3. Part 4Action governance with Microsoft's Agent Governance Toolkit: how those requirements are enforced, and where it plugs into this stack.
  4. Part 5Observability: how it is proven. What OpenTelemetry, OCSF and MLflow record, and what they cannot tell you.
  5. Part 6Rolling out a model blue/green: how the stack changes without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.

Sources
  1. Architect an open blueprint for cloud-native AI agents, Red Hat Developer, 20 July 2026
  2. Architecting the Red Hat OpenShift AI dashboard for Models-as-a-Service, Red Hat Developer, 18 August 2026
  3. Repositories read on GitHub, September 2026: Kuadrant/mcp-gateway, Kuadrant/authorino, opendatahub-io/odh-dashboard, rossoctl/rossoctl, NVIDIA/OpenShell, NVIDIA-NeMo/Guardrails, NVIDIA/NeMo-Agent-Toolkit, llm-d/llm-d, kserve/kserve

Independent work, not affiliated with Red Hat, NVIDIA or Microsoft. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.