Blueprint · part 6 of 6
Rolling out a model blue/green: changing the stack without breaking what parts 3 to 5 established
A model is the component that changes most often and that the regulations watch most closely: new weights, a new quantization, a new inference engine. This last part shows how a new version enters production on the stack of parts 1 and 2, with the two properties the previous parts demand: nothing in the evidence chain breaks during the switch, and no human signs blind.
HokonokenSeptember 2026Reading time: 9 minViews are my own, not my employer's
Three things to take away
- Two pools behind one name. On this stack, a rollout is two llm-d InferencePools and one HTTPRoute whose weights move. The agents keep calling the same model name; the version behind it changes without any client touching its configuration.
- The gate is the regulation, run as an analysis. Article 15 asks for accuracy metrics; article 9 for testing "against prior defined metrics"; FINMA for backtests and adversarial tests before and after changes. Argo Rollouts can run exactly that as an AnalysisRun, from Prometheus, from garak's probes and from replayed test sets, before the weights move.
- A model is not a stateless service. Two copies cost two GPUs, loading takes minutes, and quality is not a status code. Blue stays warm for rollback, green is warmed before it sees a request, and the judge of a quality test is never the candidate itself.
Why blue/green and not a rolling update
A rolling update replaces pods one by one and assumes any pod can take any request. A model version is not interchangeable with the previous one: its answers differ, its refusals differ, its cost per token differs. What you want is a period where both versions run, the new one can be examined under real traffic without answering anyone, and the switch is a single reversible change. That is blue/green, and the inference extension of the Gateway API was designed for it: an InferencePool per version, an HTTPRoute with weights, and a rollback that is a weight set back to one hundred.
The route, in the stack's own vocabulary
The HTTPRoute that carries the switch
HTTPRoute · Gateway API inference extensionYAML
# Gateway API HTTPRoute in front of two llm-d InferencePools: the model name the agents
# call stays the same; only the weights move (llm-d "Blue-Green Update", gateway mode)
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: gpt-oss-120b
spec:
parentRefs:
- name: inference-gateway
rules:
- backendRefs:
- group: inference.networking.x-k8s.io
kind: InferencePool
name: vllm-gpt-oss-120b # blue, the version in production
weight: 90
- group: inference.networking.x-k8s.io
kind: InferencePool
name: vllm-gpt-oss-120b-new # green, the candidate
weight: 10
Shape from the llm-d "Blue-Green Update" procedure and the Gateway API inference extension: two InferencePools, one route, weights that sum to 100. llm-d documents this for its gateway mode and advises keeping the original pool and nodes during the rollout for rollback.
The analysis before promotion
Argo Rollouts, shipped with OpenShift GitOps and available on any Kubernetes, runs an AnalysisRun at the steps of a rollout and aborts it if a metric fails. Two providers matter here. The Prometheus provider reads what the stack already exports: vLLM latency and throughput, gateway error and refusal rates, DCGM GPU seconds. The Job provider runs a container, which is how replayed test sets fit in: a golden prompt set with reference answers, garak's adversarial probes fired at the candidate through the content guardrails, a set of recorded tool-call scenarios through the agent harness with the governance policies of part 4. Each returns pass or fail.
An AnalysisTemplate with both providers
AnalysisTemplate · Argo RolloutsYAML
# Argo Rollouts AnalysisTemplate: the checks a Rollout runs before it is allowed to move
# the weights. Two providers: prometheus for measurements, job for replayed test sets.
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: model-candidate-gate
spec:
args:
- name: candidate # the green pool
metrics:
- name: p95-latency-vs-blue
provider:
prometheus:
address: http://prometheus:9090
query: |
histogram_quantile(0.95, sum(rate(vllm:e2e_request_latency_seconds_bucket{pool="{{args.candidate}}"}[10m])) by (le))
/ histogram_quantile(0.95, sum(rate(vllm:e2e_request_latency_seconds_bucket{pool="blue"}[10m])) by (le))
successCondition: result[0] <= 1.10 # no worse than blue plus 10 %
failureLimit: 0
- name: golden-set-score
provider:
job:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: eval
image: registry.example.internal/model-eval:1.4 # signed with cosign
args: ["--target", "{{args.candidate}}", "--set", "golden-v12", "--judge", "not-the-candidate"]
successCondition: result == "pass"
failureLimit: 0
- name: adversarial-probes
provider:
job:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: garak
image: registry.example.internal/garak:0.17.0 # NVIDIA garak, Apache 2.0
env:
- name: OPENAICOMPATIBLE_API_KEY
valueFrom: { secretKeyRef: { name: gateway-eval-key, key: token } }
command: ["sh", "-ec"]
args:
- |
# target: the candidate pool, through the gateway, guardrails in front
python -m garak --target_type openai.OpenAICompatible \
--target_name gpt-oss-120b --report_prefix {{args.candidate}} \
--generator_options '{"openai": {"OpenAICompatible":
{"uri": "https://gateway.internal/pools/{{args.candidate}}/v1/"}}}' \
--spec probes.promptinject,probes.latentinjection,probes.encoding,\
probes.sysprompt_extraction,probes.leakreplay
# garak exits 0 with or without hits: the verdict is the diff against blue
compare-hits --candidate {{args.candidate}}.report.jsonl \
--baseline blue.report.jsonl --allow-new-hits 0
successCondition: result == "pass"
failureLimit: 0
Illustrative: the metric names follow vLLM's Prometheus exporter, the images, the set names and the compare-hits step are placeholders; the garak flags, probe families and environment variable are the tool's own. What matters is the shape: a criterion per check, written before the run, with failureLimit zero so that one failure aborts the rollout.
What the analysis must contain, and why
Measurements under mirrored traffic
p95 latency and time to first token, error and refusal rates, GPU seconds per request, compared with blue over the same window. Article 15 asks for "an appropriate level of accuracy, robustness and cybersecurity" maintained throughout the lifecycle; FINMA §2.4 asks for performance indicators "defined in advance".
A golden prompt set
A fixed set of prompts with reference answers, versioned with the policies, scored on green and blue with the same rubric. Article 9(6) asks for testing "against prior defined metrics and probabilistic thresholds". If a model scores the answers, it must not be the candidate: a model grading itself is not a test.
Adversarial probes, with garak
garak is NVIDIA's open-source LLM vulnerability scanner (Apache 2.0, v0.17.0 in September 2026): probe families for prompt injection, injections buried in documents, system-prompt extraction, encoding tricks, jailbreaks and training-data leakage, each with its detector, aimed at any OpenAI-compatible endpoint. Fired at green through the gateway, guardrails in front, it tests the pair that will be in production. It writes a JSONL report and a hit log; the run on blue is the baseline, and the criterion is zero new hits. This is article 15(5)'s "adversarial examples" and "model evasion", and FINMA's tests "before and after changes", run rather than assumed. The report stays with the rollout, for the auditor's table of part 5.
Policy set through the agent
Recorded tool-call scenarios run through the harness with green as the model and the same policies: the ones that must be refused are refused, the ones that must succeed still succeed. A model that refuses everything fails this check too. This is how a model change is checked against part 4 rather than around it.
Cost known before promotion
Tokens and GPU seconds per request from the shadow window, against the budget set for the model name. Not a regulatory row, but the one that gets a rollout reversed a week later if it is skipped.
A person promotes, and it is recorded
The Rollout pauses before the weights move; the six results and the diff against blue are in front of whoever resumes it. Article 14 asks for the ability to intervene and interrupt; article 26(2) for authority; article 11 for documentation kept up to date, so the model card changes in the same step.
The invariants during the switch
- The model name never changes. The agents, the MaaS keys and quotas, the policies that name the model: none of them move. The version is a pool label, not a client setting.
- The policies travel with the model. Green is tested with the policy set that will be in force when it is promoted, never with the old one, or the analysis proves nothing about production.
- The records of part 5 continue across the switch. Traces, the tool-call log, the verdicts and the OCSF events all carry which pool served the request. An auditor can tell, for any call in the six-month window, which model version answered it.
- Rollback is a weight, not a redeployment. Blue stays loaded for a day. Setting its weight back to one hundred takes seconds and needs no approval beyond the on-call engineer's, because it restores the version that was already approved.
Two things to check on your cluster, not on paper
Which serving mode you are on. The two-pool rollout above is llm-d in gateway mode. On classic KServe, the canary via canaryTrafficPercent is documented for serverless mode; in RawDeployment mode, doing the same through HTTPRoute weights is an open request in the KServe project as of September 2026. Open Data Hub and OpenShift AI can be configured either way, so this is the first thing to find out.
Whether your ingress can mirror traffic. The shadow phase assumes the gateway can copy requests to green without serving its answers. Envoy can; whether your chart exposes it is another matter. If it does not, replaying recorded prompts against green is a weaker but honest substitute, and it is what the Job provider runs anyway.
Where this leaves the series. Part 1 drew the stack, part 2 rebuilt it from open projects, part 3 read what the regulations ask, part 4 showed how a governance toolkit enforces it, part 5 what gets recorded and what cannot be proven, and this part how the stack changes without losing any of it. Every component named is public and open. What is not in any of them, and what every one of the six parts ran into, is the same three things: a surface where a human oversees and approves, the local regulatory regime as content, and proof to an outsider. Those are not products yet. They are where the work is.
The series
- Part 1AI agents on OpenShift: the open blueprint, layer by layer.
- Part 2The same architecture in CNCF, NVIDIA and Microsoft projects, without a subscription.
- Part 3What the regulations actually ask for (EU AI Act, NIS2, ISO 42001, Swiss FADP and FINMA), and which layer of the stack has to answer.
- Part 4Action governance with Microsoft's Agent Governance Toolkit: how those requirements are enforced, and where it plugs into this stack.
- Part 5Observability: what the stack records about an agent, and what it cannot prove.
- Part 6This part: rolling out a model blue/green without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.
Sources
- llm-d, Blue-Green Update (gateway mode; two InferencePools, HTTPRoute weights, keep the original pool for rollback)
- Gateway API Inference Extension, InferencePool and endpoint picker
- KServe canary rollout; kserve/kserve #5335, canary weights in RawDeployment mode
- Argo Rollouts in OpenShift GitOps; Argo Rollouts documentation, AnalysisTemplate, Prometheus and Job providers
- NVIDIA/garak, Apache 2.0, v0.17.0 (9 September 2026); probe families in
garak/probes/, the OpenAI-compatible target in garak/generators/openai.py, --spec and --generator_options in reference.garak.ai
- EU AI Act, article 9, article 15, article 14, article 11; FINMA Guidance 08/2024, section 2.4
Independent work, not affiliated with the CNCF, Red Hat, NVIDIA or Microsoft. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.