Blueprint · part 6 of 6

Rolling out a model blue/green: changing the stack without breaking what parts 3 to 5 established

A model is the component that changes most often and that the regulations watch most closely: new weights, a new quantization, a new inference engine. This last part shows how a new version enters production on the stack of parts 1 and 2, with the two properties the previous parts demand: nothing in the evidence chain breaks during the switch, and no human signs blind.

HokonokenSeptember 2026Reading time: 9 minViews are my own, not my employer's

Three things to take away

  1. Two pools behind one name. On this stack, a rollout is two llm-d InferencePools and one HTTPRoute whose weights move. The agents keep calling the same model name; the version behind it changes without any client touching its configuration.
  2. The gate is the regulation, run as an analysis. Article 15 asks for accuracy metrics; article 9 for testing "against prior defined metrics"; FINMA for backtests and adversarial tests before and after changes. Argo Rollouts can run exactly that as an AnalysisRun, from Prometheus, from garak's probes and from replayed test sets, before the weights move.
  3. A model is not a stateless service. Two copies cost two GPUs, loading takes minutes, and quality is not a status code. Blue stays warm for rollback, green is warmed before it sees a request, and the judge of a quality test is never the candidate itself.

Why blue/green and not a rolling update

A rolling update replaces pods one by one and assumes any pod can take any request. A model version is not interchangeable with the previous one: its answers differ, its refusals differ, its cost per token differs. What you want is a period where both versions run, the new one can be examined under real traffic without answering anyone, and the switch is a single reversible change. That is blue/green, and the inference extension of the Gateway API was designed for it: an InferencePool per version, an HTTPRoute with weights, and a rollback that is a weight set back to one hundred.

Agents · gatewaycall "gpt-oss-120b"the name never changes Inference gatewayGateway API · inference extensionHTTPRoute: blue / green weightsllm-d picks the podCNCF BLUE pool · current versionllm-d InferencePool · vLLM pods on GPUweight 100 → 90 → 50 → 0stays warm 24 h after the switch GREEN pool · candidatellm-d InferencePool · vLLM pods on GPUnew weights, quantization or vLLMwarmed up · weight 0 → 10 → 50 → 100 GATE · ARGO ROLLOUTS ANALYSISPrometheus: p95, errors, refusals, GPU secondsJobs: golden prompts, garak probes, policy setcriteria written first (art. 9, 15; FINMA §2.4)any failure: rollout aborted, blue untoucheda person promotes; approval recorded (art. 14)and the model card is updated (art. 11) Model registry · Git · Argo CDthe candidate is declared, image signed,Argo CD creates the green pool request blue green shadow: mirrored traffic, answers never served measured deploys promotion: Argo Rollouts moves the HTTPRoute weights in steps · rollback = blue back to 100, in seconds Products: Envoy Gateway or Kuadrant + Gateway API inference extension (route, weights) · llm-d (pools, pod choice) · vLLM (engine) · KServe (lifecycle)· Kubeflow or OpenShift AI model registry · Argo CD and Argo Rollouts · cosign · Prometheus, DCGM · NeMo Guardrails, Agent Governance Toolkit (replayed sets)
Two llm-d pools behind one HTTPRoute. Green starts at zero, receives mirrored traffic, passes the analysis, then takes requests in steps while blue stays warm. The orange box is the gate: measurements and replayed sets, criteria written before the run, a person who promotes and whose approval is recorded. Mechanism documented by llm-d for gateway mode; on classic KServe, the equivalent is the canary via canaryTrafficPercent, in serverless mode only.

The route, in the stack's own vocabulary

The HTTPRoute that carries the switch
HTTPRoute · Gateway API inference extensionYAML
# Gateway API HTTPRoute in front of two llm-d InferencePools: the model name the agents
# call stays the same; only the weights move (llm-d "Blue-Green Update", gateway mode)
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: gpt-oss-120b
spec:
  parentRefs:
    - name: inference-gateway
  rules:
    - backendRefs:
        - group: inference.networking.x-k8s.io
          kind: InferencePool
          name: vllm-gpt-oss-120b          # blue, the version in production
          weight: 90
        - group: inference.networking.x-k8s.io
          kind: InferencePool
          name: vllm-gpt-oss-120b-new      # green, the candidate
          weight: 10
Shape from the llm-d "Blue-Green Update" procedure and the Gateway API inference extension: two InferencePools, one route, weights that sum to 100. llm-d documents this for its gateway mode and advises keeping the original pool and nodes during the rollout for rollback.

The analysis before promotion

Argo Rollouts, shipped with OpenShift GitOps and available on any Kubernetes, runs an AnalysisRun at the steps of a rollout and aborts it if a metric fails. Two providers matter here. The Prometheus provider reads what the stack already exports: vLLM latency and throughput, gateway error and refusal rates, DCGM GPU seconds. The Job provider runs a container, which is how replayed test sets fit in: a golden prompt set with reference answers, garak's adversarial probes fired at the candidate through the content guardrails, a set of recorded tool-call scenarios through the agent harness with the governance policies of part 4. Each returns pass or fail.

An AnalysisTemplate with both providers
AnalysisTemplate · Argo RolloutsYAML
# Argo Rollouts AnalysisTemplate: the checks a Rollout runs before it is allowed to move
# the weights. Two providers: prometheus for measurements, job for replayed test sets.
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: model-candidate-gate
spec:
  args:
    - name: candidate                       # the green pool
  metrics:
    - name: p95-latency-vs-blue
      provider:
        prometheus:
          address: http://prometheus:9090
          query: |
            histogram_quantile(0.95, sum(rate(vllm:e2e_request_latency_seconds_bucket{pool="{{args.candidate}}"}[10m])) by (le))
            / histogram_quantile(0.95, sum(rate(vllm:e2e_request_latency_seconds_bucket{pool="blue"}[10m])) by (le))
      successCondition: result[0] <= 1.10   # no worse than blue plus 10 %
      failureLimit: 0
    - name: golden-set-score
      provider:
        job:
          spec:
            template:
              spec:
                restartPolicy: Never
                containers:
                  - name: eval
                    image: registry.example.internal/model-eval:1.4   # signed with cosign
                    args: ["--target", "{{args.candidate}}", "--set", "golden-v12", "--judge", "not-the-candidate"]
      successCondition: result == "pass"
      failureLimit: 0
    - name: adversarial-probes
      provider:
        job:
          spec:
            template:
              spec:
                restartPolicy: Never
                containers:
                  - name: garak
                    image: registry.example.internal/garak:0.17.0   # NVIDIA garak, Apache 2.0
                    env:
                      - name: OPENAICOMPATIBLE_API_KEY
                        valueFrom: { secretKeyRef: { name: gateway-eval-key, key: token } }
                    command: ["sh", "-ec"]
                    args:
                      - |
                        # target: the candidate pool, through the gateway, guardrails in front
                        python -m garak --target_type openai.OpenAICompatible \
                          --target_name gpt-oss-120b --report_prefix {{args.candidate}} \
                          --generator_options '{"openai": {"OpenAICompatible":
                            {"uri": "https://gateway.internal/pools/{{args.candidate}}/v1/"}}}' \
                          --spec probes.promptinject,probes.latentinjection,probes.encoding,\
                            probes.sysprompt_extraction,probes.leakreplay
                        # garak exits 0 with or without hits: the verdict is the diff against blue
                        compare-hits --candidate {{args.candidate}}.report.jsonl \
                          --baseline blue.report.jsonl --allow-new-hits 0
      successCondition: result == "pass"
      failureLimit: 0
    
Illustrative: the metric names follow vLLM's Prometheus exporter, the images, the set names and the compare-hits step are placeholders; the garak flags, probe families and environment variable are the tool's own. What matters is the shape: a criterion per check, written before the run, with failureLimit zero so that one failure aborts the rollout.

What the analysis must contain, and why

Measurements under mirrored traffic

p95 latency and time to first token, error and refusal rates, GPU seconds per request, compared with blue over the same window. Article 15 asks for "an appropriate level of accuracy, robustness and cybersecurity" maintained throughout the lifecycle; FINMA §2.4 asks for performance indicators "defined in advance".

A golden prompt set

A fixed set of prompts with reference answers, versioned with the policies, scored on green and blue with the same rubric. Article 9(6) asks for testing "against prior defined metrics and probabilistic thresholds". If a model scores the answers, it must not be the candidate: a model grading itself is not a test.

Adversarial probes, with garak

garak is NVIDIA's open-source LLM vulnerability scanner (Apache 2.0, v0.17.0 in September 2026): probe families for prompt injection, injections buried in documents, system-prompt extraction, encoding tricks, jailbreaks and training-data leakage, each with its detector, aimed at any OpenAI-compatible endpoint. Fired at green through the gateway, guardrails in front, it tests the pair that will be in production. It writes a JSONL report and a hit log; the run on blue is the baseline, and the criterion is zero new hits. This is article 15(5)'s "adversarial examples" and "model evasion", and FINMA's tests "before and after changes", run rather than assumed. The report stays with the rollout, for the auditor's table of part 5.

Policy set through the agent

Recorded tool-call scenarios run through the harness with green as the model and the same policies: the ones that must be refused are refused, the ones that must succeed still succeed. A model that refuses everything fails this check too. This is how a model change is checked against part 4 rather than around it.

Cost known before promotion

Tokens and GPU seconds per request from the shadow window, against the budget set for the model name. Not a regulatory row, but the one that gets a rollout reversed a week later if it is skipped.

A person promotes, and it is recorded

The Rollout pauses before the weights move; the six results and the diff against blue are in front of whoever resumes it. Article 14 asks for the ability to intervene and interrupt; article 26(2) for authority; article 11 for documentation kept up to date, so the model card changes in the same step.

The invariants during the switch

Two things to check on your cluster, not on paper

Which serving mode you are on. The two-pool rollout above is llm-d in gateway mode. On classic KServe, the canary via canaryTrafficPercent is documented for serverless mode; in RawDeployment mode, doing the same through HTTPRoute weights is an open request in the KServe project as of September 2026. Open Data Hub and OpenShift AI can be configured either way, so this is the first thing to find out.

Whether your ingress can mirror traffic. The shadow phase assumes the gateway can copy requests to green without serving its answers. Envoy can; whether your chart exposes it is another matter. If it does not, replaying recorded prompts against green is a weaker but honest substitute, and it is what the Job provider runs anyway.

Where this leaves the series. Part 1 drew the stack, part 2 rebuilt it from open projects, part 3 read what the regulations ask, part 4 showed how a governance toolkit enforces it, part 5 what gets recorded and what cannot be proven, and this part how the stack changes without losing any of it. Every component named is public and open. What is not in any of them, and what every one of the six parts ran into, is the same three things: a surface where a human oversees and approves, the local regulatory regime as content, and proof to an outsider. Those are not products yet. They are where the work is.

The series
  1. Part 1AI agents on OpenShift: the open blueprint, layer by layer.
  2. Part 2The same architecture in CNCF, NVIDIA and Microsoft projects, without a subscription.
  3. Part 3What the regulations actually ask for (EU AI Act, NIS2, ISO 42001, Swiss FADP and FINMA), and which layer of the stack has to answer.
  4. Part 4Action governance with Microsoft's Agent Governance Toolkit: how those requirements are enforced, and where it plugs into this stack.
  5. Part 5Observability: what the stack records about an agent, and what it cannot prove.
  6. Part 6This part: rolling out a model blue/green without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.

Sources
  1. llm-d, Blue-Green Update (gateway mode; two InferencePools, HTTPRoute weights, keep the original pool for rollback)
  2. Gateway API Inference Extension, InferencePool and endpoint picker
  3. KServe canary rollout; kserve/kserve #5335, canary weights in RawDeployment mode
  4. Argo Rollouts in OpenShift GitOps; Argo Rollouts documentation, AnalysisTemplate, Prometheus and Job providers
  5. NVIDIA/garak, Apache 2.0, v0.17.0 (9 September 2026); probe families in garak/probes/, the OpenAI-compatible target in garak/generators/openai.py, --spec and --generator_options in reference.garak.ai
  6. EU AI Act, article 9, article 15, article 14, article 11; FINMA Guidance 08/2024, section 2.4

Independent work, not affiliated with the CNCF, Red Hat, NVIDIA or Microsoft. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.