Blueprint · part 5 of 6

Observability: what the stack records about an agent, and what it cannot tell you

Parts 3 and 4 asked for evidence: logs kept six months, a person who approved, an action that stayed within its purpose. This part follows one tool call through the stack and looks at what actually gets written down, by which component, in which format, and where it ends up. Four records exist. Together they say what happened and whether it was allowed. None of them can prove it to someone who does not trust the operator.

HokonokenSeptember 2026Reading time: 7 minViews are my own, not my employer's

Three things to take away

  1. Telemetry says what happened. It does not say whether it should have. OpenTelemetry's GenAI conventions describe the model call, the tool execution, the tokens and the latency. They carry no verdict, no approver, no sensitivity label. That second record has to come from the governance layer, and it has to share the trace id.
  2. Four records, one trace id, six months. The trace, the gateway's tool-call log, the verdict with its approval, and the sandbox's OCSF events. If they are joined by the same trace id and kept for the retention period of Article 26(6), most audit questions have an answer in one query.
  3. A Merkle tree proves integrity to the operator, not to an outsider. It shows nothing was altered after the fact. It does not show, to a regulator or a customer, that the tree existed on the day it claims. That needs a witness the operator does not control, and no component in the stack provides one.

The four records

1. The trace. The agent harness or the OGX server emits OpenTelemetry spans following the GenAI semantic conventions: one span per model invocation, per tool execution, per agent step, with attributes for the model name, the token counts, the finish reason, optionally the content. This is the richest record and the only one that shows the reasoning chain as a tree. Two caveats. The conventions were still marked "Development" in 2026, so attribute names can change between library versions. And content capture is off by default in most instrumentations: turning it on stores prompts and answers, which is a data-protection decision, not a logging one.

2. The gateway's tool-call log. The MCP gateway of part 1 writes one structured log entry per tool call: the caller's subject from the JWT, the tool, the server, the HTTP status, a request id and an anonymised session id, flagged audit=true, with the trace id attached and exported over OTLP. This is the record that answers "who called what on which system", and it is produced by the choke point itself, not by the agent. It is not in OCSF format; it is a structured log.

3. The verdict and the approval. This is the record telemetry does not produce. The governance toolkit of part 4 writes, for each action, the rule that fired, the verdict, and when a human was involved, who approved, with the digest of the exact action they saw. The entries go into a Merkle tree with a root hash and inclusion proofs, so an auditor can check that a given entry was present and unchanged. It arrives by its own path, not through the collector, and it has to carry the same trace id as records 1 and 2 or it cannot be joined to them.

4. The sandbox's security events. Below the tool call, the OpenShell supervisor emits OCSF events for what the process actually did: files opened, connections attempted, processes spawned, and the policy decision on each. OCSF is the schema SIEMs already ingest, so this record lands where the security team already looks. It answers a question the other three cannot: whether the code behind the tool touched anything it was not supposed to.

One tool callagent → gateway → servertrace id: 4bf9… FOUR RECORDS, ONE TRACE ID 1 · OpenTelemetry tracegen_ai.* spans: model call, tool execution, agent stepwhat happened, when, how long, how many tokensharness or OGX server · conventions "Development" 2 · Gateway tool-call logaudit=true · user · tool · server · status · request idwho called which tool on which serverstructured log with trace id, exported over OTLP 4 · Sandbox security eventsOCSF: process, file, network decisions of the sandboxwhat the process actually touched, below the tool callfrom the OpenShell supervisor 3 · Verdict and approval recordrule that fired · allow, deny, escalate · approvershould it have happened, and who said soits own path, same trace id OpenTelemetryCollectortraces, metrics, logsone pipeline, one trace idjoins records 1, 2, 4 Audit treeMerkle root + inclusion proofsnothing altered after the factproves it to the operator only STORES · KEPT ≥ 6 MONTHS (AI ACT ART. 26(6)) Jaeger or Tempotraces, queried by trace id Prometheus + DCGMlatency, errors, refusals, GPU seconds MLflow Tracingagent runs as experiments, approval status visible SIEM or log storethe tool-call log and OCSF events, six months Backed-up roots, six monthsthe operator keeps the root hashes with the entriesan auditor verifies inclusion against the rootbut the root is the operator's word MISSING · PROOF TO AN OUTSIDERa witness the operator does not control:transparency log, signed timestamp, receipt WHAT EACH RECORD ANSWERS1 what happened and when · 2 who did it, on what · 4 what the process really touched · 3 whether it was allowed and who approvednone of the four, alone or together, answers "prove it to someone who does not trust you": that needs a witness outside the operator's control
One tool call, four records. Records 1, 2 and 4 flow through the OpenTelemetry collector into the trace store, Prometheus, MLflow and the SIEM; record 3 goes to the audit tree by its own path, and its roots are kept with the rest. All must share the trace id and be kept at least six months. The orange box is what no component provides: a witness outside the operator's control.

What an auditor asks, and which record answers

The questionThe record that answers itThe gap
What did the agent do on 12 March at 14:07?Trace (1), queried by trace id in Jaeger or Tempo; tool-call log (2) for the systems touchedOnly if the trace store kept it: default retention of trace backends is days, not months
Who was accountable for that call?Tool-call log (2): the subject in the JWT the gateway validatedThe subject is the agent's identity; the human on whose behalf it acted is only there if the token carries it
Was it allowed, and by which rule?Verdict record (3): rule, verdict, timestampOnly for calls that went through the toolkit; a call that bypassed it leaves records 1, 2, 4 and no verdict
Who approved it, and what exactly did they approve?Verdict record (3): approver and the digest of the actionNone, if the approval echoed the digest; the toolkit refuses approvals that do not
Did the code do anything beyond the tool call?OCSF events (4) from the sandboxOnly for pattern A; the shared OGX server's isolation was still an early validation in 2026
Has this log been altered since?Audit tree (3): root hash and inclusion proofProves it to whoever holds the root; the operator holds the root
Was the model behind this call tested against adversarial inputs before it went live?The garak report kept with the rollout (part 6): a JSONL record per probe and a hit log, next to the AnalysisRun resultOnly if the rollout kept it; garak probes the model through the guardrails, not the agent's tools, and it is a test record, not a record of the call
Prove to me, an outsider, that this existed on that dayNothing in the stackNeeds a witness: a transparency log, a signed timestamp, a receipt from a party the operator does not control

Where the records are switched on

The files that create record 3, the manifest with its intervention points, the approval section, the intent, and the sector templates, are shown one by one in part 4. Two settings belong here instead, because they decide where the records go rather than what is decided. Both are read from the repository at commit 2e98345 of 21 September 2026.

Where the trace goes
agent-governance-python/agent-os/examples/pharma-compliance/docker-compose.ymlYAML
# The demos point the SDK's OpenTelemetry exporter at a collector by environment variable
environment:
  - OTEL_EXPORTER_OTLP_ENDPOINT=http://jaeger:4317

# The agent-sre Helm chart exposes the same thing as values
otel:
  enabled: true
  endpoint: http://otel-collector:4317
  serviceName: agent-sre
agent-governance-python/agent-sre/deployments/helm/agent-sre/values.yaml (enabled is false by default)
How the audit tree leaves the process
agent-governance-python/agent-mesh/src/agentmesh/governance/audit.pyPython · signatures, bodies shortened
# Two exports on the audit log: a JSON document that carries the Merkle root with the
# entries, and the same entries as CloudEvents v1.0 envelopes for a bus or a SIEM
def export(self, start_time=None, end_time=None) -> dict:
    entries = self.query(start_time=start_time, end_time=end_time, limit=10000)
    return {
        "exported_at": ...,                          # UTC timestamp of the export
        "merkle_root": self._chain.get_root_hash(),
        "entry_count": len(entries),
        "entries": [e.model_dump() for e in entries],
    }

def export_cloudevents(self, start_time=None, end_time=None) -> list[dict]:
    return [e.to_cloudevent() for e in entries]
The root travels with the export, and that is what an auditor checks the inclusion proofs against. What the export does not contain is anything signed by a party other than the operator, which is the subject of the last section.

Both defaults are conservative. The OpenTelemetry exporter in the Helm chart is off until an endpoint is given, so record 1 does not exist until someone decides where it lives. And the sample policy of part 4 keeps tool arguments out of the audit, include_tool_args: false, because arguments carry personal data: record 3 identifies the action by its digest, not by its content.

The retention problem nobody budgets for

Article 26(6) asks deployers to keep the automatically generated logs for at least six months. Trace backends are built for debugging, and their default retention is measured in days; Prometheus in weeks. Keeping six months of GenAI spans, with or without content, is a storage decision that has to be taken on purpose: a cold tier for traces, the tool-call log and OCSF events shipped to the SIEM the organisation already retains, and the audit tree backed up with its root hashes. The stack produces the records. Keeping them is the operator's job, and it is where the first audit usually finds the first gap.

What a Merkle tree proves, and what it does not

The toolkit's audit tree lets anyone with the root hash verify that an entry belongs to the tree and has not changed. That is real: an operator cannot quietly edit an entry after the fact without changing the root. What the tree does not do is prove to a third party when the root was produced, or that the tree the auditor is shown is the one that existed at the time. The operator holds the root; the operator could regenerate the whole tree. The toolkit's own code says it does not anchor to a transparency log. This is not a flaw specific to that toolkit; it is the limit of any integrity mechanism that stays inside one organisation.

The way out is known from software supply chains: publish the root, or a signed statement about it, to a log the operator does not control, and keep the receipt. Sigstore's Rekor already does this for build artefacts, the IETF's SCITT working group is standardising the receipt format, and RFC 3161 timestamps have done a simpler version of it for twenty years. None of these is wired into the agent stack of parts 1 and 2. For a regulated deployment, it is the last piece of the evidence chain, and the one to plan for before the first audit rather than after.

Net effect on the matrix of part 3. With the four records joined by trace id and retained, the logging, monitoring and incident rows are answered by the stack as it stands. The oversight row is answered by record 3 from part 4. What stays open, across every framework, is proof to an outsider; the frameworks do not name it yet, and the auditors already ask for it.

Next in the series
  1. Part 6Rolling out a model blue/green: how the stack changes without breaking what parts 3 to 5 established.
Read, not run. Everything in this series comes from reading public code and documents at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: these projects move monthly, their bugs move with them, and a component marked preview or alpha here may be stable, or gone, by the time you read this. Test it on your own cluster. When something does not match, file the issue in the project's tracker and send the fix back: that is how open code improves, and it is the only way a map like this one stays true.

Sources
  1. OpenTelemetry Semantic Conventions for Generative AI, status "Development" in 2026; CNCF graduation of OpenTelemetry, May 2026
  2. Kuadrant/mcp-gateway, tool-call audit log in internal/mcp-router/ext_proc_adapter.go, OTLP export in internal/otel/logging.go, read at commit f77d958
  3. microsoft/agent-governance-toolkit, agentmesh/governance/audit.py (Merkle tree), agentmesh/governance/trace_sink.py (no transparency log anchoring), read at commit e7f5d2b
  4. NVIDIA/OpenShell, crate openshell-ocsf; OCSF schema
  5. Constraining AI agents with Red Hat AI, Red Hat Developer, 16 September 2026, for MLflow as the audit trail a reviewer pulls
  6. EU AI Act, Article 26(6), log retention of at least six months
  7. Sigstore Rekor; IETF SCITT working group; RFC 3161, time-stamp protocol

Independent work, not affiliated with the CNCF, Red Hat, NVIDIA or Microsoft. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.