Which AI for which job · part 5 of 6

Acting within a perimeter: the AI agent with its tools, and the augmented workflow as the migration path

Part 4 ended on a question: when the worst action is no longer a disclosure but a change in a system, what has to be true before the model is allowed to make it? This row is where a language model's output is an action and a system carries it out. Two ways to build it: in the augmented workflow, an engine decides the path and a model fills one step of it; in the AI agent, the model decides the path, inside a perimeter that someone else drew. Everything the first series built exists for this row. What the row adds is the one thing that series did not have to decide: which actions a model may choose, and where a person stands.

HokonokenSeptember 2026Reading time: 20 minNot legal advice · Views are my own, not my employer's

Three things to take away

  1. The perimeter is not the prompt. A perimeter is a list of tools, an identity, a policy per action, an approval at the commit and a record of every call, all outside the model. OWASP's 2025 list names the failure "Excessive Agency" and gives it three causes: excessive functionality, excessive permissions, excessive autonomy. An instruction in the prompt addresses none of them.
  2. The augmented workflow is the migration path. Anthropic's 2024 line separates workflows, "where LLMs and tools are orchestrated through predefined code paths", from agents, "where LLMs dynamically direct their own processes and tool usage". Most of what is sold to an administration or a bank as an agent in 2026 is a workflow with one model step, and that is the right place to start: the engine keeps the path, the record and the retry, and the model cannot reach a tool the step does not offer.
  3. Read, write, commit is a label you keep, not one you receive. The Model Context Protocol lets a tool describe itself as read-only, destructive or idempotent, defaults to destructive, and says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers". The classification that decides where the approval goes is made at your gateway, per tool, by you. And the approval belongs at the commit.

What changed since part 4

Agentic RAG was already an agent: it chose what to fetch, went through a tool gateway, ran as the user, and every query was recorded, so that a read would stay a read. This row lets the same machinery write. One thing does not change: the model has no permissions of its own. It inherits them, from the person it acts for and from the identity the platform gave it, and can do nothing the perimeter does not offer. This whole row rests on that sentence.

Two things change. The worst action is now a change in a system that will act on it: a record in the case system becomes a decision the next process reads, a payment posted is money moved. A write can be undone by another write. A commit cannot be quietly undone: reversal is a new action, with its own consequences and its own approval. That is why the map places a job at level 3 or 4 by where the commit sits, not by which family the vendor named. And there may be no person reading. In the content and sources rows a person read the output before anything happened, and that reading was the control. Here the output is the action, and if the design puts the person after it, the record is the only witness.

OWASP's mitigations for "Excessive Agency" are the perimeter, item by item: "minimize extensions", "minimize extension permissions", "execute extensions in user's context", "require user approval", "complete mediation". The same page adds that logging and rate limits "can help limit damage but won't prevent the vulnerability". A record tells you what happened. It does not stop it.

Two ways to let a model act

The augmented workflow. A process engine, a BPM tool, a durable-execution runtime, or the robotic process automation many administrations already run, keeps the path: which step follows which, what is retried, what is compensated. One step, sometimes two, calls a model. The step is a function: it takes a document or a record and returns a typed result, extracted fields, a class, a draft, which the engine validates against a schema before it decides the next step. The model never sees a tool. It cannot reach the case system, the ledger or the mailbox, because the step does not offer them. The person stays where the process already put them, at the approval step, and the record is the engine's.

That is why this row is the migration path. One step changes, nothing else does, and the engine's own records show, before anyone moves the step to an agent, how often the model's output was accepted, corrected or rejected. Anthropic's 2024 note lists five workflow patterns, "prompt chaining", "routing", "parallelization", "orchestrator-workers" and "evaluator-optimizer", and every one is a path decided in code; only the fourth becomes something else, when the workers are given tools, and that is part 6. The engines are older than the models, which is their merit: Temporal, MIT, v1.32.0 of 11 September 2026; Argo Workflows, Apache-2.0, v4.1.4 of 18 September 2026, already in the first series; Robot Framework, Apache-2.0, v7.5 of 14 September 2026, for the RPA case where the "tool" is a screen. None of them knows what a language model is.

The AI agent. The same note describes the other case: "open-ended problems where it's difficult or impossible to predict the required number of steps", where the model runs a loop, picks a tool, reads the result, and continues until a stopping condition. It is plain about the price: agents require "some level of trust in its decision-making", and are to be run with "checkpoints" and "stopping conditions". The path is now the model's. What replaces the engine is a perimeter, outside the model, that says which tools exist, who the agent is, what each action may do, when a person must say yes, and what is written down.

The protocol most agents use to reach their tools has said the same since part 1 quoted it, and the current revision, 2026-07-28, keeps the sentences unchanged: tools are "model-controlled"; "there SHOULD always be a human in the loop with the ability to deny tool invocations"; and "MCP itself cannot enforce these security principles at the protocol level". The revision made the protocol stateless and moved long-running tasks into an extension; the state that matters here, the budget, the stop, the approval, lives in your runtime and your gateway.

AUGMENTED WORKFLOW · THE ENGINE DECIDES THE PATH Triggera case, an invoice, a mailENGINE STARTS A RUN Step 1 · rulefetch, validate, routeDETERMINISTIC Step 2 · the model stepextract, classify, draft: a functionTYPED OUTPUT · VALIDATED · PART 3 Step 3 · rulethreshold, four eyesDETERMINISTIC Approvala person, before the commitART. 14(4)(D) · LEVEL 3 System of recordpost, notify, payLEVEL 4 IF A RULE APPROVES COMMIT The model never sees a tool. The engine calls it as a function with a schema, checks the result against the schema, and decides the next step itself. Retry, timeout, compensation and the record belong to the engine, which is why this row is the migration path: one step of the process changes, nothing else does. Level 3 while a person approves the commit; level 4 the day a rule does. The gate of the first series reruns on every change of the model step. S1 · PART 5 (RECORDS) · PART 6 (ROLLOUT OF THE MODEL STEP) · ENGINES: TEMPORAL, ARGO WORKFLOWS; RPA: ROBOT FRAMEWORK AI AGENT · THE MODEL DECIDES THE PATH, INSIDE A PERIMETER SOMEONE ELSE DREW A taskfrom a person or a queue,IDENTITY TRAVELS with the person's identity The agenta loop: think, pick a tool, read theresult, repeat until a stopping conditionLEVEL 3 · 4OWN IDENTITY · BUDGET · A STOP Tool gatewaythe perimeter, outside the model identity checked: the agent's, the person's allowlist: the agent sees nothing else policy per action: ceiling, scope, four eyes a commit is held; digest of the exact call one log entry per call: who, what, result S1 · PART 1 · PART 4 · PART 5 MCP · EXPLICIT USER CONSENT BEFORE ANY TOOL Read a recordreadOnlyHint: true READ nothing changes; part 4 governed it Write a recordnot destructive · idempotent WRITE another write undoes it; safe to retry Commitpay, send, revoke, open world COMMIT reversal is a new action, with its own approval The label is yours. The protocol lets a tool describe itself, defaults to destructive, then says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers". The person with authoritycan deny, reverse, or stop the loopART. 14(4)(D)(E) · 26(2) approve or deny, at the commit LEVELS FROM PART 1: 3 ACTS UNDER APPROVAL · 4 ACTS, REVIEWED AFTER
the engine, deterministicthe model step, a functionthe agent, the commit, and the person who can stop itLevels from part 1: 3 acts under approval · 4 acts, reviewed after
Top: the augmented workflow. The engine decides the path; the model fills one step and returns a typed value; the approval sits where the process always had it, before the commit. Bottom: the AI agent. The model decides the path, and every tool call crosses the gateway, which is the perimeter: identity, allowlist, policy per action, a hold at the commit, one log entry per call. The three tools carry the label the gateway gave them, not the one the server claims.

The perimeter, in five parts

Each part is a control the first series built. What this row adds is the reason it is there.

The list of tools. The agent sees the tools on its allowlist and nothing else. The current protocol revision notes that a server's tool list "MAY vary by the authorization presented on the request"; that is the server's half. The gateway's half is the allowlist per agent that the first series placed in front of every MCP server. An agent cannot misuse a tool it was never shown.

Identity, twice. The agent has an identity of its own, a workload identity from SPIRE in the first series, and it acts for a person whose identity travels with the task. The gateway checks both. A shared service account is the failure this prevents: the record then says "the agent" wrote the record, and the auditor's question, which agent, for whom, has no answer.

A policy per action. The ceiling is a rule, not a prompt; part 1 said it for the spare-parts agent and this is the mechanism. The gateway evaluates a policy on every call, with the tool, the arguments and both identities as input, and returns allow, deny or hold. The first series showed two ways: a policy engine at the gateway, OPA, Apache-2.0, v1.20.2 of 3 September 2026, and Microsoft's Agent Governance Toolkit, MIT, whose manifest lists the actions that go to a person before execution and bounds arguments by schema, a maximum amount on a transfer, for instance. The toolkit's detail that matters most here: an approval must echo the digest of the exact action, or it counts as a denial.

An approval at the commit. Not at the family, not at the session, at the commit. The person who gives it needs, in the words of Article 26(2), "the necessary competence, training and authority", and must see the exact action, its arguments and its target. Article 14(4)(b)'s "automation bias" applies to approvals as much as to proposals: an agent that asks fifty times an hour is approved by habit, and the toolkit's fatigue threshold exists for that. The hold itself needs a place to live: the first series found it missing in one gateway's design document and present in the toolkit's runtime. It is a component you choose, not a property of the protocol.

A record per call. The gateway writes one entry per tool call, with the caller, the tool, the server, the status and a request id, as the first series described; the current protocol revision documents how the trace context travels in the call's metadata, so the call joins the trace of the turn that caused it. Article 12(1) asks that a high-risk system "technically allow for the automatic recording of events (logs) over the lifetime of the system"; Article 26(6) asks the deployer to keep them at least six months. For a system that acts, the log is also the reversal path.

Read, write, commit: the labels

Part 1 gave the reversibility axis three values. The protocol gives a tool four ways to describe itself, and the defaults are worth reading: a tool that says nothing about itself is, by default, not read-only, "may perform destructive updates to its environment", is not idempotent, and "may interact with an 'open world' of external entities". Assume the worst, chosen on purpose. Then the specification takes the description away from you: "clients MUST consider tool annotations to be untrusted unless they come from trusted servers". A tool that calls itself read-only is a claim, made by the server, about itself. So the classification the map needs is made by you, at the gateway, per tool, from what the tool actually does. A tool you have verified changes nothing is a read. A tool whose changes another tool can undo, and which is safe to retry, is a write. A tool that pays, sends, revokes, or reaches outside, mail, a payment network, the web, is a commit, because what leaves cannot be recalled. Article 14(4)(d) asks that the person can "reverse the output"; for a commit, that means the reversal tool is on the list too, with its own policy, or the clause has no mechanism behind it.

What the platform owes them

Take part 4's platform, whose gateway, identity and record were built for a read, and give them a write to govern.

Where the law reaches this row

Through Article 14, in full. The earlier rows met the oversight article with a person reading the output or, in part 2's level 4 wirings, with an override on a score. This row meets it, or fails to, with product features. Paragraph 1 asks that a high-risk system "can be effectively overseen by natural persons during the period in which they are in use". Paragraph 4 lists what the person must be able to do: understand its "capacities and limitations" and "detect anomalies"; interpret the output; "disregard, override or reverse the output"; "intervene in the operation" or "interrupt the system through a 'stop' button or a similar procedure". For an agent at level 4 those verbs are a hold, a reversal tool, a stop and a log, and they either exist or they do not. Paragraph 3 says the measures are built in by the provider or identified by the provider for the deployer, and Article 25(1) makes a deployer the provider when it substantially modifies a high-risk system or changes a system's purpose so that it becomes one. On that reading, an agent assembled in-house from a model, a gateway and tools, doing an Annex III job, has for provider the organisation that wired it. Counsel should confirm that before anyone relies on it.

Through the deployer's obligations. Article 26 asks the deployer to "assign human oversight to natural persons who have the necessary competence, training and authority", to "monitor the operation", to keep the logs, and, in paragraph 11, to "inform the natural persons that they are subject to the use of the high-risk AI system". Each maps to a component above; the first maps to a job description.

Through incident reporting. Article 73 asks providers to report a serious incident "not later than 15 days" after becoming aware of it, two days for the widespread cases. Article 3(49) defines one as leading, among other things, to "the infringement of obligations under Union law intended to protect fundamental rights". An agent that wrongly revokes a benefit, at scale, before anyone reads the log, is not far from that definition. The report will be written from the record.

Through data protection at level 4. GDPR Article 22 and Article 21 of the Swiss FADP apply where the agent's commit is a decision about a person with legal or similar effect, taken without a person. The level, not the family, triggers them.

Through security law and the job. NIS2 Article 21(2) asks for access control and supply-chain security; the tools an agent calls and the servers that expose them are suppliers, and the gateway's allowlist is the access-control policy the article has in mind. The Swiss financial regulator's guidance asks that results can be "understood, explained or reproduced", as read in the first series; for an agent, that is the record. And nothing in Annex III says "agent": the row inherits the risk class of the task.

Three organisations, nine jobs

The same three organisations; three of the jobs are the ones part 1 placed on this row. Each is placed by who decides the path and where the commit sits.

JobFamily, levelWorst actionWhat the platform must provideRegulatory line
A public agency renews a recurring benefit: a workflow extracts the figures from the new income document, checks them against the rules, and a caseworker approves the renewal letter.Augmented workflow, 3commit, by the caseworkerThe engine owns the path and the record; the model step returns typed fields the engine validates; the letter comes from a template; the gate reruns on every change of the model step.Annex III 5(a): systems used "to grant, reduce, revoke, or reclaim" benefits. Article 14(4)(b) on automation bias for the caseworker. Article 21 FADP not triggered while a person decides.
The same agency gives caseworkers an assistant that updates a case file on request: an address, a document, a note, each change shown before it runs.AI agent, 3write; a wrong record the next process readsTools limited to the case system's write API for the open case; runs as the caseworker; each write held and shown with its arguments; one log entry per call, kept with the case.Not Annex III unless the record feeds an eligibility decision. GDPR Article 5(1)(d) accuracy. FADP Articles 7 and 8.
The same agency, as in part 1: an agent that grants, adjusts and notifies within rules.AI agent, 4commit: a benefit revoked, a letter sentMCP gateway, agent identity, per-action policy with the thresholds as rules, an approval surface above them, a reversal tool on the list, a stop, records kept six months.Annex III 5(a), no exemption in reach. Article 14 in full. Article 26(2) and (11). Article 21 FADP. Article 73 if a wrong revocation meets Article 3(49).
A bank opens accounts: a workflow extracts the identity documents with a model step, rules run the checks, and a compliance officer approves the opening.Augmented workflow, 3commit, by the officerEngine-owned path; extraction validated against a schema; the officer sees the extracted values next to the document; the model step behind the model gateway with its own gate.Anti-money-laundering rules govern the checks. Not Annex III 5(b), which is about creditworthiness. FINMA 08/2024 inventory.
The same bank, as in part 1: an agent that reconciles supplier invoices and posts matched entries, with the payment as the commit.AI agent, 3write; commit at paymentPer-action policy with the payment step as the hold; the posting tool labelled write, reversible by a counter-entry; identity; one record per call.Not Annex III. NIS2 Article 21(2): the tools and the model are in the bank's supply chain. FINMA: results "understood, explained or reproduced".
The same bank lets a client, in a chat, block a card, raise a limit for a day, or order a replacement, and an agent does it.AI agent, 3 to 4commit: a limit raised for the wrong personThe client's identity from the bank's own authentication, never from the conversation; three tools; the block labelled write, the limit and the order labelled commit with a ceiling as a rule; the client confirms each; a stop.Article 50(1): the client is told they talk to a system. Fraud rules and contract law for the limit. NIS2. GDPR Article 22 if a refusal is automated with legal effect.
A manufacturer confirms supplier orders: RPA reads the mailbox, a model step extracts quantities and dates, rules compare them with the purchase order; matches are confirmed by rule, mismatches go to a buyer.Augmented workflow, 3 and 4commit: a wrong confirmation to a supplierEngine-owned path with the mailbox as a screen the RPA drives; the model step's output validated; the matching rule has a tolerance as a parameter; every confirmation recorded with the extracted values.Not Annex III. Contract law for the confirmation. NIS2 Article 21(2) if the manufacturer is in scope.
The same manufacturer, as in part 1: an agent that orders spare parts up to a ceiling.AI agent, 4commit: money spentPer-action policy with the ceiling as a rule, approval above it, the order tool labelled commit, an audit record per order, a stop.Outside Annex III. NIS2 if in scope. Contract law does the rest, which is why the ceiling is a rule and not a prompt.
The same manufacturer lets an agent reshuffle the production plan in the execution system when a machine goes down, within constraints, and asks a planner before it changes a delivery promise.AI agent, 4write; commit at the promiseThe solver from part 2 returns feasible plans; the agent chooses among them and writes through a tool bounded to the planning window; the promise change is a commit with a hold; the previous plan is kept for reversal.Machinery and product-safety rules for the plant. Annex III point 2 only if the plant is critical infrastructure and the agent is a safety component, which a scheduler usually is not. NIS2.

Two of the nine are high-risk, both at the agency, and in each case the reason is the job, not the family. Read the platform column instead and the pattern is the one part 1 promised: the workflow rows need an engine and a validated step, the agent rows need a gateway, two identities, a policy, a hold and a record, and the difference between a level 3 agent and a level 4 agent is a single line in the policy that says which actions are held. Move that line and the job moves rows without anyone changing the model.

What this row teaches the next

One agent, one perimeter, one person who can stop it. Every control in this part assumes those three are singular: the tool list belongs to an agent, the identity names it, the approval goes to a person who can see the whole loop. The next row breaks the assumption. An orchestrator receives a goal instead of a task, splits it, and hands the pieces to other agents, each with its own tools, sometimes in another runtime, sometimes in another organisation. The perimeter has to travel with the handover, and the question that carries over is the one the series was heading for: when the agent that acts is not the agent that was given the goal, whose perimeter applies, and who is the person with the stop button?

Next in the series
  1. Part 1The map: who decides, who acts, and how far the system goes on its own. The reference for every part that follows.
  2. Part 2AI without a language model: prediction, recommendation, perception, optimisation. What already runs everywhere, and why it is not "less" than the rest.
  3. Part 3Generating and assisting: generative AI, integrated copilots, small specialised models. The person reads, the risk is the content.
  4. Part 4Answering on your own documents: RAG, agentic RAG, graph RAG. The risk becomes access to sources, and freshness.
  5. Part 6Pursuing a goal with several agents: agentic AI, delegation, intent. The most demanding regime, and the one sold first.
This article is an engineer's reading of public legal texts and public code, checked against the versions and dates given below. It is not legal advice. For a real deployment, read the texts with counsel and with your supervisor's guidance for your sector.
Read, not run. Everything in this series comes from reading public documents and public code at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: the texts are amended, the projects move monthly, and a placement that is right for one organisation's wiring is wrong for another's. Place your own jobs on the map with the people who own them. When something here does not match what you find, tell me, or better, tell the project or the authority concerned: that is the only way a map like this one stays true.

Sources
  1. Model Context Protocol specification, revision 2026-07-28, the current revision: the "Security and Trust & Safety" principles on the overview page; the "User Interaction Model" and "Security Considerations" sections of Server features: Tools; the ToolAnnotations interface in schema/2026-07-28/schema.ts, with the defaults quoted; and the key changes since 2025-11-25. Read 24 September 2026. Part 1 quoted revision 2025-06-18; the sentences quoted there are unchanged in the current one.
  2. Building effective agents, Anthropic, 19 December 2024: workflows and agents, the five workflow patterns, and the paragraph on when to use agents. Read 24 September 2026.
  3. OWASP Top 10 for LLM Applications 2025, LLM06: Excessive Agency: the three root causes and the mitigation list. Read 24 September 2026.
  4. Regulation (EU) 2024/1689 (AI Act): Article 3(3), (4) and (49); Article 12; Article 14; Article 26; Article 73, read on artificialintelligenceact.eu on 24 September 2026. Article 25(1), Article 50(1) and Annex III points 2 and 5, as read for the earlier parts, 22 and 23 September 2026.
  5. Regulation (EU) 2016/679 (GDPR), Articles 5(1)(d) and 22; Federal Act on Data Protection (FADP, SR 235.1), Articles 7, 8 and 21; Directive (EU) 2022/2555 (NIS2), Article 21(2); FINMA Guidance 08/2024. As read for the first series and for parts 1 to 4, 22 and 23 September 2026.
  6. Repositories, release tags and dates from each repository's GitHub releases page, 24 September 2026: Kuadrant/mcp-gateway, Apache-2.0, v0.9.0, 14 August 2026; Kuadrant/authorino, Apache-2.0, v0.28.0, 22 September 2026; OPA, Apache-2.0, v1.20.2, 3 September 2026; Agent Governance Toolkit, MIT, release v4.1.0, 9 June 2026, tag v5.0.0, 27 July 2026; Keycloak, Apache-2.0, 26.7.4, 16 September 2026; SPIRE, Apache-2.0, v1.15.3, 21 August 2026; Temporal, MIT, v1.32.0, 11 September 2026; Argo Workflows, Apache-2.0, v4.1.4, 18 September 2026; Robot Framework, Apache-2.0, v7.5, 14 September 2026; OpenShell, Apache-2.0, v0.0.116, 28 August 2026.
  7. Agent stack blueprint, the first series: part 1 for the MCP gateway and identity, part 4 for the policy per action, the manifest, the digest-echoing approval and the hold, part 5 for the log entry per tool call, part 6 for the rollout gate.

Independent work, not affiliated with any regulator, court, standards body, foundation or vendor named. Not legal advice. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.