Which AI for which job · part 6 of 6
Pursuing a goal with several agents: agentic AI, delegation, intent, and the person who can still stop it
Part 5 ended on a question: when the agent that acts is not the agent that was given the goal, whose perimeter applies, and who is the person with the stop button? This last row is where a system receives a goal rather than a task, splits it, and hands the pieces to other agents, each with its own tools, sometimes in another runtime, sometimes in another organisation. It is the row every deck opens with and the one the fewest organisations run. The map's answer to the question is the sentence the whole series was written for: "agentic" is not a higher level of the same thing, it is a different regime of responsibility, in which the unit of governance is no longer the tool call but the handover.
HokonokenSeptember 2026Reading time: 20 minNot legal advice · Views are my own, not my employer's
Three things to take away
- Agentic AI is delegation, and delegation can only narrow. An orchestrator that hands part of a goal to another agent hands over a perimeter with it, and the child's perimeter must be a subset of the parent's: fewer tools, a lower ceiling, a smaller budget, a bounded depth. The one open toolkit the first series read enforces exactly that: a sub-agent's intent must be a subset of its parent's planned actions, and the code raises an error otherwise.
- Intent is declared, not configured, and the audit becomes a tree. Before it acts, an agent states what it plans to do, and the runtime checks each action against the declaration. The record of a goal is then a tree: the goal at the root, the handovers as nodes, part 5's tool calls as leaves. Anthropic's own multi-agent system needed "full production tracing" to be debugged at all, and "multi-agent systems use about 15× more tokens than chats". The budget is a rule, inherited and divided.
- The law has no row for the orchestrator. It has a person. Article 14 asks for a natural person who can stop the system; Article 26(2) asks that they have "competence, training and authority". Across organisations, the protocol built for agent-to-agent work makes the other agent "opaque" by design, so inspection is replaced by identity and contract. The stop has to reach every branch, or it is not a stop.
What changed since part 5
A goal instead of a task. Anthropic's 2024 note lists "orchestrator-workers" among its workflow patterns: "a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results". In that note the workers write text and the path is decided in code, which makes it part 3 several times over. Give the workers tools and let the orchestrator decide at run time what to delegate and to whom, and the map has no row left but this one: the model decides the split, the path and the actions, and a person sees the outcome. Level 5, delegates.
The same company described what it cost them to run one. Their research system, in June 2025, has "a lead agent" that "coordinates the process while delegating to specialized subagents that operate in parallel". Three sentences from that note belong on the map. "Agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats." "Without effective mitigations, minor system failures can be catastrophic for agents." And on where the pattern does not pay: "Most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time." Read with the map, the row is worth its cost where a goal splits into independent reads, which is research, and costly where the pieces write to the same systems in a fixed order, which is a close, an onboarding, an incident. Those are the jobs the row is sold for.
So the map draws the line between part 5 and part 6 not at the number of agents but at who drew the child's perimeter. If a person drew each agent's perimeter and a process engine calls the agents in order, that is part 5 several times, an augmented workflow whose steps happen to be agents, and it is the right way to build most of the nine jobs below. It becomes this row when an agent decides, at run time, to give a piece of the goal to another agent. Delegation is the act. Intent is what makes it governable, and it is the one control this row adds to everything the first series built.
Three shapes of delegation
Inside one runtime. The orchestrator and its workers share a platform, a gateway and an identity provider. Part 4 of the first series read the one open toolkit that governs this shape and the relevant lines are short. An agent declares an intent before it acts, a list of planned actions with bounds on their arguments; the runtime checks each action against it. A manifest sets a maximum delegation depth and requires scope narrowing; an intent created for a sub-agent must be a subset of its parent's planned actions, and the code raises an error otherwise. Each delegation is logged. Every worker has its own workload identity; a worker that reads has no write tool on its list; a worker that commits carries part 5's hold with it. Nothing in this shape is new except the check on the handover, and that check is the whole row.
Across runtimes. The Agent2Agent protocol, Apache-2.0 and under the Linux Foundation, at v1.0.1 since 28 May 2026, describes itself as "an open protocol enabling communication and interoperability between opaque agentic applications". Two objects carry the weight. The Agent Card is "a JSON metadata document published by an A2A Server, describing its identity, capabilities, skills, service endpoint, and authentication requirements". The Task is "the fundamental unit of work managed by A2A, identified by a unique ID", and "tasks are stateful and progress through a defined lifecycle". Agents collaborate, the specification says, "without needing to share their internal thoughts, plans, or tool implementations". That opacity is the design goal and the governance problem in one sentence. You cannot narrow a perimeter you cannot see. What you can do is what the protocol gives you: authenticate the agent by its card, give it a task and not the goal, bound the task, keep its lifecycle in your tree, and test that cancelling it works.
Across organisations. The supplier's agent, the other agency's agent, the counterparty's agent. The remote agent is a supplier in the sense of NIS2 Article 21(2), the message is the only record you hold, and the contract is the perimeter. OWASP's agentic threat list, at version 1.1 of December 2025, has three entries written for this shape: "Insecure Inter-Agent Protocol Abuse", where "attackers could manipulate coordination messages, memory, or protocol logic to bypass safeguards and alter agent goals or actions"; "Identity Spoofing & Impersonation"; and "Rogue Agents in Multi-Agent Systems". None of the three can be answered inside your runtime. They are answered by the card, the contract and the tree.
Delegation, bounded: four rules
Narrower only. A child gets a subset of the parent's tools, a ceiling no higher than the parent's, a share of the parent's budget, and a depth one less. The toolkit's manifest states the first two as configuration and the intent check enforces them at run time; the last two are yours to add. A child that could do more than its parent is the failure OWASP calls "Privilege Compromise", and it is the failure every "autonomous" demo quietly relies on.
Its own identity. Every agent in the tree has a workload identity, SPIRE in the first series, and no two share a credential. The remote agent is authenticated by its card and its card by whatever your organisation trusts. The record must be able to say which agent, on behalf of which parent, for which person, or the tree below is a list of events with no edges.
A budget, inherited and divided. Tokens, actions, time, money. The person at the root sets it; the orchestrator divides it; a child that exhausts its share stops and reports, it does not borrow. Fifteen times the tokens of a chat is Anthropic's number for their own system; yours will differ, and the budget is how you find out before the invoice does.
The stop propagates. Article 14(4)(e) asks for "a 'stop' button or a similar procedure". For one agent that is the end of a loop. For a tree it is a cancel that reaches every task in flight, in every runtime, including the remote one through its task lifecycle, and it is tested, not assumed. The person who presses it is one person, and the tree tells them what was in flight. OWASP's "Overwhelming Human in the Loop" is the opposite failure: a tree that sends every commit to the root drowns the person the law relies on, and the design decision of which commits bubble up, and which are held at the branch by someone with the authority for that branch, is the oversight design for this row.
The audit becomes a tree
Part 5's record was a line per tool call. This row's record is a tree, and the first series described the data structure in its fifth part: each entry chained to the previous one, a root hash that lets anyone verify that an entry belongs and has not changed, and, because the toolkit's own tree "has no external witness", a witness outside the process that holds the roots. The root of the tree is the goal and the person who gave it. The nodes are the handovers: who delegated, to whom, with which scope, when, and what came back. The leaves are the tool calls, each with the caller, the tool, the arguments, the decision and the result. The tree is also the rollback plan: what each branch did, in order, is what has to be undone when a goal is abandoned halfway, and the reversal tools of part 5 are on the lists for that reason.
Two things depend on that tree that nothing else can provide. The first is debugging. Anthropic's note is candid: "agents make dynamic decisions and are non-deterministic between runs, even with identical prompts. This makes debugging harder", and "adding full production tracing let us diagnose why agents failed and fix issues systematically". The second is the gate. You cannot test the whole by testing the parts, because the parts are chosen at run time. The gate for this row runs a set of goals and scores the tree: was the goal met, which handovers happened, did any child exceed its scope, did the stop work, what did it cost. It reruns on any change to any agent's model, prompt or tool list, and on any change to the orchestrator's prompt, which is this system's control law. OWASP's entry for the missing tree is "Repudiation & Untraceability"; Article 12(1) is the law's name for its presence, and Article 73's report is written from it.
What the platform owes them
Everything from part 5, once per agent, and five things that exist only because there is more than one.
- A runtime that knows what a handover is. Intent declared per agent, checked per action; delegation depth and scope narrowing as manifest rules; every delegation logged. The Agent Governance Toolkit, MIT, is the open reference the first series read, at release v4.1.0 of 9 June 2026 and tag v5.0.0 of 27 July 2026.
- An identity per agent. SPIRE, v1.15.3 of 21 August 2026, for workloads; the person's identity from Keycloak travelling down the tree with the task, so that a leaf can still say for whom.
- A boundary protocol, and the discipline to use it as a boundary. A2A for what crosses a runtime or an organisation, with the card verified, the task bounded, the lifecycle kept in your tree. The Model Context Protocol's current revision moved long-running tool calls into a "tasks" extension with "polling, mid-flight input, and durable handles"; a handle you can poll is a handle you can cancel, and cancellation is the property to test.
- A witness for the tree. The first series named the open building blocks for an external witness; whichever you choose, the roots leave the process that produced them.
- A budget and a stop at the root. Set by the person, divided by the orchestrator, enforced by the runtime; a cancel that fans out and is exercised in the gate.
- A gate that scores trees, and a rollout per agent. The blue/green gate from the last part of the first series applies to each agent's model, prompt and tool list as it applied to a model, with the whole tree as the thing under test.
Where the law reaches this row
One system, or several? Article 3(1) defines an AI system; it does not say how many models one contains. When several agents from several providers pursue one goal, the organisation that composes them chose the goal, the tools, the boundaries and the person, and Article 25(1) makes a deployer the provider when it substantially modifies a high-risk system or changes a system's purpose so that it becomes one. On that reading the composed tree is the system and its composer is its provider. As in part 5, counsel should confirm it before anyone relies on it.
Article 14, for a tree. Paragraph 4(a) asks that the person can "understand the relevant capacities and limitations" and "detect anomalies". For one agent that is a log; for a tree it is the tree, shown to a person who can read it, not the last message of the orchestrator. Paragraph 4(e)'s stop must reach every branch. Article 26(2)'s "authority" must extend to the remote agent's task, which, across organisations, means the contract says so.
Article 73, from the tree. A serious incident is reported within fifteen days, two for the widespread cases. Article 3(49) includes "a serious and irreversible disruption of the management or operation of critical infrastructure", which is written for the incident-response and plant jobs below, and "the infringement of obligations under Union law intended to protect fundamental rights", which is written for the onboarding one. The report names what happened, in which branch, under whose scope; that is the tree.
Data protection, purpose by purpose. A child that receives a sub-goal receives data for it, and GDPR Article 5(1)(b) asks that data be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes". A tree that hands a citizen's file to three agents for three purposes has three purposes to justify, and across agencies a legal basis for each sharing. Article 22 and Article 21 FADP apply at the leaves that decide about a person without one.
Security law, twice. NIS2 Article 21(2) names access control, supply-chain security and incident handling. In this row the remote agent is the supply chain, the tree is the access-control record, and an incident-response agent is itself part of incident handling and must not be the only one. The Swiss financial regulator asks that results can be "understood, explained or reproduced"; for an outcome produced by five agents, the tree is the only place that can be true. Annex III still says nothing about agents, and Article 50(1) still applies wherever a person talks to one.
Three organisations, nine jobs
The same three organisations. None of the nine is a new job: each is parts 2 to 5 composed. What is new in every row is the handover, and it sets the level: an orchestrator that delegates at run time is at 5 whatever its children do. Where a commit is held for a person, the hold is in the platform column; it does not lower the level, because the level is set by who decides the split.
One of the nine is high-risk, and again it is the job, not the family. Most of the nine are better built as part 5 several times, with a process engine calling agents in an order a person wrote, and they should start there; the engine's records will say when a handover can be trusted to a model. What the platform column adds, in every row, is the same short list: a handover that narrows, an identity per agent, a budget that divides, a stop that fans out, and a tree.
The map, read from the bottom
Six parts, five regimes. In the first, the risk is the quality of a number, and the person or the rule that acts on it. In the second, a person reads, and the risk is the content. In the third, the risk is what the model was allowed to read, and how old it is. In the fourth, the risk is the action, and the perimeter around it. In this one, the risk is the handover, and the tree that records it. At each step the controls changed kind, not size, and at no step did the size of the model decide anything. A logistic regression wired to refuse applications sits higher on the map than a frontier model drafting emails, and the law reads the map, not the model.
The first series built a platform layer by layer and never said when you need one. This series has now assigned each layer to the row that needs it: the model gateway and the guardrails to the content row, the index with its access list to the sources row, the tool gateway with its policy and its hold to the action row, intent, identity per agent and the audit tree to this one. A job that sits in the first two regimes needs none of the last three layers, and most jobs do.
For each job on your list, ask the two questions and then a third. Who decides. Who acts. And who can stop it. If the third has no name, the job is not ready for the row it was sold at.
The series
- Part 1The map: who decides, who acts, and how far the system goes on its own. The reference for every part that follows.
- Part 2AI without a language model: prediction, recommendation, perception, optimisation. What already runs everywhere, and why it is not "less" than the rest.
- Part 3Generating and assisting: generative AI, integrated copilots, small specialised models. The person reads, the risk is the content.
- Part 4Answering on your own documents: RAG, agentic RAG, graph RAG. The risk becomes access to sources, and freshness.
- Part 5Acting within a perimeter: the AI agent with its tools, and the augmented workflow as the migration path. The risk is the action; the perimeter is outside the model.
- Part 6Pursuing a goal with several agents: agentic AI, delegation, intent. This part.
This article is an engineer's reading of public legal texts and public code, checked against the versions and dates given below. It is not legal advice. For a real deployment, read the texts with counsel and with your supervisor's guidance for your sector.
Read, not run. Everything in this series comes from reading public documents and public code at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: the texts are amended, the projects move monthly, and a placement that is right for one organisation's wiring is wrong for another's. Place your own jobs on the map with the people who own them. When something here does not match what you find, tell me, or better, tell the project or the authority concerned: that is the only way a map like this one stays true.
Sources
- Building effective agents, Anthropic, 19 December 2024, the "orchestrator-workers" pattern. How we built our multi-agent research system, Anthropic, 13 June 2025: the lead agent and subagents, the token figures, the sentence on minor failures, the sentence on coding tasks, and the passage on tracing. Read 24 September 2026.
- Agent2Agent (A2A) protocol specification, version 1.0.0 as published, definitions of Agent Card and Task and the sentence on internal state; a2aproject/A2A, Apache-2.0, Linux Foundation, release v1.0.1 of 28 May 2026. Read 24 September 2026.
- Agentic AI: Threats and Mitigations, OWASP Top 10 for LLM Applications & Generative AI, Agentic Security Initiative. The resource page lists v1.0 of 17 February 2025; the PDF downloaded from it on 24 September 2026 is Version 1.1, December 2025, CC BY-SA 4.0, with seventeen threats, T1 to T17. Quoted: T3 Privilege Compromise, T5 Cascading Hallucination Attacks, T6 Intent Breaking & Goal Manipulation, T8 Repudiation & Untraceability, T9 Identity Spoofing & Impersonation, T10 Overwhelming Human in the Loop, T13 Rogue Agents in Multi-Agent Systems, T16 Insecure Inter-Agent Protocol Abuse.
- Agent Governance Toolkit, Microsoft, MIT: the intent, delegation and audit sections as read for part 4 of the first series (maximum delegation depth, scope narrowing, the sub-agent intent check, the audit tree without external witness); release v4.1.0 of 9 June 2026, tag v5.0.0 of 27 July 2026, checked 24 September 2026. SPIRE, Apache-2.0, v1.15.3, 21 August 2026.
- Model Context Protocol specification, revision 2026-07-28: the "Tasks" extension on the overview page and item 6 of the key changes. Read 24 September 2026.
- Regulation (EU) 2024/1689 (AI Act): Article 3(1) and (49), Article 12, Article 14, Article 26, Article 73, read 24 September 2026; Article 25(1), Article 50(1), Annex III points 2 and 5, as read for the earlier parts. Regulation (EU) 2016/679 (GDPR), Articles 5(1)(b) and 22; FADP Articles 7, 8 and 21; Directive (EU) 2022/2555 (NIS2), Article 21(2); FINMA Guidance 08/2024. As read for the first series and parts 1 to 5, 22 to 24 September 2026.
- Agent stack blueprint, the first series: part 4 for intent, delegation and the approval hold, part 5 for the audit tree and the witness, part 6 for the rollout gate.
Independent work, not affiliated with any regulator, court, standards body, foundation or vendor named. Not legal advice. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.