Which AI for which job · part 3 of 6
Generating and assisting: the person reads, the risk is the content
Two rows of the map put a language model in front of a person who reads before anything leaves: generative AI, and the copilot built into the tool you already use. A third family, the small specialised model, sits between them and part 2. Part 2 ended on a question this row answers badly: when the output is a sentence rather than a number, what is the test? Three answers, one of which is a test set, and the articles of law that reach this row through the model's provider and through disclosure rather than through the high-risk list.
HokonokenSeptember 2026Reading time: 17 minNot legal advice · Views are my own, not my employer's
Three things to take away
- Three families, one control: the person who reads. Generative writes; the copilot proposes inside the host tool; the small specialised model turns a content task back into a label. In all three the worst action is not the model's, it is the person's acceptance, and the clause written for that is Article 14(4)(b), automation bias.
- "What is the test?" has three answers, and only one is a test set. For a label, part 2's discipline returns intact. For free text, the test is a reference set, a judge that is never the model under test, a human sample and adversarial probes, and none of it yields the accuracy figure Article 15(3) wants declared. For a proposal, the test is behavioural: how often it is accepted unchanged.
- The law reaches this row through the provider and through disclosure. Article 53, since 2 August 2025, binds whoever trained the model. Article 50, since 2 August 2026, binds whoever puts it in front of people. Article 4, since 2 February 2025, binds you for the staff who use the copilot. High-risk only if the job is on Annex III, which a copilot can be without anyone noticing.
Three families, one control
Generative. A language model produces content from a prompt: a draft, a summary, a translation, a function. The output is free text and a person reads it. By itself the system is at level 2 on the scale from part 1: it proposes. It becomes a commit the moment the text leaves without that person, published, sent, merged. The risk is what the text says, and the model's provider, not you, decided what it learned.
Integrated copilot. The same model inside the tool you already use: the CRM, the case file, the IDE, the mailbox. It reads what the user can read and proposes; the user accepts or edits. This is the most sold form of AI in 2026 and the most often confused with an agent, so the test from part 1 matters: who acts? A copilot that can only propose is level 2. The moment the same product executes on the user's behalf inside the tool, sends the mail, files the ticket, it has moved to the action rows, and the vendor rarely changes the name. Part 5 takes it from there. Two things make the copilot different from plain generation. It sees your data, so the prompt is no longer only what the user typed: the document, the email, the ticket it reads can carry instructions, and the model has no way to tell a document from a command. And it proposes inside a flow designed for speed, which is exactly where a proposal gets accepted by habit.
Small specialised model. A model fine-tuned for one task: classify the ticket, extract the six fields, route the mail, flag the clause. It is a size choice, not a regime, and the reason it gets a place here is what it does to the test: it turns a content task into a label with a target column, and everything from part 2 returns. Under the AI Act it is not a general-purpose AI model either; Article 3(63) reserves that term for a model that "displays significant generality and is capable of competently performing a wide range of distinct tasks". A model that does one thing is the row above with a language model inside, and it is often the right tool where a general model is used by reflex.
What the three share is where the control is. In part 2 the wiring set the level and the person could be absent. Here the person is the wiring: the output is text, and text does nothing until someone reads it and acts. That is the strength of this row and its trap. The strength: no tool call, no perimeter, a human at every exit. The trap: a human who reads forty proposals an hour is not reading. Article 14(4)(b), for high-risk systems, asks that the overseer "remain aware of the possible tendency of automatically relying or over-relying on the output". Nothing in the stack measures that. The last band of the figure is how you would.
What is the test?
Part 2 had a test set, metrics fixed before training and a threshold. This row has three answers, in decreasing order of comfort.
For a label, the test set. A fine-tuned classifier has a target column. Hold out a test set, fix precision and recall per class before the run, choose the threshold with the people who own the queue, register the version, watch for drift, promote through the same gate. Article 15(3), "the levels of accuracy and the relevant accuracy metrics … declared in the accompanying instructions of use", is satisfiable as written. This is why the small model gets its own family: where a job can be reduced to a label, reduce it, and the test comes back.
For free text, a reference set, a judge, a sample and a probe. There is no target column for "a good summary". What exists is a reference set of your own documents with what "good" means written down as a rubric; a judge that scores outputs against the rubric, often another model, and never the model under test; a sample re-read by people, with the sample's error rate kept as a number; and adversarial probes before promotion. The open tooling is real and dated: EleutherAI's lm-evaluation-harness, MIT, v0.4.13 of 31 August 2026, for benchmark-style evaluation; the UK AI Security Institute's Inspect, MIT, for task-shaped evaluations with scorers and a log per sample; NVIDIA's garak, Apache-2.0, v0.17.0 of 9 September 2026, for the probes; NeMo Guardrails, Apache-2.0, v0.24.1 of 16 September 2026, to run the same checks at request time. Part 6 of the first series put garak and a replayed reference set into the rollout gate as an Argo Rollouts analysis; nothing changes here. What none of it yields is an accuracy. The honest declaration for a generative system is a pass rate on a named reference set, a rubric, a judge version and a date. Say so in the instructions of use, and say what the sample of humans found.
For a proposal, behaviour. A copilot's output is judged by what the person does with it, so the test is behavioural and it runs in production. Rate accepted unchanged, by user and by task. Edit distance between what was proposed and what was sent. A sample of the accepted-unchanged proposals re-read by someone other than the person who accepted them. The counter-intuitive part: a rising unchanged rate is the signal to look at, not the success metric. It is either the model getting better or the people reading less, and only the re-read sample tells you which. That sample is the evidence Article 14(4)(b) asks for, in the only form the stack can produce it.
What the platform owes them
Take the stack of the first series and keep the model side. The MCP gateway is not needed until something acts; the model gateway is, because it is where identity, token quotas and routing live, and where a copilot serving four hundred people is told what it may cost. Inference is vLLM, Apache-2.0, v0.30.0 of 22 September 2026, or llm-d, v0.9.0 of 17 August 2026, when there are several replicas and a router to keep warm; part 6 of the first series covered the rollout. Guardrails sit on both sides of the model: topic, injection and personal data on the way in; facts, tone and leakage on the way out. And the records.
- Records, and the content question. The OpenTelemetry GenAI semantic conventions define one span per model call, with the operation, the model, the token counts and the finish reason. Their status is "Development", and the attributes that carry the messages themselves, the input and the output, are "Opt-In". Switching them on is the difference between a trace that says a call happened and a trace that contains the customer's letter. Part 5 of the first series had the four records and the six-month retention; this row adds a decision that row did not need: content in logs is personal data, its retention and access are a data-protection question under the GDPR and the FADP, and it is taken before the first span is emitted, not after the first access request.
- The evaluation harness inside the rollout gate. The reference set, the judge, the probes and the sample are not a project; they are the analysis that runs before a model version is promoted and again when a prompt template or a guardrail changes. A prompt template is a deployable. Version it, and let it trigger the gate.
- A disclosure surface. Article 50(1), in force since 2 August 2026: a system "intended to interact directly with natural persons" tells them so, "at the latest at the time of the first interaction". For a chat, that is a sentence on the screen. For a copilot drafting a letter the citizen receives, the question is whether the letter is a "direct" interaction, and that is one for counsel.
- For the copilot, the host application's permissions. The model reads what the user can read. It has no permissions of its own, and that is the right design; what it inherits is the user's, including the documents that carry instructions. The controls are the host's access model, the input guardrail, and the fact that the copilot cannot act. Part 4 is about the moment it is allowed to search on its own; part 5 about the moment it is allowed to act.
- For the small model, part 2's lifecycle, plus one question. Fine-tuning a general-purpose model does not, by itself, make you the provider of a general-purpose AI model: the Commission's guidelines of 18 July 2025 set the line at a modification whose training compute is a third or more of the original's. It can make you the provider of a high-risk system: Article 25(1)(c) says a deployer who modifies "the intended purpose of an AI system … in such a way that the AI system concerned becomes a high-risk AI system" is the provider of that system. A mail router retrained to score applications has crossed that line; the model is the same size.
Where the law reaches this row
Through the model's provider. Article 53, applicable since 2 August 2025, binds providers of general-purpose AI models to technical documentation "including its training and testing process", to information for downstream providers, to "a policy to comply with Union law on copyright", and to "a sufficiently detailed summary about the content used for training". Article 53(2) lifts the first two for models "released under a free and open-source licence" with public weights and architecture, except models with systemic risk. The Code of Practice of 10 July 2025 is the way most providers will show it, in three chapters, transparency, copyright, safety and security, the last one only for the few under Article 55. None of this binds you as a deployer. All of it is what you should ask to see before you put the model in front of people, because it is the only account you will get of what the thing learned.
Through disclosure. Article 50(1) for the interaction. Article 50(2) for the provider marking synthetic content "in a machine-readable format". Article 50(4) for you: text "published to inform the public on matters of public interest" must be disclosed as AI-generated, unless "the AI-generated content has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility". Read that clause with the map: the editorial review is not a nicety, it is the exemption, and the person holding responsibility is the control this row runs on. Article 50(5): the information comes "in a clear and distinguishable manner at the latest at the time of the first interaction or exposure".
Through your staff. Article 4, applicable since 2 February 2025: providers and deployers "shall take measures to ensure, to their best extent, a sufficient level of AI literacy of their staff and other persons dealing with the operation and use of AI systems on their behalf". For a copilot rolled out to everyone, this is the only article that reaches the people accepting the proposals, and it is already in force.
Through the high-risk list, by accident. Nothing in Annex III says "copilot". It says systems "to evaluate learning outcomes", "to analyse and filter job applications", "to allocate tasks based on individual behaviour", to draft what a public authority uses "to grant, reduce, revoke, or reclaim" a benefit. A copilot given one of those jobs is on the list at level 2, and Article 6(3) offers a way out only for a "narrow procedural task", a "preparatory task", or improving "the result of a previously completed human activity", never where the system profiles people. Whether a drafted decision letter is preparatory or is the decision is, again, for counsel; the map's job is to make sure someone asks.
Through what was already there. The GDPR and the FADP for the personal data in every prompt and every log. NIS2 Article 21(2)(e), "security in network and information systems acquisition, development and maintenance", for the code assistant whose proposals end up in production. FINMA 08/2024 §2.2 for the inventory entry, whatever the risk class.
Three organisations, nine jobs
The same three organisations. Each job is placed by who reads and what the acceptance commits.
One row is on Annex III on its face, the letter-drafting copilot, unless the preparatory-task exemption holds; one more is a change of purpose away from it, the mail router. The regulatory column here is thinner than in part 2, and it is thin in a specific way: the articles that bite are the ones already in force, 4, 50 and 53, not the ones due in December 2027. This row is regulated now, lightly, and it is the row most organisations have already deployed.
What this row teaches the next
Everything above assumes the model knows two things: what was in its training run, and what the user typed or was looking at. That is why the control can be a single person reading: the model cannot reach past the screen. The next row removes that assumption. The moment the model is allowed to search your documents rather than wait for them, the risk moves from what it writes to what it was allowed to read, and to how old that is. The question that carries over: when the model can read, who decided what it may read, and when was that last true?
Next in the series
- Part 1The map: who decides, who acts, and how far the system goes on its own. The reference for every part that follows.
- Part 2AI without a language model: prediction, recommendation, perception, optimisation. What already runs everywhere, and why it is not "less" than the rest.
- Part 4Answering on your own documents: RAG, agentic RAG, graph RAG. The risk becomes access to sources, and freshness.
- Part 5Acting within a perimeter: the AI agent with its tools, and the augmented workflow as the migration path. The risk is the action; the perimeter is outside the model.
- Part 6Pursuing a goal with several agents: agentic AI, delegation, intent. The most demanding regime, and the one sold first.
This article is an engineer's reading of public legal texts and public code, checked against the versions and dates given below. It is not legal advice. For a real deployment, read the texts with counsel and with your supervisor's guidance for your sector.
Read, not run. Everything in this series comes from reading public documents and public code at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: the texts are amended, the projects move monthly, and a placement that is right for one organisation's wiring is wrong for another's. Place your own jobs on the map with the people who own them. When something here does not match what you find, tell me, or better, tell the project or the authority concerned: that is the only way a map like this one stays true.
Sources
- Regulation (EU) 2024/1689 (AI Act), Article 3(23), (63) and (66), Article 4, Article 6(3), Article 14(4), Article 15(3), Article 25(1), Article 50, Article 53, Article 55, Annex III points 3, 4 and 5; texts read on artificialintelligenceact.eu on 23 September 2026. Application dates as amended by Regulation (EU) 2026/1744, read in part 3 of the first series.
- European Commission, guidelines on the scope of obligations for providers of general-purpose AI models, 18 July 2025, as summarised on artificialintelligenceact.eu (the one-third-of-training-compute criterion); General-Purpose AI Code of Practice, 10 July 2025, three chapters. Read 23 September 2026.
- Directive (EU) 2022/2555 (NIS2), Article 21(2)(e); Federal Act on Data Protection (FADP, SR 235.1), Articles 7 and 8; FINMA Guidance 08/2024, §2.2. As read for part 3 of the first series, 22 September 2026.
- OpenTelemetry semantic conventions for GenAI, documents gen-ai-spans.md, gen-ai-agent-spans.md and gen-ai-events.md, status "Development", attributes gen_ai.input.messages and gen_ai.output.messages at requirement level "Opt-In"; read from the repository on 23 September 2026, alongside semantic-conventions v1.44.0 of 4 August 2026.
- lm-evaluation-harness, MIT, v0.4.13, 31 August 2026. Inspect, MIT, UK AI Security Institute, repository read 23 September 2026. garak, Apache-2.0, v0.17.0, 9 September 2026. NeMo Guardrails, Apache-2.0, v0.24.1, 16 September 2026. vLLM, Apache-2.0, v0.30.0, 22 September 2026. llm-d, Apache-2.0, v0.9.0, 17 August 2026. Release tags and dates from each repository's GitHub releases page, 23 September 2026.
- Agent stack blueprint, the first series, parts 1, 5 and 6, for the model gateway and guardrails, the records, and the rollout gate.
Independent work, not affiliated with any regulator, court, standards body, foundation or vendor named. Not legal advice. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.