Which AI for which job · part 4 of 6
Answering on your own documents: what the model may read, and how old it is
Part 3 ended on a question: when the model can read, who decided what it may read, and when was that last true? This row is where a language model answers from your documents rather than from its training run. Three ways to do it, three different deciders, and one thing they share that the previous rows did not have: an index, which is a copy of your data with an access list and a date. The risk moves from what the model writes to what it was allowed to read.
HokonokenSeptember 2026Reading time: 16 minNot legal advice · Views are my own, not my employer's
Three things to take away
- Retrieval puts your documents where the training run was. The model's memory becomes two things: the weights, which are the provider's, and the index, which is yours. Everything part 3 said about content still holds. What is new is that the answer draws on a copy of your data, and the copy has an access list and a date.
- Three RAGs, three deciders. Classic: the retriever decides, by similarity, in one pass. Agentic: the model decides what to search, where and how many times, so retrieval has become an action and part 5 begins here. Graph: the ingest decided, before any question was asked, which entities and relations exist; the risk moved to ingest time.
- The index is a copy, and the law treats copies as data. Accuracy and storage limitation under GDPR Article 5, erasure under Article 17: a deletion that does not reach the index is not a deletion. And access has to be checked at query time, as the user, or the model becomes the way around the permission system.
What changed since part 3
The 2020 paper that named the technique described a model that combines "pre-trained parametric and non-parametric memory": the weights on one side, "a dense vector index" reached through a retriever on the other. That second memory is the whole point of this row. It is yours. You fill it, you can empty it, you can say what is in it and when it was put there. Nothing about the weights gives you any of that.
Four consequences, and the rest of this part is about them. The model can be right about your data without anyone retraining it. It can be confidently wrong about last quarter's policy, in the present tense, with a citation, because the index is as old as its last ingest. It can answer a user's question from a document that user was never allowed to open, because the retriever looked at similarity and not at permissions. And, for the first time in this series, the answer can point at its source: the chunk, the page, the version. The citation is not a courtesy in this row; it is the control.
Three RAGs, three deciders
Classic RAG. The question is embedded, the nearest chunks are fetched, the model writes an answer from them, a person reads it. The retriever decides what the model sees, by similarity, once. The person is still the control, as in part 3, and the system sits at level 2. What the person cannot see is what was not retrieved: a missing chunk produces a fluent answer with a gap in it, and the citation shows only what was used.
Agentic RAG. The 2025 survey that named it describes "embedding autonomous AI agents into the RAG pipeline" so that they "dynamically manage retrieval strategies, iteratively refine contextual understanding, and adapt workflows", using "reflection, planning, tool use, and multi-agent collaboration". Read that with the map: the model now decides what to search, in which store, how many times, and when to stop. Retrieval has become an action chosen by the model, and every question from the action rows applies to it: through which gateway, under whose identity, with what record. The system is at level 4 even though nothing is written anywhere, because its reads happen without a person and are read about after, and OWASP's 2025 list has an entry for the failure, "Excessive Agency". Part 5 starts exactly here; this part only asks what the retrieval action needs.
Graph RAG. The 2024 paper from Microsoft Research starts from a limitation: conventional retrieval "fails on global questions directed at an entire text corpus, such as 'What are the main themes in the dataset?'". Its answer is to build "an entity knowledge graph from the source documents" at ingest and to "pregenerate community summaries for all groups of closely related entities", so that a global question is answered from the summaries rather than from nearest chunks. What that does to the map is subtle. The decision about what the model will draw on was taken before any question, by a model, at ingest time. An extracted relation is an assertion about your data, made by a model, on a date. When it is wrong, it is wrong for every question that touches it, and the person reading the answer has no chunk to check against, only a summary of a summary. The control moved to the ingest, and so must the review.
The index is a copy
Four properties follow from that sentence, and each one is a requirement on the ingest pipeline, not on the model.
Access. The model has no permissions of its own; part 3 established that for the copilot and it holds here. What it inherits is the user's, and the retriever must honour them at query time: the same user, the same identity from the platform's identity provider, evaluated against the access list that travelled with each chunk. A retriever that searches the whole index and filters afterwards has already leaked; a retriever that searches an index built without access lists cannot filter at all. OWASP's 2025 list names both halves: "Sensitive Information Disclosure" and "Vector and Embedding Weaknesses", the latter written for systems "relying on semantic search and retrieval". The first series had identity at the ingress and at the tool gateway. This row needs it at the index.
Freshness. Every chunk carries the version of its source and the date it was ingested, and the answer states them. That is the whole freshness axis from part 1, made concrete. A policy replaced on Monday is answered from the Friday version until the next ingest, and nothing in the model knows. GDPR Article 5(1)(d) asks that personal data be "accurate and, where necessary, kept up to date"; the index is where that clause is met or missed for this row.
Erasure. A deletion in the source system that does not propagate to the index is not a deletion. Article 17 gives the data subject "the erasure of personal data concerning him or her without undue delay", and Article 5(1)(e) limits storage to "no longer than is necessary for the purposes". Whether an embedding of a paragraph about a person is personal data is a question I will leave to counsel; the paragraph next to it in the chunk store certainly is. The ingest pipeline needs a delete path that is as reliable as its insert path, and a record that it ran.
Integrity. The index is fed by a crawler, and a crawler reads what is placed where it reads. A document that contains instructions is, once retrieved, an instruction to the model. OWASP calls it "Prompt Injection" when it arrives through the prompt and "Data and Model Poisoning" when it arrives through the data; the AI Act's Article 15(5) names "data poisoning" among the attacks a high-risk system must resist. The defences are dull: a list of what the crawler may read, a review of what enters a graph, an input guardrail between the retrieved chunk and the model, and a record of which chunk was in the context when the answer went wrong.
What the platform owes them
Take part 3's platform and add the index and everything that feeds it.
- An ingest pipeline with lineage. The parser from part 2, Docling for layout and OCR, is the data stage of this row. Each chunk or node records its source, version, ingest date and access list; each run records what it inserted and what it deleted. This is the same Article 10(2) discipline part 2 had, applied to a store that changes every night.
- An index you can govern. Milvus, an LF AI & Data project under Apache-2.0, v3.0.2 of 20 September 2026. Qdrant, Apache-2.0, v1.19.1 of 4 September 2026. pgvector, under the PostgreSQL licence, v0.8.6 of 22 September 2026, which puts the vectors in the database you already back up, permission and audit. Weaviate, v1.39.6 of 22 September 2026. The first series placed the vector store behind the OGX server's Vector_IO in part 1; whichever you choose, the two questions are the same: does it evaluate a per-document access filter inside the query, and does it log what it returned. For graphs, Microsoft's GraphRAG, MIT, v3.1.2 of 21 August 2026, is the reference implementation of the paper; the graph it builds is an artefact to version and review like a model.
- Identity at the retriever. The question arrives with the user's identity from Keycloak or the workload identity from SPIRE, as in the first series, and the retrieval runs under it. For agentic RAG this is not optional: the model's search is a tool call, it goes through the MCP gateway of part 1, the tool is on an allowlist, and the query is recorded with the caller.
- Records that include what was read. The GenAI semantic conventions from part 3 have an "embeddings" operation, so the embedding call is traced. The retrieval itself is a database call; nothing in the conventions records which chunks went into the context. That record is yours to make, and it is the one an auditor will ask for: for this answer, which chunks, from which versions, retrieved under whose identity. Keep it with the answer, under the six-month floor of Article 26(6) if the system is high-risk, and under the data-protection decision from part 3 either way, because the chunks are content.
- Two more tests in the rollout gate. Part 3's reference set and judge apply. This row adds two measures the judge can score: whether the right chunk was in the context at all, and whether the answer is supported by the chunks it cites rather than by the weights. And a new trigger: the gate reruns on every index rebuild, not only on every model change, because the index is part of the system under test.
Where the law reaches this row
Through data protection, first. This is the first row where the law's main handle is not the AI Act. GDPR Article 5(1) gives the index its five constraints in one paragraph: purpose limitation, data "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes"; minimisation, "limited to what is necessary"; accuracy; storage limitation; and integrity, "protection against unauthorised or unlawful processing". Article 5(2) adds that the controller must "be able to demonstrate compliance", which is what the ingest records are for. Article 17 is the erasure path. The Swiss FADP carries the same principles; the first series read its Articles 7, 8 and 12, protection by design, data security and the record of processing activities, and an index of your documents is a processing activity that belongs in that record.
Through the AI Act, lightly. Article 50(1) for any assistant a person talks to. Articles 10 and 15(5) if the job is on Annex III, in which case the index is training-adjacent data and poisoning is a named attack. Article 26(6) for the logs. Nothing in Annex III says "retrieval"; the row inherits the risk class of the job it serves.
Through security law. NIS2 Article 21(2) asks for "access control policies and asset management" and "the use of cryptography". An index of everything an organisation knows is a single asset that holds all of it, at rest, in a form built to be searched; it belongs on the asset list with the encryption and the access policy that implies. The same article's supply-chain clause covers the embedding model and the store.
Through the licence. An index is a reproduction of the documents in it. For your own documents that is your decision. For licensed content, manuals, standards, market data, the licence says whether a copy built for machine retrieval is allowed, and most were written before anyone asked. Read it before the crawler does.
Three organisations, nine jobs
The same three organisations. Each job is placed by who decides what is read, and the last column is what the copy brings with it. Article 50(1) applies to every row where a person talks to the assistant and is not repeated below.
None of the nine is high-risk under the AI Act on its own, and every one of them carries a data-protection or confidentiality obligation that already applied to the documents before anyone indexed them. That is the pattern of this row: the law did not change, the copy did. The three agentic rows are the ones where the platform column grows a gateway and an identity, and those are the columns of part 5.
What this row teaches the next
Agentic RAG is an agent that is only allowed to read. It chooses what to fetch, it goes through a tool gateway, it runs as the user, and every query is recorded; if it were not, the file the caseworker may not open would be one similarity search away. All of that machinery exists so that a read stays a read. The next row lets the same machinery write. The question that carries over is the one the reversibility axis was made for: when the worst action is no longer a disclosure but a change in a system, what has to be true before the model is allowed to make it?
Next in the series
- Part 1The map: who decides, who acts, and how far the system goes on its own. The reference for every part that follows.
- Part 2AI without a language model: prediction, recommendation, perception, optimisation. What already runs everywhere, and why it is not "less" than the rest.
- Part 3Generating and assisting: generative AI, integrated copilots, small specialised models. The person reads, the risk is the content.
- Part 5Acting within a perimeter: the AI agent with its tools, and the augmented workflow as the migration path. The risk is the action; the perimeter is outside the model.
- Part 6Pursuing a goal with several agents: agentic AI, delegation, intent. The most demanding regime, and the one sold first.
This article is an engineer's reading of public legal texts and public code, checked against the versions and dates given below. It is not legal advice. For a real deployment, read the texts with counsel and with your supervisor's guidance for your sector.
Read, not run. Everything in this series comes from reading public documents and public code at a stated date, not from running them in production. Treat it as a map to test, not a result to trust: the texts are amended, the projects move monthly, and a placement that is right for one organisation's wiring is wrong for another's. Place your own jobs on the map with the people who own them. When something here does not match what you find, tell me, or better, tell the project or the authority concerned: that is the only way a map like this one stays true.
Sources
- Patrick Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv 2005.11401, submitted 22 May 2020, v4 of 12 April 2021. Darren Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization, arXiv 2404.16130, submitted 24 April 2024, latest version 19 February 2025. Aditi Singh et al., Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, arXiv 2501.09136, submitted 15 January 2025. Abstracts read 23 September 2026.
- OWASP Top 10 for LLM Applications 2025, entries LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM04 Data and Model Poisoning, LLM06 Excessive Agency, LLM08 Vector and Embedding Weaknesses; read 23 September 2026.
- Regulation (EU) 2016/679 (GDPR), Article 5 and Article 17, read on gdpr-info.eu, 23 September 2026. Federal Act on Data Protection (FADP, SR 235.1), Articles 7, 8 and 12, as read for part 3 of the first series on Fedlex, 22 September 2026; the principles article was not re-read for this part and is not quoted.
- Regulation (EU) 2024/1689 (AI Act), Article 4, Article 10, Article 15(5), Article 26(6), Article 50; Directive (EU) 2022/2555 (NIS2), Article 21(2); FINMA Guidance 08/2024, §2.2 and §2.5. As read for the earlier parts, 22 and 23 September 2026.
- Milvus, LF AI & Data, Apache-2.0, v3.0.2, 20 September 2026. Qdrant, Apache-2.0, v1.19.1, 4 September 2026. pgvector, PostgreSQL licence, v0.8.6, 22 September 2026. Weaviate, v1.39.6, 22 September 2026. GraphRAG, MIT, v3.1.2, 21 August 2026. Docling, MIT, v2.130.0, 22 September 2026. Release tags and dates from each repository's GitHub releases or tags page, 23 September 2026.
- OpenTelemetry semantic conventions for GenAI, the "embeddings" operation, read 23 September 2026. Model Context Protocol specification, revision 2025-06-18, server primitives: resources, prompts, tools; the current revision, 2026-07-28, keeps the three (see part 5).
- Agent stack blueprint, the first series, parts 1, 3 and 5, for the vector store behind the OGX server, identity, the regulatory matrix and the records.
Independent work, not affiliated with any regulator, court, standards body, foundation or vendor named. Not legal advice. Product names belong to their owners. Views are my own and do not represent the position of my employer. Text and diagrams: CC BY 4.0; quoted code and documents stay under their own licences.