Connecting AI to policies and evidence (not to its memory)
How to connect AI to the company's internal regulatory base and audit trails so it answers by citing the real policy, with the General Data Protection Law (LGPD) as the permanent case running through almost every compliance question.
You ask AI whether a new customer data flow complies with the General Data Protection Law (LGPD) and ask it to cite the article backing the analysis. It answers "compliant with article 7, no relevant concerns," confident, with the look of a closed opinion. You almost attach it straight to the controls report. Article 7 covers a different legal basis for processing, and nobody opened the text of the law to check before reporting.
You ask public AI which internal controls cover the new refund process, pasting the entire finance policy into the prompt, with approval thresholds and approver names, to give it more context. It hands back a list of controls, with regulation number and everything, but it made it up, because it had never seen your policy before that paste. The control you were going to report to internal audit doesn't exist, and the company's approval threshold just went outside the environment controlled by finance. Connecting AI to the real policy, inside a controlled environment, solves both ends: it cites the real control, and the threshold data never leaves the house.
You ask AI what the data protection clause of the firm's template contract with suppliers says, without connecting the real template. It answers with a standard market clause, tidy, but not the one your firm actually uses, with a retention period different from what's in your real document. You almost send the draft for the client's review like that. Connecting AI to the firm's real, indexed template makes it answer citing the actual clause, deadline and all, instead of inventing a market standard.
You ask public AI whether a campaign claim needs regulatory backing, pasting the brand's history of complaints with the advertising self-regulation body into the prompt, so it understands the context. It answers with a list of regulatory requirements that seem correct, but it made it up, because it had never seen your marketing compliance guide before that paste. The requirement you were going to follow doesn't exist the way it described, and the brand's complaint history just went outside the controlled environment. Connecting AI to the real advertising compliance guide solves both ends: it cites the real requirement, and the sensitive history never leaves the house.
You ask public AI which labor clauses cover sharing an employee's health data, pasting a specific employee's medical history into the prompt to give it more context. It hands back a list of legal requirements, with article and everything, but it made it up, because it had never seen your HR data policy before that paste. The requirement you were going to communicate to the committee doesn't exist, and the employee's health data just went outside the HR system. Connecting AI to the real HR data protection policy solves both ends: it answers citing the real policy, and the sensitive data never leaves the house.
You ask public AI whether the new geolocation feature needs a specific privacy notice, pasting the entire PRD with real users' behavioral data into the prompt so it understands the request. It answers with a notice requirement that seems right, but it made it up, because it had never seen your real privacy policy before that paste. The requirement that was going into the PRD doesn't exist the way it described, and users' behavioral data just went outside the controlled environment. Connecting AI to the real privacy policy solves both ends: it cites the real requirement, and the data never leaves the house.
You ask public AI whether you can offer an extra discount in a negotiation with a public agency, pasting the account history with amounts and the public official's name into the prompt so it understands the context. It answers with a discount limit that seems right, but it made it up, because it had never seen your real anti-corruption policy before that paste. The limit you were going to apply in the negotiation doesn't exist the way it described, and the public official's name just went outside the environment controlled by sales. Connecting AI to the real anti-corruption policy solves both ends: it cites the real limit, and the sensitive data never leaves the house.
You ask public AI what the correct procedure is for disposing of a document with personal data, pasting the operation's entire data inventory, with volume and system for each item, into the prompt so it understands the context. It hands back a procedure with deadline and disposal method, but it made it up, because it had never seen your real retention policy before that paste. The procedure you were going to follow doesn't exist the way it described, and the operation's data inventory just went outside the controlled environment. Connecting AI to the real retention policy solves both ends: it cites the real procedure, and the inventory never leaves the house.
You ask public AI to map the controls covering the new billing process, pasting the company's entire risk map into the prompt, with system names and sensitive data volumes, so it understands the context. It puts together an impeccable list, with articles and owners, but it made it up, because it had never seen your policy or your risk map before that paste. The control you were going to report to the audit doesn't exist, and the company's risk map just went outside the controlled environment. Connecting AI to the real policies, inside an environment the company controls, solves both ends: it answers with traceability, and the risk map never leaves the house.
You ask public AI whether the new data pipeline meets the LGPD's data minimization principle, pasting the entire architecture design into the prompt, with database names and sensitive data volumes, so it understands better. It answers with a compliance assessment that seems right, but it made it up, because it had never seen your real impact assessment before that paste. The assessment that was going to the architecture committee doesn't exist the way it described, and the sensitive pipeline design just went outside the controlled environment. Connecting AI to the real impact assessment solves both ends: it cites the real assessment, and the design never leaves the house.
You ask public AI whether the consent screen for the new flow complies with the privacy policy, pasting entire session recordings with participants' personal data visible into the prompt to give it more context. It answers with a compliance assessment that seems right, but it made it up, because it had never seen your real policy before that paste. The assessment that was going to justify the launch doesn't exist the way it described, and the session participants' personal data just went outside the controlled environment. Connecting AI to the real privacy policy solves both ends: it cites the real assessment, and the participants' data never leaves the house.
You ask public AI whether entering a new international market requires any specific regulatory registration, pasting the entire expansion plan into the prompt, with financial projections and the target country's name, so it understands the context. It hands back a list of requirements, with regulator name and everything, but it made it up, because it had never seen your real plan before that paste. The requirement that was going into the board memo doesn't exist, and the confidential expansion plan just went outside the controlled environment. Connecting AI to the company's real regulatory mapping solves both ends: it cites the real requirement, and the plan never leaves the house.
Whoa, notice something: the risk in compliance has two faces, and they pull in opposite directions. On one side, AI left on generic mode invents legal articles and cites internal policies that don't exist. On the other, once you give it real context, comes the temptation to dump the entire risk map or sensitive data inventory into some public cloud tool, and the problem stops being invention and becomes exposure of data that the compliance function itself should be protecting. This lesson is about doing both things at once: anchoring AI in the truth of the regulation AND keeping what's sensitive inside the environment the company controls.
The core idea of this lesson. Connecting well takes AI out of "thinking it knows the regulation" and anchors it to YOUR sources (internal policies, technical standards, audit trails), answering with a traceable citation. And there's a case that runs through almost every compliance question: the General Data Protection Law (LGPD). It shows up in almost every process that touches personal data, and that's why it becomes this module's permanent case, always cited literally, article by article, never from memory. Connecting wrong invents a regulation or exposes sensitive evidence; connecting right does both things at once: it anchors the answer and protects the data.
01Why AI on its own invents the regulation
Generic AI answers off the top of its head. It has never read your internal data retention policy, never opened your control map, never seen the audit trail for the process under review. So it does what a rushed analyst would do under pressure: it talks a good game and guesses with confidence.
Think about the difference between two compliance analysts. The first answers from memory and sometimes cites an article that doesn't exist or an internal policy item that says something else. The second, before saying a word, gets up, goes to the regulatory repository, pulls the applicable policy, the audit trail, and the control map, and answers with the document open in front of them, pointing to where each piece came from.
Connecting AI to the regulatory base is turning it into the second analyst. Instead of "thinking it knows the regulation," you anchor it in your sources and it answers by citing the source. The gain is direct: fewer reports with phantom citations, checkable origin, and less rework redoing what AI made up. Fair?
02What "connecting" means: AI starts working WITH your policies and evidence
Connecting isn't magic. It's retrieving the right passage before generating the answer, now pointed at the company's regulatory archive. You index the sources once (internal policies, technical standards, procedures, audit trails, control map) and, with every question, AI searches only the right passages and answers anchored to them, with the citation at the bottom.
The difference from a memory-based summary is what makes this useful in compliance. A question like "is the billing process compliant with the purpose limitation principle" won't find anything with a literal keyword search if the policy uses different words to say the same thing. Meaning-based search finds it by proximity of meaning, so it brings the right passage even if it's worded differently, and it still points to which document and which section it's in.
03The permanent case: the General Data Protection Law (LGPD)
There's a regulation that runs through almost every workflow in this module, and it deserves special treatment: the General Data Protection Law (LGPD), Law No. 13,709/2018. Almost every compliance process touches personal data at some point (the supplier you audit processes customer data, the ethics channel collects data from whoever reports, the training records data on who took part), and that's why it becomes the permanent case we keep coming back to throughout this track.
The practical rule here is strict and non-negotiable: every citation of the General Data Protection Law (LGPD) is literal. Article name, item number, exact text. Never paraphrased from memory. If AI says "in accordance with article 7, item I," you open article 7, item I, and check that it actually says that. It's common for AI to cite a real article with the right number but describe a legal basis that isn't the one in your process, so getting the number right is no guarantee at all that the content matches.
Where this fools you: it seems silly to check a law "off the top of your head" when the article is famous, like article 7 on legal bases. But precisely because it's famous, AI has more training material about it and more confidence to cite it wrong, mixing the right item with the wrong basis. Familiarity isn't the same as precision. Always check, especially on what seems obvious.
04Audit trails: evidence is what gets checked, not what gets summarized
The second source this module connects is audit trails: the log of who accessed what, the approval history for an exception, the record that a control ran. Here the most common error isn't AI inventing a whole document, it's it summarizing the evidence in a way that sounds plausible but doesn't exactly match what the log shows.
The difference between "AI summarized the evidence" and "AI is connected to the evidence" is that, in the second case, every statement in the report points to the exact record: which log line, which date, which system. If the audit report says "the dual-approval control ran on all transactions above the threshold," that needs to be traceable to the log that proves it, not a general impression AI formed skimming through it.
The same practical takeaway that applies to the legal library from memory applies here: a question about the content of a single policy is solved with meaning-based search (Vector RAG), the pattern most people start with. But when the control map has controls that depend on each other (control 12 is only valid if control 4 already ran), the question becomes "follow the connections," and that's where the graph, which stores those relationships between controls and evidence, helps more. Start simple with meaning-based search; when questions start crossing several interlinked controls, flag that map as a graph candidate.
Do it now
Think of a real case from your compliance day where AI answered off the top of its head and got it wrong, or where you were afraid to paste a sensitive document into the tool: your real task.
- List 3 to 5 of your sources that would contain the right answer (internal policy, technical standard, audit trail, control map, article of the General Data Protection Law (LGPD)).
- For each source, mark a label: "can go out anonymized," "stays only in the controlled environment," or "public with no restriction." Justify in one sentence who the sensitive data belongs to there.
- Write the rule your connection would need to have to never let a regulation citation enter a report without pointing to the source document and section.
- If the case involves the General Data Protection Law (LGPD), write the exact article and item number you'll check against the official text before citing it.
You've just designed a compliance connection that anchors AI in the real regulation without letting sensitive data escape.
Practice
1. What is the central gain of connecting AI to internal policies and audit trails (compliance RAG) instead of letting it answer off the top of its head?
2. Why does the General Data Protection Law (LGPD) work as this module's permanent case, and what is the rule for citing it?
3. What is the difference between AI 'summarizing' an audit trail and AI being 'connected' to it?
For the board
On the unconnected AIleft generic it invents statutes and cites internal policies that do not exist.
On connectingthe AI answers from the source and you check the origin. That is what kills the phantom citation.
On evidencean audit trail is verified down to the record that proves it, not summarised into something that sounds reasonable.
Thanks for the feedback. It helps sharpen the next lesson.