Knowledge Engineering · Governance

You Governed the Agent. You Forgot the Data Layer Beneath It.

Every AI agent project focuses on two things: how the agent works (orchestration) and what keeps it in line (guardrails). Both are the agent itself, the part you can see. Beneath it sits the data layer, where your regulated risk actually lives: what the agent can reach, what may leave your organization, and what you can prove months later. Almost no one is governing it.

Infozense Knowledge Engineering · CIO · CISO · GRC · Head of Data · ~5 minutes

This is the Governance piece of the Infozense Knowledge Engineering library. The data layer beneath your agent has five gates: Input (who may ask) · Egress (what may leave) · Output (safe and true) · Action (who approves) · Operational (prove it later). The rest of this piece goes through each one.

Walk into any AI agent project today, and you'll find everyone looking in the same direction.

They're looking at the agent. How it plans and routes is the orchestration work (tools like LangGraph and Dify). What keeps it in line is the guardrail work: staying on topic, stopping jailbreaks. This work is real and necessary, and there are plenty of tools for it.

But that is only the agent. Everyone watches it so closely that no one looks underneath, at the data layer the agent stands on. That layer gets ignored for a simple structural reason: it isn't “the agent,” so it sits outside every framework built to govern the agent. Orchestration decides how the agent works. Guardrails keep it in line. Neither one decides what it can reach, what leaves your organization, or what ends up in the record. And that is what a regulator checks: not the words the model says, but the layer beneath.

Who this is for

Most organizations do not need a governed data layer. If your AI answers from public information, a model guardrail is enough. You can stop here.

You hold data that must not leak: patient records, contracts, regulated files. You want real AI working on that data, not just on the information anyone can find on your public website. But the best AI does not live in your building. It runs in the cloud, on computers you do not own. And whichever way you handle that, someone will eventually ask you to prove where your data went and who saw it. That is when you need a governed data layer.

Without that layer, every request becomes a judgment call. Send it to the cloud AI and you get the best answer, but if it carried something sensitive, that data now sits on someone else's computers and you cannot get it back. Send it to the AI on your own computers and nothing leaks, but if it was an ordinary request with nothing to hide, you gave up the best answer for no reason. So every request has to be judged before it goes anywhere, and every judgment written down. You can put that in a policy. In practice it is impossible. An agent sends more requests in an hour than a person could review in a day, so the judging either happens in seconds with no record, or it does not happen at all.

The governed layer takes that decision away from the person and gives it to the data. Every piece of your data carries its own classification, set once when it enters. Before a request goes anywhere, the layer reads that classification and routes on it: ordinary data goes to the best AI in the cloud at full power, sensitive data goes to the AI on your own computers and never leaves. You get the strongest answer the data allows, every time, and no person has to make that judgment in the moment. The same rule applies to every request, and every routing decision is written down. That is one of five gates this layer runs, and you will see all five below.

A guardrail watches words, not data

Here is a distinction that often gets blurred.

A guardrail like NVIDIA's NeMo Guardrails wraps the model. It sits between your application and the model, and it checks the conversation: the prompt going in, the answer coming out, and whether a tool call is correctly formed. It does that job well. But everything it governs is the model's words and the model's calls.

A whole layer sits outside what a model guardrail is built for. A model guardrail does not decide who is allowed to ask, or for whose data. It does not decide what classification of data may leave your organization. It does not keep an audit record no one can change. It does not require a human to approve an irreversible action. None of those are about the words the model produces. They are about the data, the identity, and the record.

That is the difference between a guardrail and a gate. A guardrail watches the model's words and judges each message: on topic or not, a jailbreak or not, a tool call formed correctly or not. A gate does not judge the message. It enforces a fixed rule about the data and the person behind the request: who may ask, what may leave, what gets written to the record, applied the same way every time. Your organization's regulated risk lives with the gates, and there are five of them.

The five gates

A cross-section: a crowd on the roof studies an AI on a pedestal, while five labeled vault gates below govern the data flowing through: Input (who may ask), Egress (what may leave), Output (safe and true), Action (who approves), Operational (prove it later).
Five gates on the data layer beneath the agent.

Input: who is allowed to ask? One agent often serves many teams. It must not let an HR user pull Finance's records, or let one client's data appear in another client's answer. Orchestration assumes the person asking is allowed. The input gate checks first: who is asking, and for whose data, before it touches a single document.

Egress: what is allowed to leave? The moment your agent calls an outside model, your data is on someone else's server. The egress gate reads the classification of what is about to leave and decides where it goes. Confidential data stays on your own model, inside your organization. Only cleared data leaves, and only after it is redacted. This is the gate that stops a patient record or a deal memo from quietly becoming training data for someone else's model.

Output: is the answer itself safe and correct? Even when the data never leaves, the answer can still cause harm. It might show a private name it should have hidden, or cite a source that does not exist. The output gate hides what should not be shown, and checks that every citation points to a real passage. The answer cannot leak, and it cannot make things up.

Action: who approves before something cannot be undone? Reading data can be undone. Changing or deleting it cannot. The operations that matter here are not exotic: permanently deleting someone's records, erasing their data to honor a “forget me” request, or removing files when a retention deadline arrives. Each one is permanent. Each one should wait for a named person to approve it. The action gate holds these changes until they are approved. Nothing destructive happens on its own.

Operational: can you prove it, months later? When an auditor asks what the agent did back in March, “we think it was fine” is not an answer. The operational gate keeps a permanent record of every question, every routing decision, and every action. It lives in a log that no one, not even you, can change afterward. Governance you cannot prove is governance you do not have.

The bottom line

You didn't get less governance than you thought. You aimed it at the agent, not the layer beneath. Orchestration and agent guardrails govern how the agent works and what it says. Keep them. You need them. The five gates govern what it can reach, send, and prove. One keeps the agent in line. The other protects your organization. No one sold you the second because everyone was busy with the first.

You don't build this. You configure it.

Five gates can sound like a year of engineering. They are not, because you do not build them. Infozense Knowledge Engineering ships the governed layer with all five already in place: who may ask, what may leave, what is shown, what needs approval, and what is recorded. You set the policy for each one against your data and your rules, and the layer enforces it from the first request. You configure the gates. The plumbing is already there.

That is the real head start: governance on day one, not at the end of a project you never quite finish.

Infozense Knowledge Engineering

Putting an agent into a regulated process, and not sure the layer beneath it is covered?

That's the conversation we have best. Bring one real workflow, and we'll walk through it gate by gate.

Let's talk →

contact@infozense.com  |  +66-82-242-4008  |  Bangkok, Thailand