Knowledge Engineering · Egress

Your Best Model Is in the Cloud. Your Best Data Cannot Go There.

You want a frontier model's answers on the data that actually matters, and that data is the one thing you cannot send out. So you pick a side: send it and lose control, or wall everything off and let one rule govern all of it. Why? Because the decision is being made once, for a whole workspace, by a person. A governed data layer moves it: the data's own classification decides, for every single request, which model it is sent to for processing.

Infozense Knowledge Engineering · CISO · CIO · GRC · Head of Data · ~6 minutes

This is the Egress piece of the Infozense Knowledge Engineering library. The data layer beneath your agent has five gates: Input (who may ask) · Egress (what may leave) · Output (safe and true) · Action (who approves) · Operational (prove it later). This piece is about the second: Egress, the gate that decides what is allowed to leave your walls, and for which model.

One workspace holds material of mixed sensitivity. A single rule, set to the strictest item present, sends everything to your own hardware, and the frontier model is unreachable even for the published manual.
The published manual is locked down as tightly as the most sensitive thing beside it. Nothing is safer, and the restriction bought nothing.

Somewhere in your organization there is a decision that was made once, quietly, and has governed everything since.

Either your AI platform may call an outside model, or it may not.

  • If the answer was yes. A great deal of your knowledge is now reachable by a system you do not run.
  • If the answer was no. You built inside your own walls, and every question your people ask is answered without the model you would rather have used.

Both answers are defensible. What is not defensible is that the decision was made once, for everything, by a person, in a meeting, before anyone knew which questions would be asked.

Who this is for

Most organizations do not need to think about this. If nothing you hold would matter if it left, then send it out and enjoy the better answers. That is the right call and this piece is not for you.

You need this when two things are true at once: some of your data genuinely cannot go to an outside model, and you want the best available answers on the rest of it. That is most regulated businesses. It is a bank whose product documentation is dull and whose customer records are not. It is a hospital whose clinical guidelines are published and whose patient notes are not. It is any organization where a single body of knowledge contains material of very different sensitivity, which is to say almost all of them.

Walling everything off does not make your answers bad

There is a version of this argument that says the sovereign choice means settling for less, and it is wrong. You can run a large, genuinely capable model on your own hardware. Plenty of organizations do. If someone tells you that keeping data inside your walls means accepting a poor answer, they are selling you something. What they are selling is the idea that you only get two choices: everything out, or everything in.

Your data types are not uniform. The rule you wrote is.

When one decision governs all the data in a workspace, it has to be set to the strictest thing that workspace holds. For example, the product manual, the published policy, the release notes, the training deck, the whole non-sensitive majority of your corpus: stored alongside sensitive material, all of it is grouped at the same highest sensitivity level.

That is not a security gain. Nothing is safer because the product manual was treated carefully. The cost is two things. The first is opportunity: the better model was safe to use on that material, and one decision put it out of reach. The second is spend: your own hardware has to carry work that never needed to be inside your walls at all.

The waste is invisible, which is why it survives. Nobody files a ticket saying the answer about the published policy could have been better. They just get a worse one, forever, and never learn what they missed.

The decision belongs to the data, not the deployment

One request assembles knowledge of mixed sensitivity. The strictest classification present decides the destination: public material may reach the frontier model, confidential material stays on your own hardware, and restricted material is refused. Here the strictest is confidential, so the request stays on your own hardware.
The strictest thing in the request decides. No person sorts them, because no person could.

The fix is not choosing between walling everything in and sending everything out. It is moving the decision that chooses whether data is processed outside or inside.

In a governed data layer, sensitivity is a property of the knowledge itself, carried on the material rather than on the deployment. A single request may assemble material carrying different classifications, and the governed data layer uses the strictest one present as the classification for the whole request. Not the average, not the first, not the workspace's setting. The maximum of what is actually in front of it, this time.

That single value chooses where the request is allowed to go. An operator maps each level of sensitivity to a destination once: public material may go to the frontier model, confidential material stays on your own hardware, and the most restricted material is not sent to any model at all. Every request is then routed to the destination its own classification sets.

The result is that both things you wanted are true at the same time. The question about the published policy reaches the best model available. The question that touches your most sensitive material never leaves the building. No human sorted them, because no human could, at the speed and volume an AI is asked questions.

That also settles the second cost. Each level of sensitivity points at its own destination, so the model on your own hardware only has to be sized for the material that genuinely has to stay there. You are not buying capacity beyond what you actually need.

What happens to the thing that may not leave

The interesting case is the last one, and it is where most systems quietly fail.

A request whose material is too sensitive for any destination has to be refused. Not sent somewhere less good. Not stripped down and sent anyway. Refused, as a real outcome that the design plans for rather than an error it stumbles into.

Two properties make that trustworthy, and both are the opposite of what a hand-rolled version usually does.

Unlabeled is treated as most sensitive, not as safe. Knowledge that arrives without a classification, or with one the layer does not recognize, or with something malformed where the label should be, is handled as the most restricted thing it could be. A gap in your metadata makes the layer more careful, never less. That is the single most important line in the design, because in every system that has ever leaked, something was unlabeled and something assumed that meant fine.

A mis-configured route cannot widen anything. Each destination carries its own ceiling for what it will accept, and that ceiling is enforced underneath the operator's map. If someone maps a sensitivity level to the wrong destination, the destination refuses the material anyway. The map can make the system stricter than intended. It cannot make it looser.

The bottom line

You were never really choosing between a better model and safer data. You were choosing how often to decide.

Deciding once, for everything, is what forces the trade. It makes your least sensitive material live under your most sensitive material's rules, and it calls that security when it is really just a blunt instrument left in place because nothing finer was available.

Decide per request and the trade-off is gone. The data's own sensitivity picks the destination every time, without a person in the loop, and the answer you get is the best one its classification allows.

You don't build this. You configure it.

This is not a model you train or a rule engine you write. Infozense Knowledge Engineering ships the governed layer with the Egress gate already in place. You classify your knowledge, you map each level of sensitivity to a destination once, and every request from then on is routed to the destination its classification sets: the strictest thing in the request decides, unlabeled material is treated as restricted, and a destination refuses anything above its ceiling regardless of what the map says.

Your best model stays available. Your data that cannot leave does not leave. Nobody has to choose again.

Infozense Knowledge Engineering

Holding data you cannot send to an outside model, and answers you wish were better?

That's the conversation we have best. Bring one real workspace, and we'll walk through what would route where.

Let's talk →

contact@infozense.com  |  +66-82-242-4008  |  Bangkok, Thailand