Est.
FeaturesLong read

Enforcing Data Use Boundaries Between AI Features at Runtime

Runtime enforcement separates AI policy control from the model itself.

Editor-at-Large · · 10 min read
Cover illustration for “Enforcing Data Use Boundaries Between AI Features at Runtime”
Features · October 9, 2026 · 10 min read · 2,261 words

A written policy says what an AI system should do, but that's before anyone has built the system that does it. That is the entire problem this piece sets out to map: data use boundaries between AI features hold only when enforcement happens at runtime, at the level of the individual action, because a document drafted before deployment cannot anticipate how features chain together, what context a session carries, or what an autonomous agent decides to do once it is live. A policy is a static artifact. A production AI system is not. Features call other features, and agents hand tasks to other agents, so each step happens with information the policy's authors never had in front of them when they wrote the rules down.

If you write guardrails as paragraphs inside a system prompt, the model may treat them as suggestions, not enforceable boundaries, because a prompt injection or a misbehaving agent can reach and manipulate that same surface. AI agents decide, chain, and execute actions while a session is still open. The threat surface is generated while the system runs rather than fixed in advance, so any control point that only evaluates the prompt rather than the live request is evaluating the wrong thing.

Two documented failures show what this looks like outside the abstract. The EchoLeak vulnerability in Microsoft 365 Copilot showed that a zero-click prompt injection could silently access and exfiltrate enterprise data across feature boundaries, and no user had to do anything to trigger it. The exploit moved data out of the system by exploiting how Copilot's features passed information to one another, which is exactly the kind of chaining a pre-deployment policy has no way to foresee. System prompt leakage tells a related story. It ranks among the most exploitable bug classes across models currently in production, and the reason is structural rather than incidental: system prompts routinely contain API endpoints, internal tool names, role definitions, and the access boundaries meant to keep one feature from reaching into another's data. When that prompt leaks, the map of the system's internal boundaries leaks with it.

None of this stays confined to engineering postmortems. The gap between what a policy says and what a system does at runtime is a financial and legal exposure sitting inside production infrastructure, not a defect to be logged and scheduled for the next sprint.

The regulatory and operational pressure that makes closing this gap urgent now

You can no longer treat runtime AI governance as something to put in place eventually. The regulatory clock already started. Providers of general-purpose AI models have carried obligations under the EU AI Act since August 2, 2025, and the European Commission's enforcement powers, along with those of national authorities in member states, became active on August 2, 2026. One provision says you have to record events automatically across the system's lifetime, so logging has to support traceability for risk identification, post-market monitoring, and ongoing operational oversight. You can't assemble a log file after an incident and call that requirement met. The recording has to happen as the system runs.

Frameworks built outside the EU, including NIST's AI Risk Management Framework and ISO 42001, overlap structurally with the Act's requirements and share the same underlying demand: evidence that enforcement is actually happening, not a binder of policies describing what should happen. Passive monitoring cannot produce that evidence. Governance tools have to enforce which AI tools are permitted, how those tools can be used, and what data they're allowed to touch, and they have to generate the records that regulators will ask to see. A major platform processes millions of interactions each year, so a single cross-tenant data exposure in a regulated industry is not a contained incident, because the same architectural gap that allowed it sits across every tenant the platform serves. That is the backdrop against which the rest of this piece makes its case: the mechanisms described below are not best practices for a well-run AI team. They are what the current regulatory floor requires.

Runtime enforcement means the enforcement layer as a separate trust domain

With runtime enforcement, you put the control layer outside the model and outside whatever trust domain an attacker or a misbehaving agent can influence. The enforcement layer needs a vantage point the attacker cannot reach by manipulating the model's input.

In practice, you see this as a proxy that sits between the application and the model, inspecting each request at several distinct points as a check on the model. Input checks examine prompts before they ever reach the model. Before what the model generates reaches a user or passes to a downstream system, output checks evaluate it. If an agent can take action, the layer extends further, and it intercepts tool calls before the tool executes them. The architecture resembles a policy decision point paired with a policy enforcement point: evaluate the context surrounding a request, then decide to block or allow it before the action happens, not after the fact.

Fail-closed design has to be the default setting for this layer. If a request arrives without sufficient context, or triggers a system fault along the way, it gets denied automatically, not passed through on the assumption that it's probably fine. The need for that default is not theoretical. Coralogix's analysis found that 97 percent of organizations that had an AI model or application breach lacked proper AI access controls at the time of the breach. The enforcement layer has to evaluate a minimum set of inputs on every single action it reviews: which agent is acting, what policy applies to the customer or workload involved, what resource the action targets, and what operation is being attempted. Changing any one of those four inputs can change the correct answer to "should this be allowed.

This has to happen before execution rather than after because an agent that deletes a record does its damage in the act of deleting it, and no output filter reviewing the result afterward can undo that deletion. For any action carrying that kind of risk, pre-execution interception is the only design that actually works, and every mechanism described in the sections that follow is a different way of making that interception possible.

Scoped identity and per-action authorization as the foundation of feature-level boundaries

Static permissions assume a system that behaves consistently across a session, and agentic AI systems do not behave that way. The question an enforcement layer needs answered is what this specific agent is permitted to do right now, for this specific action, against this specific resource, not whether a user logged in with valid credentials an hour ago.

Task-scoped, just-in-time permissioning follows directly from that question. A related design choice separates what an agent can attempt from what it can complete: an agent may be allowed to propose or prepare an action, drafting the request and assembling its parameters, while still being blocked from executing that action until the enforcement layer confirms the request falls within bounds. That separation reduces accidental overreach, because a prepared-but-unexecuted action gives the system a checkpoint it wouldn't otherwise have, and it makes the approval logic auditable, since you now have a discrete decision recorded between proposal and execution.

Multi-agent systems and agent-chaining environments raise a version of this problem that is specific to chaining. Each agent in a chain needs its own distinct, attributable identity with its own granular permission scope, because if it doesn't have one, a compromised output from one feature can propagate downstream with nothing checking it at the next hand-off. Governance tooling has started tracking MCP-specific enforcement directly, and some platforms now pair MCP policy enforcement with audit logs as part of broader agent governance.

The strongest version of this argument appears in the EHV architecture, which applies a zero-trust model by separating three functional roles across distinct trust boundaries. It is a verified property attached to the message itself. Policy bundles carry this down to the rule level: each bundle packages identity rules, data classification rules, and route-level rules into a single versioned unit, and when a request arrives, the enforcement engine matches it against the active bundle and records the decision made, the specific rule that produced it, and a reason code explaining why.

Information flow control: tracking sensitive data as it moves across feature boundaries

Scoped identity answers who is acting and what they're permitted to do, but it does not answer a separate question: what happens to the data itself once a feature has legitimately received it and passes it along to another feature downstream. An agent can be fully authorized to read a record and still pass that record somewhere it should never go, simply because nothing attached to the data itself carried a rule about where it was allowed to travel next. The data needs its own label, one that moves with it through the system and constrains what any downstream feature is allowed to do once that label is attached.

Information flow control, or IFC, is the mechanism that does this. CaMeL, published by Debenedetti and co-authors in 2025, extracts the control flow and data flow out of a trusted query into an explicit program, so that untrusted data entering the system cannot influence what the program does next, and attaches capabilities to that program that gate which tool calls are allowed. Microsoft Research's FIDES, from Costa and co-authors, takes a related approach with a planner that tracks confidentiality and integrity labels as execution proceeds, enforces security policy against those labels, and selectively reveals or hides information across the agent's working context depending on what the labels permit. A separate system, VIGIL, enforces behavioral specifications for AI agent skills at runtime, and it operates at the boundary of the skill itself to catch violations as they happen, rather than trying to anticipate and block every possible bad input before the agent ever runs.

None of this is free of friction, and the friction has a name: label creep. A 2026 survey of the field notes that IFC defenses work well when you have full visibility into a single agent, but it stays unclear how to generalize that approach to orchestrations built from multiple black-box agents running on proprietary commercial models, where no single party can see every step. A related failure mode compounds the first: strict IFC without any recovery mechanism permanently strands downstream execution the moment an agent ingests untrusted data, because labels in most of these systems only ever become more restrictive as execution proceeds, never less, and recovery gets pushed into fragile workarounds like retrying the prompt or ad-hoc patches in application code. APPA, from Kravchenko and co-authors in 2026, treats this as a design problem rather than an acceptable cost, building explicit recovery semantics into its own IFC system as a core architectural feature.

A deeper requirement underlies all of this: without it, none of it holds. Every document a retrieval system pulls into context has to be treated as untrusted input, full stop, because an attacker who can write to a page that a retriever indexes has effectively written directly into the prompt. IFC enforcement that only starts at the tool-call boundary misses this. It has to begin at the point where data first enters the context, which in most retrieval-augmented systems is earlier than most teams assume.

Enforcement location: application-layer versus OS-level interception

Application-layer guardrails, including CaMeL, FIDES, and the broader family of harness-boundary approaches, all intercept at the edge of the agent framework itself. They see harness-mediated requests, the calls and responses that pass through the agent's own execution loop, but none of them see what happens at the system level once a tool call has started executing outside that loop.

ActPlane, a 2026 proposal, argues that closing this gap requires pushing enforcement down to the operating system kernel, where every execution path eventually has to pass regardless of which application layer initiated it. ActPlane lets agents declare their own policies, but it enforces those policies in the OS kernel, and it pairs that enforcement with semantic feedback and isolation, so it closes the space between tool-level checks and system-level consequences.

The strongest objection to treating OS-level enforcement as sufficient on its own points back to where this piece started. No matter how an in-process system is built, it enforces its rules within the same trust domain that an attacker can influence through the model itself, and none of them bind a cryptographic per-message caller identity the way the EHV architecture does. A gateway sitting in a genuinely separate trust domain complements OS-level enforcement rather than duplicating it, since the two catch different classes of failure: one sees system-level effects the application layer misses, the other verifies identity in a way no in-process check can forge.

The adaptive intelligence architecture for real-time AI policy enforcement is one attempt to resolve this rather than leave it as a standing disagreement between two camps. It separates policy decisions from the inline enforcement points that carry them out, and it draws on multiple context sources and audit records so it can support accurate verdicts and keep tuning those verdicts over time. That separation, between deciding and enforcing, is the same principle this entire piece has traced from a different angle in every section before it: a policy document cannot enforce itself, and only a layer built to sit outside the system it governs, evaluating real actions as they happen, can close the distance between what an organization says its AI will do and what it actually does once it's running.

Sources

  1. AI Guardrails in 2026: How They Work and How to Implement
  2. Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architecture for Agentic AI Systems
  3. Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
  4. VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills
  5. ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses

More in Features