Est.

Confidential Computing Enclaves as the Technical Backstop Against AI Vendor Prompt Exposure

Hardware-enforced enclaves prevent cloud providers from reading AI prompts during inference.

Editor-at-Large · · 10 min read
Cover illustration for “Confidential Computing Enclaves as the Technical Backstop Against AI Vendor Prompt Exposure”
Zero-Trust and Provable Data Privacy · October 8, 2026 · 10 min read · 2,309 words

Confidential computing enclaves solve a problem that standard encryption cannot touch: the moment data is unlocked for processing. An enterprise that wants an AI vendor to never see its internal prompts and queries needs hardware-enforced isolation during inference itself, not just encryption before and after it, and that is precisely what Trusted Execution Environments like Intel TDX, AMD SEV-SNP, and attested GPU stacks are built to guarantee. This piece lays out how that guarantee works, where production systems like EnclaveX and Talaria already deliver it, and where the honest limits of the technology sit.

The standard encryption stack leaves AI prompts exposed at the moment they matter most

Encryption at rest and encryption in transit are mature problems. Neither protects it during the interval that actually matters for an AI system: the moment the prompt is decrypted into memory so a model can read it, tokenize it, and generate a response. Every inference call walks through that interval, and standard encryption has nothing to say about what happens there, because by design it has already handed the data over in plaintext.

In a cloud-hosted LLM deployment, you don't control that infrastructure, and you can't inspect it either. Research on confidential inference for cloud-based LLMs treats prompt and response confidentiality as a deployment prerequisite for regulated industries precisely because of this gap, and the restrictions that JPMorgan, Citi, and Goldman Sachs have placed on external AI tools serve as market evidence that the risk is not abstract. Those restrictions were driven by compliance concerns over third-party software and the exposure of sensitive data, and they reflect an operational judgment that the processing layer cannot currently be trusted.

Regulators have reached a similar conclusion from a different direction. U.S. state privacy laws such as California's CPRA and Colorado's CPA are pushing toward similar reviews, and FTC enforcement is addressing AI data practices under existing consumer protection authority. The regulatory pattern treats prompts as legally significant data, not incidental traffic.

The threat does not require an external attacker. The fix has to live in the processing layer too.

Confidential computing enclaves and hardware enforcement that changes the trust equation

A Trusted Execution Environment builds a hardware-enforced encrypted zone inside the processor itself, and code and data stay sealed there even from the host operating system, the hypervisor, and the cloud administrator who manages the machine. The boundary is enforced by the chip itself, below the layer where operating systems and administrators operate, and no sufficiently privileged user can override it.

Three implementations have reached broad deployment. Intel TDX provides encrypted memory isolation at the confidential VM level and is available on Google Cloud and Microsoft Azure, though not on AWS, which uses its own Nitro Enclaves approach instead. AMD SEV-SNP offers confidential VM protection, but on major cloud platforms you have to opt in. Confidential GPU capability, covered later through NVIDIA's H100 and H200 lines, extends the same principle to the accelerators that actually run model inference.

Hardware enforcement only matters if it can be verified, and that verification comes through remote attestation. A report signed by the hardware manufacturer's root of trust cryptographically proves the identity and integrity of the software running inside the TEE before any sensitive data is released to it. An enterprise does not need to trust the cloud provider's word that a workload is isolated. It can cryptographically confirm it before handing over anything sensitive.

The specific threat that enclaves are designed to foreclose, vendor-side prompt exposure in AI-as-a-service

AI-as-a-service creates a trust problem that runs in both directions, and no software-only solution resolves it cleanly. If you want a useful response, you must share sensitive prompts with the model provider. Providers, in turn, must expose valuable, often proprietary model weights to the environment running inference. Each side needs assurance that the other cannot walk away with what it brought to the transaction, and historically that assurance has rested only on contracts and reputation.

Confidential inference resolves both directions of that problem at once. Model weights are encrypted with a key that is released only to a verified TEE, so a user cannot extract the proprietary model even while using it. Prompts, the key-value cache, tokenizer state, and responses are sealed inside the same TEE, so the provider cannot read what the user sent or what the model returned. A production LLM enclave isolates the complete inference context this way: prompts, model weights, tokenizer state, KV-cache, evaluator code, and responses, all shielded from the host operating system, the hypervisor, and the cloud provider operating the hardware. Neither party has to trust the other's intentions, because the enclave enforces the boundary regardless of intention.

A working system illustrates what this looks like assembled end to end. Attestation in this stack operates at three levels simultaneously: VM level through Intel TDX, GPU level through the NVIDIA H200, and application level through SCONE, with each layer verified before any sensitive material is released. That layered verification is what separates a genuine zero-trust architecture from a vendor simply asserting that its infrastructure is secure.

Talaria's protection of responses as well as prompts when model weights stay in the cloud

Even a well-built CVM approach has historically left half the conversation exposed. The prompt goes into the secure environment, but the inference output travels back out through cloud GPU infrastructure the provider controls. The response, the half of the exchange that often contains the most sensitive synthesis of information, remains visible to the provider.

Talaria addresses this by splitting the inference pipeline rather than trying to seal the entire thing in one place. Sensitive, weight-independent operations run inside a Confidential Virtual Machine, and the client controls it. Weight-dependent computations, the parts that require the provider's proprietary model, are offloaded to cloud GPUs. The exchange between the two environments is secured by what Talaria's authors call a Reversible Masked Outsourcing protocol, which obscures intermediate data before it ever leaves the client-controlled environment.

The output of this arrangement is identical to what the unprotected model would produce, while token reconstruction accuracy, the measure of how much an adversary could reconstruct of the original data from what the cloud GPU processes, drops from a near-perfect baseline to an average of 1.34%. The cloud provider is still doing the computation. It cannot read what it is computing.

Talaria's own framing of the problem is useful well beyond Talaria itself: model privacy, model performance, and model efficiency function as three requirements in tension, and a system that sacrifices output quality or introduces overhead too large to tolerate is not something an enterprise can actually deploy. That tension is the right lens for evaluating any confidential inference system a vendor proposes, and it is why response-level protection has become a live and actively advancing area of the field alongside prompt-level protection.

Performance overhead across current confidential inference configurations

Overhead is the objection people raise most often against confidential inference, so it deserves a direct answer. It is not a fixed cost. It varies by hardware generation, by how the stack is configured, and by whether avoidable inefficiencies have been eliminated before benchmarking even begins.

On current-generation hardware, benchmarking confidential inference on NVIDIA B200 GPUs paired with Intel TDX found that a correctly configured stack adds only a low single-digit percentage of throughput overhead, a cost comparable to the kind of variation that shows up from routine infrastructure changes anyway. That is the optimized case, and it demonstrates that confidential inference does not have to impose a meaningful tax on production workloads.

The unoptimized case tells a different story. NVIDIA's Hopper H100 confidential computing capabilities show a performance impact that sits within a range enterprises have found workable for production deployment.

No single overhead figure should be taken as representative of what an enterprise will experience. The responsible approach is to benchmark the specific workload, batch size, and stack configuration against the hardware actually intended for deployment, because the distance between a poorly configured stack and a well-configured one is large enough to decide whether a deployment is viable. Performance, in other words, has become an engineering variable to optimize. The honest objections to confidential computing now live elsewhere.

The limits enclaves cannot overcome, what TEEs do not protect against

An enclave seals memory and attests that specific code is running. It cannot interpret what that code is being asked to do. A prompt-injected instruction executes exactly as written inside a perfectly attested enclave, because attestation proves the identity and integrity of the software, not the trustworthiness of the inputs handed to it. It is outside the scope of what an enclave was ever designed to check.

The EchoLeak vulnerability in Microsoft 365 Copilot illustrates the point concretely. A zero-click prompt injection came through a single email, and it was able to silently exfiltrate enterprise data, reaching internal files, emails, and documents well beyond the email that carried the attack. That attack would have run identically inside a correctly attested enclave, because the enclave would have faithfully executed whatever instructions the model was given, with no way to judge whether those instructions were legitimate.

The hardware itself is not immune to physical attack. WireTap, presented at ACM CCS 2025, showed that a physical interposer monitoring the DDR5 memory bus could extract ECDSA attestation keys from Intel's Provisioning Certification Enclave, the keys that underpin the entire SGX and TDX attestation chain. Side-channel leakage adds another category entirely: timing patterns, cache and bus contention, page-fault behavior, GPU residual-state leakage, and traffic-analysis signals can all expose information to a capable adversary without the enclave's memory boundary ever being breached.

None of this erases the value of hardware isolation. It narrows the claim to its proper scope. TEEs foreclose vendor-side access to plaintext prompts and responses during inference, which is the specific threat this piece is built around, and they do so with a rigor no access-control policy can match. They do not, on their own, foreclose prompt injection, physical-access attacks on the hardware supply chain, or microarchitectural side channels, and any enterprise evaluating confidential inference should understand that distinction before treating a TEE as a complete security program rather than the backstop for one particular, serious problem.

PIM-Enclave and the direction of research addressing TEE side-channel exposure

Conventional enclave design still requires data to move between the CPU and memory over a bus, and that bus can be observed. Encrypting the memory itself does not hide the pattern of movement across it, and that movement pattern is what an attack like WireTap exploits.

PIM-Enclave proposes addressing this at the architectural level. The idea is to bring confidential computation inside memory itself, placing processing logic adjacent to data storage so that sensitive computations never need to traverse the external memory bus. That removes the observable channel rather than trying to obscure what travels through it. Evaluation of PIM-Enclave finds that it can deliver side-channel-resistant secure computation offloading while running data-intensive applications, with negligible performance overhead compared to the baseline processing-in-memory model, and that matters for AI workloads because they are inherently data-intensive.

PIM-Enclave is a research direction, not a deployed product, and it should be read as evidence of where the field is headed. Confidential computing is not a static technology that has reached its ceiling. The limits documented in the previous section, bus-level side channels chief among them, aren't left as permanent tradeoffs; people are working on them at the hardware architecture level. An enterprise adopting confidential infrastructure now is betting on a technology still improving, which is a materially different proposition than adopting one that has already settled into its final form.

πCreds and the extension of the TEE trust model to verifiable claims over sensitive data

The trust model built around TEEs, hardware isolation plus cryptographic attestation, was developed to answer a narrow question: can a piece of code be proven to be running, untampered, inside a sealed environment? Confidential inference applies that question to a single use case, protecting prompts and responses during a model's forward pass. But the same underlying mechanism, a verifiable claim about what happened inside a protected boundary, generalizes well beyond inference.

That generalization matters because enterprises do not only need to protect what a model sees during a single query. They need to make verifiable claims about sensitive data more broadly: that a dataset was processed under a given policy, that a credential was issued based on specific underlying attributes without exposing those attributes themselves, or that a computation over regulated data complied with a rule without revealing the data used to satisfy it. πCreds represents this extension of the TEE trust model, moving from "this code ran untampered" toward "this claim about sensitive data can be verified without exposing the data itself." The mechanism is the same hardware-rooted attestation that underlies confidential inference. The application widens from a single inference call to the broader category of claims an organization needs to prove about information it is obligated to protect.

Hardware-enforced trust boundaries and privacy-preserving architecture serve the same goal from different angles: both eliminate the need for users to trust vendor claims about data isolation, instead building systems where sensitive data either never enters an uncontrolled processing environment or is cryptographically sealed from the provider itself. In confidential inference deployments, prompts and model weights are isolated inside the TEE from the provider. In privacy-first AI architectures, the same isolation is achieved through design, and both approaches converge on the same outcome: the vendor operating the infrastructure is structurally unable to access the sensitive data flowing through it, whether that data is a prompt, a response, or a credential being proven true, which is the structural guarantee Confidant AI was built around from the start. The question for any enterprise evaluating a vendor is no longer whether such isolation is achievable. It is whether a given vendor has actually built it in, verifiably.

Sources

  1. EnclaveX: End-to-End Confidential AI with CPU/GPU TEEs Robert Schambach
  2. $\pi$Creds: Privately Inferred Credentials
  3. PIM-Enclave: Bringing Confidential Computation Inside Memory
  4. Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
  5. Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
  6. Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study

More in Zero-Trust and Provable Data Privacy