Every serious enterprise AI contract now promises that the vendor will not train on your data. Far fewer vendors can prove that nobody, including themselves, can read that data while the model is working on it.
That gap between a promise and a proof is where the next round of enterprise AI procurement will be decided. On 29 September 2026, OpenAI put a name on it at DevDay: Private Intelligence, a programme built around a technology called confidential inference.
This guide explains what confidential AI inference is, how it differs from the privacy clauses already in your contracts, and how to decide which of your workloads actually need it.
What is confidential AI inference?
Confidential AI inference means running an AI model inside hardware-protected memory, called a trusted execution environment, so that prompts, documents and answers stay encrypted even while the model processes them. The cloud provider, the model vendor and their staff cannot read the data, and the customer can verify this cryptographically.
Most data protection in enterprise IT covers two states: data at rest, encrypted on disk, and data in transit, encrypted on the network. Confidential computing covers the third state, data in use, which has historically been the exposed moment.
When a large language model answers a question about a client contract, the contract text must sit in plain form in memory for the computation to happen. Confidential inference keeps that memory sealed inside the processor and, increasingly, inside the GPU itself.
The practical result is a shift in what you are trusting. Instead of trusting a vendor's policy, you are trusting a chip's hardware isolation and a signed report proving which software is running.
Why did confidential inference become a board-level issue in 2026?
Confidential inference moved onto board agendas because frontier models now handle source code, legal files and unpublished research, and enterprise buyers started demanding technical proof rather than contractual promises. OpenAI's Private Intelligence announcement at DevDay on 29 September 2026 made the issue mainstream for any organisation evaluating AI vendors.
According to VentureBeat’s coverage of DevDay, OpenAI split Private Intelligence into two parts: Zero Data Retention with Private Safety Processing, available now, and Private Inference, a preview combining confidential computing with verifiable controls, due this autumn.
The same report cited The Information's finding that Palantir, Nvidia and Booz Allen Hamilton had restricted model use or sought stronger guarantees over fears that proprietary data could benefit AI providers. When companies of that calibre push back, procurement teams everywhere take note.
Analyst framing points the same way. Gartner’s Top Strategic Technology Trends for 2026 lists confidential computing as a core architecture trend and predicts that by 2029, more than 75% of operations processed in untrusted infrastructure will be secured in use by confidential computing.
For a Hong Kong COO, the takeaway is simple. Within two procurement cycles, "can you prove you cannot see our data?" will be a standard RFP question, and vendors without an answer will be at a disadvantage.
How does confidential AI inference actually work?
Confidential inference works in three steps. The model runs inside an isolated, encrypted region of the CPU and GPU. The hardware produces a signed attestation report describing exactly what software is loaded. Only after verifying that report does the customer's system release encryption keys, so data is never decrypted anywhere unverified.
The building blocks, in plain terms--- Trusted execution environment (TEE): a hardware-sealed area of memory that the operating system, hypervisor and cloud administrators cannot read. CPU families from Intel and AMD offer this, and NVIDIA's data-centre GPUs since the H100 generation add a confidential computing mode.
--- Remote attestation: the chip signs a report of the exact model, runtime and configuration it is running. Your systems check that signature before sending anything.
--- Key release: encryption keys are released only to an environment that passes attestation. If the software changes, the keys stay locked.
--- Verifiable controls: logs and policies that an auditor, not just the vendor, can check.
This is not purely an OpenAI idea. Apple has used a similar architecture for Private Cloud Compute since 2024, and Microsoft announced a preview of Azure AI Confidential Inferencing the same year. What changed in 2026 is that a frontier model vendor is now promising it for its own flagship models.
There is a performance cost. Encrypting memory traffic between CPU and GPU adds overhead, and the size of that overhead varies by workload. Any vendor quoting a figure should show its own benchmark for a workload similar to yours.
How is confidential inference different from zero data retention and no-training clauses?
Zero data retention and no-training clauses are contractual promises about what a vendor will do with your data after processing. Confidential inference is an architectural control over who can read your data during processing. The first relies on trust and audit rights; the second relies on hardware isolation and cryptographic verification.
Most enterprise AI agreements already include three privacy commitments: no training on customer data by default, limited retention, and restricted staff access. These are valuable, but they are enforced by policy and contract.
OpenAI's own design illustrates the nuance. Under Private Safety Processing, content flagged by a safety classifier is encrypted and written to storage the customer controls, such as an AWS S3 bucket, with a 30-day time-to-live so automated safety checks can review it inside a hardware-attested runtime.
That means "zero data retention" does not mean no encrypted copy exists anywhere. It means the vendor does not hold a readable copy. Your compliance team needs to account for that customer-side storage in retention schedules and records policies.
A useful way to brief your board: contracts tell you what the vendor promises; confidential computing tells you what the vendor is technically able to do. Mature AI governance needs both.What does confidential inference mean for Hong Kong enterprises?
For Hong Kong enterprises, confidential inference matters in three ways: it strengthens the case for using overseas cloud AI under the PDPO, it gives regulated sectors a stronger answer to data-handling questions, and it forces a rethink of the private-versus-cloud decision. Availability, however, remains the binding constraint locally.
Availability comes first. OpenAI has not served Hong Kong first-party since July 2024, and PTS Consulting’s 2026 review notes that this had not changed as of mid-2026. Hong Kong organisations will therefore meet confidential inference mainly through hyperscaler platforms and through private deployments, not directly through OpenAI.
Regulation is the second driver. The Personal Data (Privacy) Ordinance does not dictate where servers sit, but the PCPD's 2024 Model Personal Data Protection Framework for AI expects organisations to assess risk and protect personal data across the AI lifecycle. Our analysis of the PCPD’s agentic AI guidance covers the latest expectations.
Confidential inference strengthens the evidence you can present. Instead of explaining a vendor's policy, you can show an attestation record. That is particularly useful for banks under HKMA supervision, insurers, law firms and healthcare administrators, where cross-border processing questions arise in every audit. It complements, rather than replaces, a clear data residency position.
The third effect is strategic. Until now, the only way to guarantee nobody outside could read sensitive prompts was to run models on your own infrastructure. Confidential inference creates a middle path: cloud scale with private-deployment-like guarantees, at a price yet to be disclosed.
Which enterprise AI workloads actually need confidential inference?
Not every workload needs confidential inference. Use a four-tier sensitivity test: public content needs nothing extra; internal operational data needs standard enterprise terms; personal or client-confidential data needs strong contracts plus residency controls; and trade secrets, regulated records or privileged material justify confidential inference or private deployment.
The four-tier workload test--- Tier 1, public: marketing drafts, public research summaries. Standard enterprise AI terms are sufficient.
--- Tier 2, internal: SOPs, meeting notes, internal FAQs. Enterprise terms with no-training and retention limits.
--- Tier 3, personal or client-confidential: customer records, HR files, client correspondence. Add residency controls, access logging and a documented PDPO assessment.
--- Tier 4, crown jewels: M&A documents, unreleased financials, source code, privileged legal advice, medical records. Require confidential inference with attestation, or a private deployment you control.
Consider a Hong Kong logistics group with 400 staff. Its customer-service drafting sits in Tier 2, its payroll analytics in Tier 3, and its pricing models for a pending tender in Tier 4. Treating all three the same either overspends on the first or underprotects the last.
Most organisations find that Tier 4 represents a small share of AI use cases but a large share of risk. That is where confidential inference budgets belong.
What questions should you ask AI vendors about confidential computing?
Ask five questions: which hardware provides the trusted execution environment, whether attestation reports are available to you or only to the vendor, which models and endpoints are covered, what happens to logs and safety review data, and what the performance and price premium is. Vague answers indicate marketing rather than architecture.
The five-question vendor checklist--- Hardware: Which CPU and GPU platforms run the TEE, and are GPUs in confidential mode or only the CPU?
--- Attestation access: Can our security team verify attestation independently, or do we receive only a vendor statement?
--- Coverage: Which models, regions and API endpoints are protected? Does it cover chat products as well as the API?
--- Side channels: Where do logs, telemetry, safety reviews and metadata go, and who can read them?
--- Cost: What is the price premium and latency overhead for our workload, with a benchmark?
These questions matter because even OpenAI has not yet published the answers for Private Inference. As VentureBeat noted, which models, which hardware, what customers can independently verify and how much it costs all remained open at launch.
What mistakes do enterprises make when evaluating private AI inference?
The most common mistakes are treating a no-training clause as a privacy guarantee, ignoring metadata and logs, assuming private deployment is automatically more secure, applying the highest protection to every workload, and approving a vendor's roadmap promise as if it were a shipped capability. Each one creates cost or exposure without real protection.
--- Confusing training with access. A vendor can promise not to train on your data while staff or systems can still read it.
--- Forgetting metadata. Usage logs, file names and prompt lengths can reveal deal activity even when content is protected.
--- Equating private with secure. A self-hosted model on a poorly patched server can be less safe than a well-run attested cloud service. Private shifts responsibility; it does not remove risk.
--- Over-protecting everything. Applying Tier 4 controls to Tier 1 workloads slows adoption and inflates budgets.
--- Buying the preview. Private Inference was announced as a preview. Approve architectures that work with what is generally available today, and add confidential inference when it ships with verifiable documentation.
What should enterprise leaders do about confidential inference in the next 90 days?
In the next 90 days, classify your AI use cases into the four sensitivity tiers, add the five confidential-computing questions to every AI RFP, review retention schedules for customer-side storage, and identify one Tier 4 workload as a candidate for attested or private deployment. That positions you to act as vendor offerings mature.
Confidential inference will not replace contracts, governance or good security operations. It does change the question your board can ask, from "do we trust the vendor?" to "can the vendor prove it?"
For Hong Kong organisations, where first-party access to some frontier models remains limited and PDPO scrutiny is rising, that shift rewards leaders who prepare their classification and procurement frameworks before the technology becomes standard.
Technology should make your team more confident, not more anxious. We understand AI. We understand you. With UD by your side, AI never feels cold.
Reviewed by the UD enterprise AI team. UD has supported Hong Kong enterprises with cloud, security and AI infrastructure for 28 years.Map Your Sensitive AI Workloads
Now that you have the framework, the next step is identifying which of your AI workloads need private or attested processing. We'll walk you through every step, from workload classification and data residency planning to private deployment on infrastructure you control.