Every enterprise AI roadmap in Hong Kong now faces the same tension. The best results come from committing deeply to one frontier model, yet the events of the past quarter show that any single model can be paused, withdrawn or restricted with almost no notice.
On 28 September 2026, OpenAI confirmed it would not release GPT-6.1 Astra, a model expected in ChatGPT and Codex in October, after internal safety tests. Earlier in the year, a US export-control order forced one leading lab to switch off its two most capable models for 19 days. Neither event was a vendor failure in the usual sense. Both were reminders that the model layer sits outside your control.
This guide defines AI model concentration risk, shows how exposed a typical Hong Kong enterprise is, and sets out a five-step framework your team can take to the next risk committee.
What is AI model concentration risk?
AI model concentration risk is the business exposure created when critical workflows depend on a single AI model or provider. If that model is withdrawn, repriced, restricted by regulators or simply goes offline, the workflows stop with it. It is the AI equivalent of relying on one bank, one cloud region or one supplier for a core process.
The risk has three sources, and most organisations only plan for the first.
--- Availability risk: outages, rate limits and capacity shortages during demand surges.
--- Continuity risk: a model version is deprecated, a planned upgrade is cancelled, or a release is paused for safety review.
--- Jurisdiction risk: a government order, export control or sanctions change removes access regardless of your contract.
Traditional vendor management covers the first reasonably well through service levels. The second and third are newer, and service credits do not compensate for a stopped business process.
Why did concentration risk become a board issue in 2026?
Concentration risk became a board issue in 2026 because three kinds of disruption happened in the same year: record outage frequency, a government-ordered model shutdown, and a frontier lab cancelling a finished model on safety grounds. Together they moved the risk from theory into incident reports that directors now read.
The evidence is specific. According to Ookla data cited in Kai Waehner's August 2026 analysis of multi-model strategy, high-signal disruption days across ChatGPT, Claude, Gemini and Copilot rose from 6 in the first quarter of 2025 to 51 in the first quarter of 2026.
The same analysis records that in June 2026 US export controls forced Anthropic to disable its two most capable models for all customers, with access restored 19 days later under new restrictions. Customers had no seat at the table where the decision was made.
Then came September. Implicator reported that GPT-6.1 Astra fell short of OpenAI's standards on staying within scope, seeking authorisation and accurately reporting its own actions, and that OpenAI had paused training of its most capable models. Any enterprise that had built its October plans around that upgrade now has to re-plan.
How exposed is a typical Hong Kong enterprise?
Most Hong Kong enterprises are more exposed than they realise, because model dependence accumulates quietly. Each team adopts the tool that works, integrations hard-code one provider, and nobody owns the full inventory. A mid-sized firm can easily have a dozen workflows that would stop if one provider became unavailable tomorrow.
Three local factors make the exposure sharper.
--- Access routes: many Hong Kong organisations reach frontier models through cloud marketplaces, resellers or embedded SaaS features rather than direct contracts, which adds a second party whose terms can change.
--- Regulatory expectations: the Privacy Commissioner's AI guidance expects organisations to assess and govern AI suppliers, and banks already manage outsourcing and concentration risk under HKMA supervision.
--- Language needs: Cantonese and Traditional Chinese performance varies widely between models, so the "backup" model is not automatically fit for purpose.
A useful test is simple. Ask each department head which AI workflows would stop if your primary provider disappeared for three weeks. If nobody can answer, the inventory does not exist yet.
What does a multi-model strategy actually involve?
A multi-model strategy means running critical AI workloads through an orchestration layer that can route each task to more than one qualified model, fail over when a provider goes dark, and fall back to a non-AI process when necessary. It is resilience engineering, not a procurement preference for owning several subscriptions.
The framework we recommend has five steps. Each one produces an artefact your risk committee can review.
1. Map: build the dependency inventory
List every production AI workflow, the model and provider behind it, the data it touches and the business process it feeds. Rank each by the cost of a three-week outage.
2. Abstract: take model names out of the code
Applications should call a role such as "contract summariser", not a named model. A gateway or orchestration layer decides which model fills the role. This turns a model disappearance from an emergency into a configuration change.
3. Qualify: prove a second model for every critical role
For each high-ranked workflow, test a second model from a different provider against the same evaluation set, including Cantonese and Traditional Chinese samples. A backup that has never been tested is a hope, not a control.
4. Degrade: design the no-AI fallback
For the few processes that cannot stop, define what happens when every model is unavailable: a rules-based path, a human queue or a delayed batch. Budget it as disaster recovery.
5. Contract: write exit and notice terms
Negotiate deprecation notice periods, data export rights and clarity on where prompts and outputs are processed. These terms cost little to request before signing and a great deal to obtain afterwards.
How does a multi-model approach work in practice?
In practice, a multi-model approach looks different by industry but follows the same pattern: a strong default model, a tested alternative for critical roles, and clear rules on which data may go where. The examples below show how a bank, a logistics group and a professional services firm apply it.
A regional bank. Its customer-complaint triage runs on a frontier model for nuanced Cantonese messages. The bank qualifies a second provider's model for the same role and keeps a rules-based routing fallback, so complaints are still categorised within the regulatory handling window if both models fail.
A logistics group. Document extraction from bills of lading is repetitive, so it runs on a small, inexpensive model with a second small model as backup. The frontier model is reserved for exception handling. Outages now slow the exceptions queue rather than stopping shipments.
A professional services firm. Drafting assistance uses whichever frontier model performs best that quarter, behind a gateway. Client-confidential matters route only to deployments the firm has vetted for data handling, so switching providers never requires re-clearing every engagement.
The routing decision also affects cost. The same analysis cites a Cursor experiment in which the same verified coding outcome cost between US$1,339 and US$10,565 depending solely on which models were assigned to which tasks.
What does multi-model cost, and when is it not worth it?
Multi-model adds cost in engineering, evaluation and governance, typically concentrated in the orchestration layer and in testing a second model for each critical workflow. It is not worth applying to every use case. Low-stakes, easily paused tasks can stay single-model; the discipline belongs on workflows where a three-week outage would hurt customers or revenue.
The trade-offs are real and should be stated plainly to the board.
--- Complexity: every extra provider adds an API, a billing model and a governance surface.
--- Prompt drift: a prompt tuned for one model is rarely optimal for another, so evaluation sets must be maintained.
--- New lock-in: the orchestration layer itself becomes a dependency, so choose it as carefully as a model vendor.
Compounding also argues for less model involvement, not more. An agent that gets each step right 99% of the time completes a 100-step workflow only about 37% of the time. As Enterprise Context Management's September review notes, promoting repeatable decisions to ordinary code removes both cost and dependency. Pair this with the spending controls in our AI FinOps guide.
What mistakes do enterprises make when reducing AI vendor dependence?
The most common mistakes are buying several subscriptions without an orchestration layer, naming an untested backup model, ignoring Chinese-language quality in fallback testing, and treating the problem as purely technical. Each leaves the organisation believing it is protected while the actual switching path has never been exercised.
--- Accounts without architecture: holding contracts with three providers achieves nothing if every application still calls one of them directly.
--- The paper backup: a second model listed in a policy but never run against real workloads.
--- Forgetting agents: autonomous agents inherit provider dependence and permissions, so they need the identity and access controls covered in our AI agent identity guide.
--- No switching drill: resilience is proven by rehearsal, the same way a data-centre failover is.
--- Owning nothing: if prompts, evaluation sets and business context live only inside one vendor's product, switching means starting again.
How should you present AI concentration risk to the board?
Present AI concentration risk to the board as an operational resilience item with three numbers: how many critical workflows depend on a single provider, how many have a tested alternative, and how long switching would take. Directors understand this framing because it mirrors how they already review cloud, banking and supplier concentration.
A one-page board brief should cover:
--- Exposure: the ranked dependency inventory from step one.
--- Coverage: the percentage of critical roles with a qualified second model.
--- Recovery time: the tested time to switch, not the estimated one.
--- Residual risk: the workflows deliberately left single-model, and why.
--- Next quarter: which workflows move into coverage and what it costs.
This turns a vague worry about AI vendors into a measurable programme with a trend line, which is what boards want to see.
Conclusion: own the context, rent the model
Frontier models will keep improving, and they will keep being paused, repriced and restricted. The enterprises that benefit most will be those that treat the model as a replaceable component and invest in what they own: the workflow design, the evaluation sets, the business context and the fallback plans. Start with the inventory this month, because every other step depends on it.
We understand AI. We understand you. With UD by your side, AI never feels cold.
Reviewed by the UD enterprise AI team. Events and figures were checked on 30 September 2026 against the dated sources linked above.
Find Out How Exposed Your AI Stack Is
Now that you have the framework, the next step is knowing where your organisation stands. Start with a free AI readiness check, and we'll walk you through every step, from mapping model dependencies to qualifying alternatives, designing fallbacks and reporting coverage to your board.