What accumulated, and how
No board approves a single-supplier strategy for its model layer. The concentration arrives one application at a time. The first composed application calls one provider's software development kit because that is the kit that was open on the day. The second copies the first, because copying the first is what composition is for. By the fourth, the pattern has become a house style, and the house style has become a dependency.
This is the same sentence as a failure mode already named in ordinary governance: the moment a composed application holds data that exists nowhere else, it has quietly become a system of record. Nobody set out to build it; it accumulated. Model concentration has the same shape, one layer down — and it is harder to see, because there is nothing in the architecture diagram to point at. The dependency is a line in a configuration file.
The cheapest thing you built is holding the deepest dependency.
Two dependencies, two urgencies
These are routinely conflated, and conflating them produces the wrong budget. One is a productivity question with a tolerable failure mode. The other is a business-continuity question with a silent one.
| Human-facing — chat, projects, assistants | Programmatic — composed applications, agentic workflows |
|---|---|
| Failure modeA person waits, or switches tool for an hour | A pipeline stalls or errors quietly; nobody notices until a deliverable is late |
| What is exposedIndividual productivity | Business processes, and the commitments made on top of them |
| DetectionImmediate — the person tells you | Whenever somebody happens to look |
| The remedyNone required. Keep the tool | Stop calling a provider directly from anything the business depends on |
Human-facing — chat, projects, assistants
- Failure mode
- A person waits, or switches tool for an hour
- What is exposed
- Individual productivity
- Detection
- Immediate — the person tells you
- The remedy
- None required. Keep the tool
Programmatic — composed applications, agentic workflows
- Failure mode
- A pipeline stalls or errors quietly; nobody notices until a deliverable is late
- What is exposed
- Business processes, and the commitments made on top of them
- Detection
- Whenever somebody happens to look
- The remedy
- Stop calling a provider directly from anything the business depends on
The remedy is not to leave the provider. The remedy is narrower than that, and cheaper.
NUVAI'S OWNThe human-facing side is a genuine asset and the model in use may simply be the best one for the work. Nothing here argues for moving it. What survives: the tool, the habit, and the productivity that came with both. The remedy is narrower and cheaper than switching providers — stop calling a provider directly from anything the business depends on.
Why no uptime figure appears in this piece
It would be easy to open with a status-page number. Those numbers move weekly, every one of them is contestable, and a reader who disputes the figure will dispute the argument attached to it. The argument does not need one. It requires only that you agree you never negotiated an availability commitment for the thing your workflow now depends on. That takes thirty seconds to check on a pricing page, and the answer is nearly always the same.
What is worth looking at is the shape of the arithmetic rather than any vendor's position on it. Enterprise contracts routinely assume three nines. Consumer and developer tiers of frontier model services generally commit to nothing at all.
Half a percentage point of availability is a working day per quarter
Arithmetic on a 91.25-day quarter (2,190 hours). No provider's observed availability is asserted here — the point is the shape of the curve, not anyone's position on it.
NUVAI'S OWNThe test you already own
Any sourcing doctrine worth the name already carries the question that settles this: could we change supplier within six months without losing the capability? If the answer is no, the dependency is too deep to be called a partnership. Most companies apply that test to their integrator, their platform vendor and their agency. Almost nobody applies it to the model layer, because the model layer never looked like a supplier. It looked like a feature.
Apply the second rule and the answer follows: buy what is commodity, build what differentiates, compose the long tail. A router that sits between your applications and the model providers is commodity. Nobody will value your company for having written one, competent products exist in the category, and the engineering required to maintain one is engineering not spent on the thing you are actually paid for.
What an LLM gateway actually does
It sits between your applications and the model providers and exposes one interface. Behind it, you route by cost, latency, task type or provider health; you fail over automatically when a provider degrades; and you add or swap a provider by changing configuration rather than code.
One interface in front of many providers, changed by configuration
The programmatic path is routed; the human path is left alone. Governance sits at one point instead of in every application.
NUVAI'S OWNFour classes, each with its named cost:
- Hosted marketplace. One API in front of many providers, run as a service by somebody else. The fastest start and the thinnest governance — your traffic leaves your boundary, and a percentage sits on spend.
- Open-source, self-hosted. The router as a component you run. Full control of residency and configuration — and engineering ownership, permanently, on something that is not your product.
- Managed, governance-first. The router as a product, with roles, guardrails and audit built in. For teams that need the audit trail more than the flexibility — at the price of a subscription and one more vendor relationship.
- Hyperscaler-native. Model access and routing inside a cloud tenancy you already contracted. For estates already consolidated on one cloud — and the depth of the dependency moves from the model vendor to the cloud vendor.
The choice turns on one question, and it is not a technical one: where does the traffic have to sit, and who has to be able to audit it? Answer that and the class picks itself. Answer it after the pilot and you will run the pilot twice.
The cheapest first step is not the gateway
Before any selection, there is a step available this quarter at close to no effort. Most frontier models are also served through at least one hyperscaler's model service, on separate infrastructure from the provider's own interface. Same model, same weights, different failure domain, different contract — and usually inside a cloud agreement you have already signed, with the residency and audit terms you already negotiated.
Configuring that as a secondary path is a configuration change and a credential. It survives a provider-side interruption without altering your prompts, your evaluations or your output quality.
A second path to the same model is not architecture. It is a week of insurance, and it is available now.
— What survives: no cost routing, no governance, and no relief from concentration in the model itself — only in the infrastructure serving it. It buys time. Spend the time on the gateway.
The governance answer costs a column
The objection arrives immediately from anyone who has watched a governance programme consume a year: this had better not be a new committee. It is not. Every part of the answer extends something that already exists in a company running composed applications properly.
| OWNER | DATA TOUCHED | CLASSIFICATION | MODEL PROVIDER(S) | FALLBACK TESTED |
|---|---|---|---|---|
| Ops lead | Order data (read) | Limited risk | Provider A, Provider B | 2026-07-14 |
Two columns. The AI Act inventory already existed; this is the same table.
NUVAI'S OWNTwo fields in the registry. It already carries an owner, the data touched and a classification per composed application, and it already doubles as the AI Act inventory. Add the model provider or providers used, and whether a fallback is configured and when it was last tested. That is a column, not a process.
One gate. Extend the graduation rule: anything composed that becomes business-critical needs a tested fallback path before it graduates, alongside the security and architecture sign-off. Not a documented one. A tested one — the difference is the whole point.
One number. Give the steering group the share of business-critical composed workflows with a tested fallback, reported quarterly against a dated target.
Two columns, one gate, one number. No parallel programme, no policy document nobody reads, and — the part that matters commercially — an inventory that answers a diligence question before it is asked.
Why model concentration is a P&L item, not an IT one
Supplier concentration has been a diligence question for as long as there has been diligence. A buyer asks what happens if the largest supplier raises prices, changes terms, or stops for a day. Nothing about that question is new. What is new is that the model layer is now a supplier — of an input that sits inside the product rather than beside it — and in most companies it has not yet been added to the list.
Two consequences follow. The first is ordinary: an unpriced supplier dependency eventually gets priced by somebody, and that somebody is usually the buyer, in the direction you would expect. The second is sharper. A reviewer reading an architecture diagram is looking for one pattern above all others — a product that is a thin interface over one public model service, with no fallback, no proprietary workflow and no owned data. Your company may be nothing of the kind. The intellectual property may be owned outright and the data genuinely proprietary. The composition layer can still read that way, because it is the newest, least documented and most visible part of the estate.
You will price this dependency, or your buyer will price it for you.
Turning that reading around costs a gateway subscription and two columns in a register. Not turning it around costs whatever a reviewer decides it costs, at the one moment in the company's life when you have no standing to argue.
The honest limits
- A gateway adds latency and one more component to operate. Real cost — and the same cost already accepted for every other integration layer in the estate.
- Multi-provider routing does not equalise model quality. One model may simply be better at your task. The gateway's job is optionality under stress, not diversification for its own sake.
- It does not fix ownership. If the capability sits in one person's head, or the intellectual property in a partner's repository, that closes on a hiring plan and a contract — not a software purchase.
- It does not make a weak application worth keeping. Routing a badly chosen workflow to three providers produces a badly chosen workflow with better uptime.
And yes, that includes us
The argument applies to this practice before it applies to anyone else. NUVAI stays model-agnostic for exactly the reason set out above: the durable assets are the client relationship and the domain model, not a badge from a platform. A practice that brands itself around one provider has made the same mistake it is being paid to find.
If knowledge work with an invoice attached is contestable, ours is too. It should be. The difference is what you hold when it ends: not a retainer that renews on inertia, but a working system, its registry entry — now with two more columns in it — and a team that runs it without us.
We are the last invoice of this kind you buy for that workflow.