Skip to main content
All notes

AI Audit & Compliance

Governance is not the fifth layer

NVIDIA and Red Hat describe the AI factory as a five-layer cake, with governance at the top. In a regulated or sovereign environment, that is precisely where AI programmes go to die.

Nitish Hejmadi · Principal Architect · October 2, 2026 · 9 min read

There is a useful frame doing the rounds. On the NVIDIA AI Podcast, Red Hat CTO Chris Wright and NVIDIA’s Justin Boitano describe the enterprise AI factory as a five-layer cake: accelerated hardware, hybrid cloud infrastructure, models, agents, and production-grade governance. It is a good frame. It is the first vendor articulation I have seen that treats enterprise AI as an infrastructure problem rather than a model problem, and it is right about the thing most organisations get wrong — that you cannot pilot your way to production.

I want to take it seriously enough to disagree with one part of it.

First, why the stack exists at all

The economics are the reason this is urgent rather than interesting. Red Hat’s published position is that the unit cost of a token is falling roughly 75 to 90 percent a year, while enterprise token consumption is rising more than 500 percent a year. Their own projection has total consumption rising twenty-four-fold by 2030.

−75–90%

Annual fall in the unit cost of a token

+500%

Annual rise in enterprise token consumption

24×

Projected growth in consumption by 2030

Red Hat’s published figures. See provenance below.

Read those together and the conclusion is uncomfortable: falling prices will not save you. Consumption is outrunning deflation, which means efficiency stops being a procurement question and becomes an architectural one. That is the real argument for having a stack at all — not elegance, but the fact that an unstructured estate gets expensive faster than the market gets cheap.

The amendment

Governance is not a layer. It is a property of all five, and the organisations that treat it as the top of the stack are the ones whose pilots never ship.

I say this from a specific vantage point. My work spans defence and a Big Five Canadian bank. In both, the question that stops a promising AI system is almost never “does it work.” It is “can you evidence what it did, on what data, under whose authority — and can you show the control existed before the incident rather than after it.” That question cannot be answered at the top of the stack, because the evidence it requires is generated further down, or it is not generated at all.

Here is what governance looks like at each layer if you build it in rather than bolt it on.

Layer 01 — Accelerated hardware

Jurisdiction is decided in the rack, not in the policy

Where the machine physically sits determines which country’s law reaches the workload running on it. That is not a legal footnote; it is an architectural constraint that is expensive to reverse. For a federally regulated institution it drives third-party and concentration risk. For defence and government it sets the accreditation boundary — and that boundary starts at the metal, not at the application.

The hardware layer is also where confidentiality from the operator is won or lost. Extending trusted execution environments from the CPU to the accelerator is what lets you run a sensitive workload on infrastructure you do not exclusively control. If that is not designed in at layer one, no amount of policy at layer five recovers it.

The question to ask

Can you name, for every accelerator in your estate, the jurisdiction it sits in, who can physically access it, and whether the workload is protected from the platform operator?

Layer 02 — Hybrid cloud infrastructure

Shared accelerators mean a shared blast radius

The economic case for this layer is utilisation: queueing and sharing GPUs across teams so they are not idle. That is correct, and it quietly creates a multi-tenancy problem. Once teams share accelerators you have to answer whose workload can pre-empt whose, what one tenant can observe of another, and what happens to a regulated workload when an experimental one saturates the cluster.

This is also where configuration drift lives. Drift is usually filed as an operations annoyance. In a regulated estate it is a governance failure with a precise definition: the growing gap between the environment you had accredited and the one that is actually running. Every week that gap widens, your assurance evidence describes a system that no longer exists.

The question to ask

If an experiment and a regulated production workload compete for the same accelerator at 3am, which one wins — and is that a design decision or an accident?

Layer 03 — Models

A model nobody registered is a model nobody validated

Model sprawl is the governance problem of this layer, and it is structural rather than cultural. Capable open models are released weekly, and the distance between a team trying one and a team depending on one is very short. Shadow AI is not people being reckless. It is the predictable outcome when the sanctioned path is slower than the unsanctioned one.

This is why curated model catalogues exist — not as a convenience, but because provenance is a control. For Canadian financial institutions it stops being optional shortly: OSFI’s Guideline E-23 takes effect on 1 May 2027 and expects an enterprise-wide model inventory with lifecycle oversight proportionate to risk. An inventory is not a document you assemble the quarter before. It is a by-product of a serving layer that makes registration the path of least resistance.

The question to ask

Is your model inventory generated by the platform, or maintained by hand in a spreadsheet that is accurate on the day it is circulated?

Layer 04 — Agents

The layer where our identity model runs out

This is the least solved layer and the one I would watch most closely. An agent is a non-human actor that takes consequential action, and almost every enterprise identity model was built on the assumption that actions trace to a person. An agent needs three things we are not yet good at issuing: an identity of its own, an authorisation scope narrower than the human who launched it, and an execution trace detailed enough to reconstruct a decision after the fact.

The industry is converging on real answers — workload identity frameworks such as SPIFFE and SPIRE, kernel-isolated sandboxes, standardised tool interfaces. What has not caught up is the accountability question underneath. When an agent takes an action that turns out to be wrong, the organisation must be able to say who authorised it to act, within what limits, and on whose behalf. That is an architecture problem long before it is a policy problem.

The question to ask

Does every agent in your environment have its own identity and scope — or does it borrow the credentials of the person who started it?

Layer 05 — Production-grade governance

Where governance becomes visible, not where it happens

The top layer is real and it matters. It is the assurance surface: the inventory, the monitoring, the thresholds, the reporting a board or a regulator actually reads. My objection is not that it exists. It is the implication that it is where the work is done.

Everything the fifth layer presents was produced by the four beneath it. The inventory is only as true as the serving layer that populates it. The audit trail is only as complete as the agent runtime that emitted it. The residency attestation is only as good as the decision made in the rack. Build layers one to four without asking what evidence they must produce, and the top of the stack becomes a reporting exercise describing a system you cannot account for.

The question to ask

For your last AI governance artefact — was it generated by the platform, or written about the platform?

Why this lands differently in Canada right now

Two things are converging. Guideline E-23 arrives on 1 May 2027 and pulls model risk management out of the quantitative-validation corner and across the whole enterprise model estate, AI included. At the same time, sovereign AI capacity is being built here — HIVE’s BUZZ HPC has announced a 320-megawatt AI data centre in the Greater Toronto Area.

The constraint on that build-out will not be manufacturing, and it will not be power, though both are hard. It will be the buyers — banks, departments, agencies — who cannot put a workload onto new capacity until they can evidence how it is governed. The governance question is not a compliance tax applied at the end of the programme. It is a gating factor on the whole market, and it is decided at every layer of the stack.

Navies have a way of describing a procedure nobody has ever rehearsed: it is not a control, it is a document.

The same is true of an AI governance framework that sits on top of four layers never built to produce evidence. The fifth layer is not where governance happens. It is where governance becomes visible — and you cannot make visible what was never generated.

Provenance

Nitish Hejmadi is Principal Architect at Aegis Systems. He works on security and AI architecture for defence and federally regulated financial services.