The Observation
I've sat in enough post-mortems now to recognize the pattern before anyone says it out loud. Someone opens the meeting with a version of the same sentence: “We don't understand how this got so out of hand. It worked so well when Sarah's team was using it.”
That sentence is the whole problem, compressed. The tool didn't fail. The team didn't do anything wrong. What failed was the assumption sitting quietly underneath the rollout. The assumption that something proven at ten people, on one workflow, with one accountable owner, would behave the same way at two hundred people, across a dozen workflows, with no single owner in sight.
It won't. It never does. And the gap between those two states has a name I've come to rely on in almost every AI governance conversation I have with clients now: Governance Debt.
Today’s issue is my attempt to give that gap a shape you can actually work with — not just a warning, but a diagnostic, a framework, and a step-by-step playbook for closing it before it closes in on you.
I've built it around real incidents (not hypotheticals dressed up to look real), a scoring model you can run on your own AI tools this week, and a rollout protocol I use with teams moving from pilot to department scale.
If you lead an AI initiative, sit on a governance committee, or are the executive who just asked “can we roll this out everywhere,” this is written for you.
The Pattern
The Story Every AI Leader Recognizes
Here's a composite scenario I've watched play out, in some variation, at more organizations than I can count.
A five-person content team adopts a generative writing assistant to speed up first drafts. One person champions it, tests it obsessively, learns exactly where it's reliable and where it isn't, and builds a set of informal rules. They never use it for client-facing numbers, always fact-check names and dates, never paste in anything under NDA. The team ships faster, quality holds, and everyone's happy.
Six months later, a VP sees the results and asks a reasonable question: “Why doesn't the whole department use this?” The tool gets a department-wide license. Two hundred people across content, comms, sales enablement, and customer support get access within a week.
Within a month, someone in customer support pastes a customer's account details into a prompt to write a better response. Someone in sales enablement uses it to draft claims about product capabilities that legal never reviewed. Someone in comms publishes a stat the tool fabricated with total confidence.
Three different teams build three different sets of ad hoc rules, none of which match. IT gets a security question they can't answer. Legal asks who approved this. Nobody can say.
Nothing about the tool changed between month one and month six, except for the exposure.
What Governance Debt Actually Is
I borrow the term from technical debt on purpose, because the mechanism is identical. Technical debt is the gap between the fast, working shortcut you shipped and the properly engineered version you'd build if you had time. It doesn't announce itself. It accrues silently, and you only feel it when you try to change something and discover how much has been quietly resting on the shortcut.
Working definition
Governance debt is the accumulated gap between the informal, team-level controls that let an AI tool work safely for a small group, and the formal, organization-level controls required for that same tool to work safely at scale. It is invisible at team scale precisely because team-scale usage never triggers the conditions that expose it — broad data exposure, inconsistent judgment calls, external liability, and reputational visibility.
Three things make governance debt different from ordinary risk management, and worth treating as its own category rather than folding into AI risk generally:
It's created by success, not failure. A tool has to work well enough at the team level to earn the right to scale. Which means the same story that gets you funded is the story that hides the debt.
It's invisible to the people best positioned to see it. The champion who built the informal rules doesn't write them down, because they don't experience them as rules. They experience them as judgment. Judgment doesn't scale by license count.
It's discovered by whoever gets hurt first, not by whoever is responsible for preventing harm. That's almost always the wrong order, and it's why the first sign of governance debt is usually a legal, security, or PR incident rather than an audit finding.
Why This Keeps Happening: Three Structural Reasons
1. Adoption is bottom-up; governance is top-down. They rarely arrive on the same timeline.
AI tools spread through organizations the way spreadsheets and Slack did before them — laterally, through word of mouth and visible results, well ahead of any formal review. By the time governance functions become aware a tool exists, it's already generating value someone doesn't want to give up, which changes the politics of every subsequent conversation from “should we approve this” to “how do we not kill something that's already working.”
2. Team-level success conditions are the opposite of department-level success conditions.
A small team succeeds with an AI tool because of high context and tight informal feedback loops. Everyone knows the failure modes, everyone can course-correct in real time, and one person can hold the whole risk picture in their head. A department can't run on any of that. It needs the risk picture written down, distributed, and enforced, which is exactly the infrastructure a successful pilot never had to build.
3. Incentives reward speed to scale, not speed to governance.
The person who champions a pilot is measured on adoption and results, not on how well-documented the guardrails are. Scaling faster looks like leadership. Slowing down to build a data handling policy looks like friction until the incident that makes everyone wish someone had insisted on it.
Real-World Cases and the Governance Debt Lesson in Each
These aren't abstractions. Each of the following surfaced publicly, and each maps cleanly onto one or more layers of the framework in the framework section below.
Case 1 — The data-exposure case: a generative AI assistant, used department-wide
In 2023, engineers in Samsung's semiconductor division started using a public generative AI chatbot to help debug code and summarize internal documents — useful, individually reasonable choices. At department scale, several employees pasted in proprietary source code and confidential meeting notes to get faster answers. The company subsequently restricted use of external generative AI tools on corporate devices while it built internal alternatives and clearer policy.
Governance debt layer: Data
A tool that was safe on a single laptop with careful judgment became a data exfiltration risk the moment two hundred people had access and no shared rule existed about what could and couldn't be pasted in. There was no data classification layer between the tool and the user — just individual discretion, which does not scale.
Case 2 — The accountability case: AI-drafted legal filings
In a widely reported 2023 U.S. federal court matter, attorneys submitted a brief containing case citations generated by a chatbot that turned out not to exist. The filing had passed through the firm's normal drafting process because no step in that process was built to catch fabricated citations from a tool nobody had used at that scale before. The court sanctioned the attorneys involved.
Governance debt layer: Accountability & Process
The informal check that would have caught this in a two-person team (“I personally verify anything the tool gives me before it goes out”) had no equivalent once the tool moved into a broader workflow with more filings, more drafters, and less individual ownership of any single output.
Case 3 — The liability case: a customer-facing chatbot
In 2024, a Canadian airline's website chatbot gave a customer inaccurate information about bereavement fare policy. When the customer relied on it, and the airline later tried to disclaim responsibility for its own chatbot's statements, a tribunal rejected that argument and held the airline accountable for what its AI tool had told the customer.
Governance debt layer: Accountability
The organization treated the chatbot as a support convenience rather than as an agent making binding representations on the company's behalf. Nobody had defined who owned the accuracy of what it said, which meant no one had built a process to keep it current or to review its answers against actual policy.
Case 4 — The reversal case: AI customer service at scale
A well-known fintech company automated a large share of its customer service interactions with an AI system in 2024, publicly citing the equivalent workload of hundreds of human agents. In 2025, the company walked part of that back, reintroducing human agents after acknowledging that service quality had suffered at the volume and complexity the automation was actually handling.
Governance debt layer: Process
Success at a contained pilot scope (simple, high-frequency queries) does not predict success once volume and complexity both rise — and no organization can tell in advance without a process for measuring degradation as scope expands, not just measuring cost saved.
Notice what these four cases have in common. None of them are stories about the AI not working. Every one of these tools functioned exactly as designed. What was missing, in every case, was the organizational infrastructure to keep working as designed once the number of people, the sensitivity of the data, or the stakes of the output moved beyond what the original pilot ever had to withstand.
The Framework
The Governance Debt Curve
The reason governance debt is so consistently underestimated is that it is genuinely invisible for most of the scaling journey, and then it isn't. Perceived value rises quickly and then plateaus as a tool proves itself.
Governance debt rises slowly at first, then compounds sharply the moment a tool crosses from a single team into a cross-functional group, because that's the exact point where the informal controls that worked (one champion's judgment, one team's shared context) stop covering the actual population using the tool.

Figure 1. Perceived value and governance debt accumulate on very different curves — and cross well before most rollouts notice.
The practical implication: The moment to build formal governance is not when a problem appears. It's at the inflection point on this curve (the transition from single team to cross-functional group), which is usually well before anyone feels enough pain to justify the effort. That mismatch, between when governance is needed and when it feels needed, is the core timing problem this publication exists to solve.
Same Tool, Different Blast Radius
It helps to be concrete about exactly what changes between team and department scale, because more people understate it. Six dimensions typically shift at once, and each one independently raises the cost of the same failure mode:

The A-D-P-A Framework: Four Layers of Governance Debt
Every governance debt incident I've examined (including the four above) traces back to a gap in one or more of four layers. I use this breakdown because it turns we need better AI governance, which is too vague to act on, into four specific, assignable questions.
Access — Who can use it, and how?
At team scale, access is implicitly self-limiting. Only the people who need the tool ask for it. At department scale, access is a deliberate control decision (role-based permissions, device restrictions, authentication requirements) or it is a decision made by default, which is to say, made badly.
Data — What it can see, store, and retain?
The question is not “is this tool secure” in the abstract. It's “what categories of data are people, in practice, going to put into this tool once two hundred of them have access,” and whether there is a classification and handling rule that answers that question before the paste happens, not after.
Process — How work gets reviewed, escalated, and approved?
A team's informal review (“I check anything before it goes to the client”) is a process, even if no one ever called it one. Scaling requires making that review explicit, assigning it to a role rather than a person, and defining what happens when volume outpaces the reviewer's capacity — which is exactly where the fintech case above broke down.
Accountability — Who owns outcomes when the tool is wrong?
This is the layer most often skipped, because at team scale it doesn't need naming; the champion owns it by default. At department scale, if no one has explicitly been assigned ownership of the tool's outputs, the honest answer to “who's accountable” is NO ONE, right up until a court, a regulator, or a customer decides it's the organization.
Why Traditional IT Governance Doesn't Close This Gap on Its Own
It's worth being direct about this, because it's the objection I hear most from teams who've already got governance. A standard IT security review, procurement checklist, or annual risk assessment is built to catch a different kind of problem — static systems with stable behavior, evaluated once before purchase.
AI tools are neither static (their behavior shifts with prompts, models, and usage patterns) nor stable in their risk profile (the risk changes with who is using them and for what, not just what the tool technically does).
While traditional governance asks “is this tool approved?” Governance debt requires asking “is this tool still governed for how it's actually being used today?” A question with no natural expiration date, which is exactly why it needs a recurring process rather than a one-time gate.
The Diagnostic
Before you can close a governance debt gap, you need to know how large it is. This is the assessment I run with clients on any AI tool being considered for scale-up, or just as often, on a tool that's already sprawling and needs a retroactive audit. It takes about 20 minutes per tool and requires input from whoever actually uses it day-to-day, not just whoever approved it.
The Governance Debt Self-Assessment
Score each statement 0–2:
0 = Not True/Doesn't Exist.
1 = Partially True/Informal.
2 = Fully True and Documented.
Answer honestly for how the tool is actually used today, not how it's supposed to be used.
Access (max 8)
☐ We know the complete, current list of everyone with access to this tool.
☐ Access is granted by role or need, not by whoever asks.
☐ We can revoke access for one person without disrupting the whole team.
☐ Access to sensitive functions (e.g., customer data, financial data) is restricted beyond general access.
Data (max 8)
☐ We have a written rule for what categories of data may and may not be entered into this tool.
☐ Users have been trained on that rule, not just told it exists.
☐ We know what the vendor does with data submitted to the tool (training, retention, third-party sharing).
☐ We have a way to detect or audit when the data rule is broken.
Process (max 8)
☐ There is a defined review step between AI output and anything that reaches a customer, regulator, or the public.
☐ The review step is assigned to a role, not dependent on one specific person being available.
☐ We have tested what happens when volume exceeds the reviewer's capacity.
☐ We track quality or error rates over time, not just adoption or usage volume.
Accountability (max 8)
☐ A named owner is accountable for this tool's outputs — not the vendor, not “the team.”
☐ That owner has the authority to pause or restrict use of the tool if something goes wrong.
☐ Legal, compliance, or risk functions know this tool exists and what it's used for.
☐ There is a documented escalation path for when the tool produces a harmful or incorrect output.
Reading Your Score

Run this per tool, not once for AI as a category. A department typically has three to eight AI tools in active use at any time, often with wildly different scores.
I've seen organizations score a 28 on their sanctioned enterprise tool and a 6 on the free version of a consumer chatbot three teams are quietly using instead because it's faster. The second number is usually the one that matters more.

Figure 4. A scorecard like this, run per tool, makes the weakest layer visible at a glance — in this illustrative example, Data and Accountability are the gaps to close first.
The Playbook
A diagnostic tells you where you stand. This section is what you actually do about it; whether you're about to take a team pilot department-wide, or you're staring at a tool that's already sprawled and needs to be brought under control after the fact.
The Graduation Protocol: Taking a Tool from Team to Department
Six steps, run in order, before any tool moves from a contained pilot to a broader rollout. I've deliberately kept this lightweight enough to run in one to two weeks. The goal is to close governance debt before scale, not to build a bureaucracy that makes teams route around it.
Step 1 — Inventory before you scale
Document exactly how the pilot team actually uses the tool today: which workflows, what data goes in, what informal rules the champion has been enforcing in their head. Write these down explicitly. This is the moment the invisible judgment calls become visible requirements.
Step 2 — Classify the risk tier
Not every AI tool needs the same level of governance. Classify based on two questions: how sensitive is the data it touches, and how directly do its outputs reach customers, regulators, or binding decisions?

Step 3 — Assign an accountable owner
Before scale, not after. This person is not necessarily the technical admin. They are the individual who can answer “who do we call when this goes wrong” without hesitation, and who has the standing to pause the tool if it does.
Step 4 — Rebuild the governance layer for scale
Using the A-D-P-A framework above, convert every informal team-level judgment call identified in Step 1 into an explicit, written control appropriate to the risk tier from Step 2. This is the actual debt paydown. Everything else in this protocol supports getting this step right.
Step 5 — Pilot the governed version
Before full rollout, run the newly governed version with a slightly larger group than the original pilot. Enough to stress-test the process step and the data rule against real, less-curated usage, but small enough that a failure is still contained and correctable.
Step 6 — Institutionalize
Fold the tool into standard onboarding, publish the usage policy somewhere people will actually find it, and set a review cadence so governance keeps pace with usage as it evolves rather than becoming another one-time gate that quietly goes stale.
RACI for AI Tool Scaling

R = Responsible, A = Accountable, C = Consulted, I = Informed.
A 30-60-90 Rollout Template
Days 1–30 — Foundation
Complete the Governance Debt Self-Assessment on the pilot tool.
Run Steps 1–3 of the Graduation Protocol.
Draft the data handling rule and get it reviewed by legal/compliance for Tier 2–3 tools.
Days 31–60 — Controlled expansion
Run Step 4 (rebuild governance layer) and Step 5 (governed pilot) with a small expanded group.
Track error rate, escalations, and any data-rule violations weekly, not monthly.
Adjust the process step based on what actually breaks under slightly higher volume.
Days 61–90 — Full rollout and institutionalization
Complete Step 6: onboarding, published policy, review cadence set.
Re-run the self-assessment. Confirm the score has moved into the Scale-ready band before declaring the rollout complete.
Brief legal, security, and the executive sponsor on the final governance posture, not just the adoption numbers.
The Governance Debt Paydown Plan For Tools Already Sprawling
If a tool is already in wide use without having gone through anything like the protocol above (which, if you're reading this, is probably why you're reading this), don't try to build perfect governance retroactively before doing anything.
Triage first:
Run the self-assessment today, as-is. Resist the urge to fix anything before you've scored the current state. You need the honest baseline.
If the score lands in Critical or High debt, immediately restrict the highest-risk usage (typically: block entry of customer or regulated data) while the fuller fix is built. This is a scope reduction, not a shutdown. Keep the tool running for lower-risk use.
Identify the single weakest A-D-P-A layer and fix that one first. Don't try to rebuild all four simultaneously. In every incident I've reviewed, one layer was doing almost all of the damage.
Notify legal/compliance proactively, before an incident forces the conversation. This single step does more to protect the organization, and the team that built the tool's success in the first place, than almost anything else on this list.
Run the full Graduation Protocol retroactively, treating the current usage as the pilot you're now formally scaling.
Governing for scale, Going Forward
Closing today's governance debt is necessary but not sufficient. The same pattern will repeat with the next tool a team adopts, unless the organization changes how it handles the moment between pilot and scale. This section is about building that muscle once, rather than re-running crisis response every time.
A Lightweight AI Tooling Intake Process
The goal here is not to gate every experiment. That just pushes adoption further underground, which is worse for visibility, not better. The goal is a low-friction way for governance to find out a tool exists while it's still small, so the Graduation Protocol can run before scale rather than after an incident.
A single-question intake form: What AI tool are you piloting, and what does it touch? It takes under two minutes, routes automatically to the accountable owner function.
A standing quarterly scan of expense reports, SaaS spend, and IT access logs for AI tool signatures. This is how most shadow AI usage actually gets discovered, and it should be routine, not reactive.
A default answer of “Yes, and here's the risk tier” rather than “No.” The intake process should make it easier to do the right thing than to go around it.
Two-Speed Governance
Not every stage of a tool's life needs the same rigor, and pretending otherwise is exactly what pushes teams to work around governance rather than with it.
I use a two-speed model:

The point of naming these as two explicit, legitimate speeds rather than one slow process everyone quietly bypasses is that it gives teams a fast, sanctioned path for exactly the kind of small-scale experimentation that produces genuine innovation, while keeping the crossing-the-threshold moment as the one place real rigor kicks in.
Metrics to Track Ongoing Governance Health

I opened this with a sentence I've heard in a dozen variations: “It worked so well when it was just one team.”
I want to leave you with the sentence I'd like to hear instead, in the same rooms, a year from now: “We scaled it, and we knew exactly what we were building before we built it.”
That's not a higher bar than most organizations can clear. It's a different sequence. Governance built alongside adoption instead of after it, at the one predictable moment (the crossing from team to department) where the informal controls that made the pilot succeed stop being enough on their own.
Everything in today’s issue exists to help you catch that moment on purpose, rather than finding out about it in a post-mortem.
