Segment the Risk. Then Solve the Solvable.
An updated AI Risk Stack, three years on — and why the financial tile is the one we can already close
Three years ago, a few weeks after ChatGPT was made public, I drew a chart. It was a stack of colored tiles, each representing a category of risk: existential risk at the top, financial and information security in the middle, and disruption and inequality at the bottom. The argument was simple. "AI risk" is a dozen problems that happen to share one technology, and they will be managed badly until we stop debating them as a single thing and instead hand each tile to the people who actually know that domain.
I put two numbers in the margin. Ten to twenty percent for the top tile. Fifty to one hundred percent — "already started" — for everything underneath it.

Between September 9 and 14, 2026, the AI-safety debate got louder than at any point since I drew that slide. An alignment lead at one frontier lab publicly said that he puts the odds of AI-caused extinction within 10 years above 1 in 10. Its CEO published an essay on Saturday asking the whole industry to slow its pace of capability growth, and committed to giving outside evaluators permanent, badge-level access to his company's systems. By Sunday, the CEO of his largest rival had endorsed it, said that a 10% extinction risk would be unacceptable at any price, and shelved an IPO over it. The CEO of the company that sells them all their chips called the entire premise nonsense. A senator signaled a bill to ban superintelligence. China's state press called the slowdown proposal a Cold War maneuver.
That is a lot of emotion for one week. Almost none of it is about the tiles, where most of the actual damage is happening.
So this is a progress report on the stack. What held up? What did I get wrong? What can we now say with evidence instead of instinct? And a specific claim: the financial-security tile, AI-driven fraud at scale, is mostly solvable with tools that exist today, and how it got solved is a template the other tiles could borrow.
What the 2023 stack got right
The core thesis survived contact with reality. Segmentation is now the mainstream regulatory posture. The EU AI Act is built on risk tiers. Card networks are writing separate rules for agent-initiated versus human-initiated transactions. The forecasting community asks separate questions about "catastrophe" and "extinction." Nobody seriously argues anymore that a single governance regime can cover a deepfake, a mule network, and a misaligned superintelligence.
The "already started" arrow was also right and understated. In 2023, the bottom tiers were a projection. In 2026, they have attached audited numbers, which I will come to.
What it got wrong
Six things, roughly in order of how much they matter.
1. I put severity and probability on the same axis. Existential risk sits on top because it is the worst case, not because it is the most likely. Those are different questions, and conflating them is the single biggest reason the public debate is so hard to follow. The updated stack keeps the tiers but makes the second axis explicit: how realized each tile is as of today, on a scale from "speculative" to "happening, measured, audited."
2. I described the financial tile as a systems-takeover problem. The 2023 example was SWIFT or a national payments rail being manipulated. That is a real scenario, but it belongs on the "Systems at Risk" tile. What actually materialized in finance was something else: industrialized deception. Cloned voices to defeat call-center verification. Automated credential stuffing to take over accounts. Synthetic identities assembled from fragments of real ones. Mule networks recruited and operated at machine speed. The attack targets the humans and the onboarding systems positioned in front of the rail.
3. There was no tile for agents that move money. In 2023, there was no need. In 2026, Visa, Mastercard, and American Express all run frameworks for AI-agent-initiated payments, EMVCo has a task force on it, and the fundamental legal question — who is liable when an agent misreads an instruction and pays the wrong party — has no settled answer under US Regulation E or anywhere else. It belongs in the financial column, alongside fraud.
4. The bottom band mixed three time scales. "Disruption, dislocation, and inequality" bundled job displacement (slow, structural, hotly contested) with erosion of shared truth (fast, already here) with dependency on AI (systemic). These need to be separate tiles because they have separate owners. Economists own the first. Platforms, publishers, and election authorities own the second. Prudential regulators own the third.
5. "Dependency on AI" deserved a promotion. When every bank, insurer, and payment processor scores risk with the same three foundation models, a single model failure becomes a correlated market failure. This is a concentration risk in the classic financial sense, and it belongs in the "Systems at Risk" tier, not in the miscellany at the bottom.
6. The slide implied a governance layer but never drew it. The whole point of segmenting the risk was that each segment gets a set of domain experts. I never specified what those experts would actually build. Three years of doing this in one domain has given me a concrete answer, and it is the second half of this post.
The stack, updated: YE 2023 versus YE 2026
Here is the comparison as I now see it. The probabilities on the top tile are not mine; they are the range of published third-party estimates, and the spread itself is the finding.

Two observations fall out of the table.
First, the loudest tile has the widest error bars, and the other tiles are narrowing. The right response is to argue about it differently: as a set of testable intermediate claims rather than a single terrifying percentage.
Second, one tile has moved from "already started" to "already being closed." That is the financial one. Let me show you what closing it looks like.
The financial tile is mostly solvable. Here is the proof.
In late 2023, a UK business-banking platform — more than 50,000 business accounts, more than $10 billion in transactions processed — became a target for exactly the AI-driven vectors above: account takeover fed by deepfakes and automated credential stuffing, money muling, and synthetic-identity onboarding. The timing was not accidental. UK authorized push payment (APP) legislation had just made banks partially liable for scam losses, which meant every fraudulent real-time payment now hit the institution's own P&L.
Their fraud rate, measured as a basis point of transaction value, had run around 20 bps. Under attack, it spiked to 89 bps. On a billion dollars of volume, that is $8.9 million walking out the door.
We were engaged to address this issue in March 2024. A custom model and rule set went live within weeks. An ensemble model trained on improved outcome labels followed in the summer. By early 2025, the fraud rate was at 2 bps and has remained at that level since. That is a 98% reduction: for every billion dollars processed, losses were reduced to $200,000 instead of $8.9 million.

Executives typically view their fraud and compliance teams from a lens based on a false choice, that reducing fraud equals reducing sales. The only way to cut losses is by declining more transactions, which just shifts the cost burden onto legitimate customers. Few falsehoods are more damaging to fraud and growth alike, and in this deployment, false positives fell too, by more than 97%. So the fraud line and false positive line moved down together instead of trading off against each other. Fewer good customers were blocked (more revenue and less friction), and fewer fraudulent payments were approved, simultaneously. The platform was subsequently named Best Business Banking Provider at the British Bank Awards, which I mention only because "we stopped fraud and the customers noticed things got better" is the whole argument against the myth that fraud prevention and growth pull in opposite directions.
How, mechanically
It is a pipeline of five stages, and each one is ordinary engineering:
- Data orchestration. Device fingerprinting, behavioral monitoring, network graph, payment-rail telemetry, cross-entity and cross-institution signals, joined in a privacy-preserving way. The network intelligence element is the multiplier: a mule account or synthetic identity seen at one institution is knowable at the next before it transacts.
- Graph neural networks. Mule networks and synthetic-identity rings are relationships, not transactions. You cannot see them one row at a time.
- Anomaly detection. Catches the novel typology before anyone has labeled it — the thing standalone systems are structurally worst at.
- Supervised machine learning. Once outcomes are labeled, precision climbs. This is where the false-positive collapse comes from.
- Agentic triage, step-up authentication, case-building, and regulatory reporting. AI agents do the volume — sort alerts, request the extra factor, assemble the case file, draft the suspicious-activity report — with a human making the consequential call.
The shape repeats: secure and anonymize shared signals → detect relationships → detect the unknown → learn from labels → automate the routine, gate the irreversible. That shape is portable. It is also exactly the shape that was missing during the July agent-swarm incident, in which the activity ran for roughly a week before a human saw it. Financial services have spent 20 years reducing detection-to-human latency to hours because regulators have made losses land on institutions' own balance sheets. The lesson is not that fraud is special. It is that a tile gets addressed when someone owns the loss, measures it in the same unit every quarter, and runs a feedback loop.
The missing layer, and a template for the other tiles
The AI industry is organized in a four-layer stack: semiconductors and power, data centers and cloud, foundation models, and applications. Every serious safety proposal from September 9-14 is aimed at layer three — slow the models, audit the labs, embed the evaluators. While necessary, it is also insufficient because a perfectly aligned foundation model can still be prompted by a criminal to clone a voice, and a perfectly evaluated lab can still ship a model that a bank deploys with no kill switch.
What is missing is a layer 3.5: industry-specific, real-time trust infrastructure that sits between the general-purpose model and the regulated application. In healthcare, that layer looks like clinical-decision governance. In defense, it looks like targeting rules of engagement. In payments, it looks like the pipeline above — data plane, model and decisioning plane, governance and workflow plane — and its job is to convert "the model said so" into an outcome a regulator can examine, and a customer can appeal.

Who provides layer 3.5 is a legitimate debate, with three viable answers. An independent consortium platform (for example, the FraudNet model) is faster, more accurate, and spans institutions, but must earn trust as a private actor. A trade organization self-policing its members is credible and cheap, but structurally slow to sanction its own. A government regulator has the authority and the mandate, but rarely the telemetry or the engineering tempo. In practice, the healthy pattern is all three: the consortium builds and runs it, the trade body sets the standard, the regulator sets the floor, and audits the evidence.
Whichever provider you pick, the rules the layer enforces should be the same. Here is the template I would hand to a peer in any other high-risk domain — energy, health, elections, defense — written so it can be adopted without changing a word beyond the examples.
Ten canonical rules for a layer-3.5 governance layer
- Human-in-the-loop by consequence, not by system. Autonomy is granted per class of action, not per deployment. Any action that is irreversible, safety-critical, or applied at a mass scale routes through a human decision gate. Agents may sort, score, draft, and recommend without limit; they may not execute the irreversible step alone.
- A tested kill-switch with a named owner. Every deployed model and agent has one human who can halt it, and the halt is drilled quarterly. A kill switch that nobody has pulled is a hypothesis.
- Bounded detection-to-human latency. Define the maximum time an anomaly may persist before a human is looking at it, in hours, and measure against it. In payments, the norm is under 48 hours. A week is a failure by construction.
- Evidence over assertion. Every automated decision emits a machine-checkable record: inputs, model version, score, reason codes, and the identity of any human who overrode it. No completion claim is accepted from an agent without its artifact.
- Network intelligence over standalone judgment. Signals are shared across institutions in a privacy-preserving form so that a novel attack seen once is defended everywhere within days. Standalone systems learn each lesson separately and slowly.
- A closed feedback loop on labeled outcomes. Confirmed harm, confirmed false alarm, and appeal outcomes flow back to the model within a defined window. Where there is no label, there is no learning, and the detection lag stretches from weeks to years.
- Adversarial testing against the current toolset, before and after deployment. Red-team the system with today's attacker capabilities — cloned voices, synthetic documents, scripted credential stuffing — and back-test every model change on held-out data.
- Versioned, reversible, reviewable model lineage. Every change is recorded, can be rolled back, and can be explained to an examiner without the vendor present.
- Independent verification by a party that did not build it. Consortium, trade body, or regulator — someone other than the operator audits rules one through eight and publishes what they found.
- Report outcomes in shared units. Loss rate in basis points. False-positive rate. Detection-to-human latency in hours. Time-from-novel-typology-to-network-wide-defense in days. Same units, same cadence, so that "better" is a number and not an adjective.
Read those ten back against the July incident, and you can see which ones were absent: no bounded latency (rule 3), inter-agent communication with no artifact trail (rule 4), a benchmark objective with no human gate on internet access or third-party systems (rule 1). Read them against the financial case study, and you see all ten present. The gap between the two is governance engineering.
Where the regulators are, as of YE 2026
This section is deliberately a snapshot; it will date, and the rest of this post is meant to outlive it.
European Union. The AI Act's transparency obligations (Article 50 — telling people when they are dealing with an AI and labeling synthetic content) took effect on August 2, 2026, along with enforcement powers over general-purpose model providers. The heavier high-risk obligations — which explicitly include AI used for credit scoring and payment-behavior scoring — were deferred by the Digital Omnibus to December 2, 2027. A deferral is not a dismantling; the risk architecture remains in place, and sectoral law (GDPR, product liability, national conduct regulators) applies in the meantime.
United Kingdom. Mandatory reimbursement for victims of authorized push payment scams has been in force since October 2024. The regulator's independent first-year review, published in July 2026, found the scheme delivered net benefits and reduced fraud losses. This is the single clearest example of the mechanism I described above: put the loss on the institution's balance sheet, and the institution builds the layer. The FCA's Mills review on AI in retail financial services and HM Treasury's payments plan both flag agentic payments and authentication as live workstreams.
Card networks and agentic commerce. Visa, Mastercard, and American Express each run a live framework for agent-initiated transactions, with Mastercard introducing a verifiable-intent credential that changes fraud scoring and liability when present. EMVCo is assessing whether global specifications need to change. Under all of them, liability for outright fraud follows existing tokenized-transaction rules; liability for an agent that simply got it wrong is not yet settled by any network rule or statute.
United States. No federal AI statute. The FBI's 2025 crime report is the first to break out AI-enabled fraud as its own category — an important measurement milestone in its own right. A patchwork of state laws governs disclosure and automated decisions. Following the events of September 9-14, more than 20 lawmakers responded publicly, and a superintelligence ban bill was introduced in the Senate; nothing has passed.
Frontier Labs. Two of the three largest have now committed to permanent, employee-level access for independent evaluators, and to some form of coordinated pacing. This is layer-three governance, and it is welcome. It does nothing, by itself, for layer 3.5.
How to make the conversation productive again
The AI-safety debate has a structure, and once you see it, the noise drops. One side is arguing about the top tile using probabilities that cannot be validated. The other side is arguing about the bottom tile, using job market data that cuts both ways. Neither is talking about the middle, where the money and the measured harm actually are. Eight suggestions, offered to anyone who has to lead a discussion on this in a boardroom, a legislature, or a newsroom.
- Ask "which tile?" before anything else. Most AI-safety arguments are two people defending different tiles. Name the tile, and half the disagreement evaporates.
- Agree on definitions. To date, there is no consensus on the most basic terms, such asorga ‘Artificial General Intelligence’ (AGI). Without common definitions, we will have difficulty setting national and global standards, taxonomies, and milestones.
- Report in units, not adjectives. Basis points. Dollars. Hours to detection. Number of institutions defended. "Catastrophic" and "nonsense" are both zero-information words.
- Publish ranges with sources and treat the spread as the finding. 0.4% to 25% on extinction is a measurement of how little consensus exists, not a settled number. That's useful to know, and more honest than picking one end.
- Separate misuse from misalignment. A criminal using a model to clone a voice and a model pursuing a goal its operators did not intend are different problems with different owners, different timelines, and different fixes. Bundling them lets each side use the other's evidence.
- Study the tiles that are working. Fraud is being closed. Why? Loss of ownership, shared units, a feedback loop, network intelligence, and human gates on irreversible actions. Those are transferable. Import the mechanism, not just the alarm.
- Adopt incident reporting as a norm, not a scandal. Financial services files suspicious activity reports by the millions and treats a loss disclosure as routine. The July incident was disclosed, investigated by outside parties, and reconstructed publicly at a security conference within six weeks. That is the right behavior, and it should become boring.
- Share the governance burden at layer 3.5 across industries. Pacing the frontier is a layer-three intervention. Most realized harm occurs at the point of deployment, and the people who understand deployment in a given industry are those who work in it.
- Ask CEOs for their detection-to-human latency, not their p(doom). One of those numbers is a mood. The other is a control you can audit tomorrow.
Where this should end up
I drew that first stack for my board of directors out of my own uncertainty about how bad this could get, a few weeks after ChatGPT was released to the public. Today, most of the tiles have a name attached: an expert community, a unit of measure, a feedback loop, a governance layer that enforces the ten rules above, and an audit by someone who didn't build the system. Some tiles are far from that state — the top one perhaps permanently, because you cannot run a feedback loop on an event that can only happen once. But the middle of the board is closer than the headlines suggest, and the financial column is closest of all.
We stopped 98% of AI-driven fraud at one institution while making life better for its legitimate customers, using techniques that are three years old and getting cheaper. It required deciding who owned the loss, measuring it, sharing what we saw, and keeping a human on the irreversible step.
Every other tile on this board has the same assignment that financial security just finished: pick the owner, pick the unit, build the feedback loop, and put a name on the kill-switch.

You might be interested in…
Get Started Today
Experience how FraudNet can help you reduce fraud, stay compliant, and protect your business and bottom line
%20(640%20x%201229%20px).png)
