Line chart on a dark slate ground, headed "The Token Tax", reading "AI is not the problem. The wrong architecture creates the tax." A flat blue line labelled tokens runs along the bottom, marked visible. It passes through a green gate labelled architecture, after which three curves rise away from it, labelled governance, operations and trust. A footer line reads: unbounded authority creates costs the architecture cannot contain.

The Token Tax

The Hidden Cost of Giving AI the Wrong Responsibility

AUG 24, 2026 · GEORGE WILLIAMS


Recent posts

Chart comparing two lines over time. A rising amber curve labelled task speed climbs steeply and then flattens, while a flat red line labelled the close stays level throughout. Task speed improved; the length of the month-end close did not.

AI Made Every Part of the Month-End Close Faster. The Close Didn't.

The first wave of AI made every task faster. The next will have to own the process, not just help with it.

AUG 12, 2026 · GEORGE WILLIAMS

Title card labelled "probabilistic versus deterministic AI", reading "Two Kinds of AI. Only One Belongs on Your Ledger." Below it: the AI in the headlines is the wrong tool for the close.

Two Kinds of AI. Only One Belongs on Your Ledger.

Is it safe to automate your close with AI? Only if it follows rules instead of guessing.

JUL 30, 2026 · GEORGE WILLIAMS

Line chart headed "More revenue. Not more headcount." From today to tomorrow, a revenue line climbs steeply while a headcount line stays nearly flat. Subtitle: AI is increasing what each team can deliver.

AI Is Breaking the Channel's Oldest Rule

AI is severing the link between revenue and headcount. The partners who rebuild pricing, talent, and delivery around it will protect their margin.

JUL 23, 2026 · GEORGE WILLIAMS

Funnel on a dark slate ground, labelled "Bank of America, Q2 2026 disclosure", reading "300 approved. 34 running." Three narrowing stages: approved, more than 300; live, 114; running, 34. Below it: your AI program looks like it is compounding, why don't your results?

300 Approved. 34 Running.

The most useful thing any executive said about AI this year came from a bank CEO telling shareholders his AI program would not lift margins.

JUL 16, 2026 · GEORGE WILLIAMS

Diagram on a dark ground labelled "deployment versus proof", headed "What crosses the line without a record cannot be reconstructed later." Live events move along timelines into a capture boundary. What passes through becomes a verifiable record carrying event, context and time, and a decision with its trace. What bypasses it leaves no trace. A footer line reads: evidence you did not capture cannot be captured later, instrumentation is part of deployment, not a forensic step after it.

You Are Deploying AI Faster Than You Can Prove What It Did

Companies describe agentic AI as a capability in their marketing and as an exposure in their SEC filings. The filings are the more informative document.

JUL 9, 2026 · GEORGE WILLIAMS

Cover on a deep blue ground with an amber field welded to the left edge labelled "evidence", and a kicker reading "x402 on Base, 280 days". Headed "AI payment volume is not proof." Below it: 136,708,672 settlements, and in white, $187,861 provably real. The exact figure is $187,861.35.

The Agent Economy Booked $44 Million. Two Thirds of Its Transactions Were With Itself.

A population-scale measurement found 136 million settlements and could prove $187,861 of them reached an outside party. The gap is not fraud. It is a missing layer.

JUL 2, 2026 · GEORGE WILLIAMS

Video call grid on a deep navy ground, headed "$25 million." Below it: one real person on the call, everyone else was generated. Six tiles, five drawn in dashed red and marked generated, labelled CFO and colleague, and one drawn in solid blue and marked real, labelled him. Beneath: HK$200 million approved after the call. A footer line reads: the control was recognition, and recognition is now a forgeable artifact.

$25 Million. One Real Person on the Call.

A finance worker checked the request, joined a video conference, recognized his colleagues and approved the transfers. The colleagues were generated.

JUN 18, 2026 · GEORGE WILLIAMS

Two bars on a deep navy ground, headed "You use 2 percent." Below it: of the access you already hold, your agent will use all of it. The upper bar, labelled human identity permissions actually used, is nearly empty and reads 2 percent used, 98 percent idle. The lower bar, labelled an agent holding the same grant, is filled completely and reads all of it. Beneath: 51,000 permissions grantable across the major clouds. A footer line reads: latent access becomes active access the moment the holder stops being human.

You Use 2 Percent of Your Access. Your Agent Will Use All of It.

The permissions nobody exercises were harmless while only people held them.

JUN 4, 2026 · GEORGE WILLIAMS

Diagram on a deep navy ground, headed "700 companies." Below it: one chatbot’s credential, no password, no phishing, no prompt. A single amber box labelled one token connects by curved lines to a dense field of 700 blue dots labelled 700 plus organizations, data exported, with nobody was phished beneath. A footer line reads: an identity you cannot revoke is not an identity.

A Chatbot's Token Opened 700 Companies.

Nobody was phished. The attacker used a credential that belonged to an AI integration, and most organizations cannot revoke one.

MAY 21, 2026 · GEORGE WILLIAMS

Flow diagram on a deep navy ground, headed "No click required." Below it: one email, your files, CVE-2025-32711, severity 9.3. Three boxes joined by arrows read email unopened, the assistant, and outside, with your files beneath the last and reads everything in one stream above the middle. A bordered panel below reads privileges required none, user interaction none. A footer line reads: the model did exactly what it was told, not by you.

One Email. No Click. Your Files.

A critical vulnerability in Microsoft 365 Copilot needed no user interaction at all. The model was working correctly.

MAY 7, 2026 · GEORGE WILLIAMS

Two panels on a deep navy ground, headed "The exam has errors." Below it: 6.49 percent of the questions are wrong, and the grader changes its mind. On the left, a grid of small squares with 22 marked in red, labelled 6.49 percent of questions contain errors, over a note reading MMLU, 5,700 questions re-annotated. On the right, three rows labelled run one, run two and run three, each with a marker at a different position, headed the same work graded three times, and labelled same work different mark, agreement with itself approached zero. A footer line reads: a benchmark score is evidence about a laboratory, not about your process.

The Exam Has Errors. The Grader Changes Its Mind.

Every AI decision you have approved rests on an evaluation. Three findings say the evaluation is shakier than the model.

APR 23, 2026 · GEORGE WILLIAMS

Fan diagram on a deep navy ground, headed "80 different answers." Below it: one thousand identical questions, temperature zero, nothing was random. A single blue line labelled one thousand identical requests runs from the left into a green gate labelled token 103, after which it fans into many faint lines. Two are labelled, 992 Queens New York in blue and 8 New York City in red, with 80 distinct completions above them. A footer line reads: the nondeterminism is not in the model, it is in the load.

One Thousand Identical Questions. Eighty Different Answers.

Temperature zero exists to make a model repeatable. It does not, and the reason is not randomness.

APR 9, 2026 · GEORGE WILLIAMS

Curve on a deep navy ground, headed "Eleven of thirteen." Below it: models advertising 128,000 tokens, measured at 32,000. A line starts at 99 percent on the left and falls away as the axis extends right, with a marker at 32,000 tokens labelled below half of baseline. A footer line reads: the window is a capacity, not a competence.

128,000 Tokens. Eleven of Thirteen Failed at 32,000.

The advertised context window is a capacity. What the model can actually use is a much smaller number, and nobody prints it.

MAR 26, 2026 · GEORGE WILLIAMS

Three grouped columns on a deep navy ground, headed "14 ways to fail." Below it: 1,600 traces across seven frameworks. The columns are labelled system design, inter-agent misalignment and task verification, with tally marks distributed across them. A footer line reads: not one of the three categories is model capability.

1,600 Traces. 14 Ways to Fail. None of Them Is the Model.

Multi-agent systems break on specification, coordination and verification. That list is an org chart, not a research problem.

MAR 12, 2026 · GEORGE WILLIAMS

Distribution curve on a deep navy ground, headed "The tails go first." Below it: 74 percent of new pages contain AI content. A bell curve is drawn twice, the second with its outer edges faded away and the centre unchanged, labelled generation one and generation four. Beneath the faded edges a label reads the rare cases. A footer line reads: 25.8 percent of new pages were purely human.

74 Percent of New Pages Contain AI. The Tails Go First.

Training on generated data does not produce nonsense. It produces a model that has forgotten the rare cases.

FEB 26, 2026 · GEORGE WILLIAMS

Grid of 100 small key symbols on a deep navy ground, headed "Most of them still work." Below it: 23.8 million secrets reached public repositories, and three years on most of the keys still open something. The grid is labelled 100 keys leaked in 2022. Seventy of the keys are drawn solid in amber and thirty are faded almost to nothing, over a line reading 70 of them still open something. To the right, 23.8 million leaked in 2024, up 25 percent, and AI assistant repositories leak 40 percent more. A footer line reads: a leaked credential is a condition, not an event.

23.8 Million Secrets. Most of Them Still Work.

A leaked credential is not an incident. It is a condition, and the assistant writing your code makes more of them.

FEB 12, 2026 · GEORGE WILLIAMS

Two bars on a deep navy ground, headed "The tools you did not buy." Below it: one in five breaches involves AI nobody approved. A lower bar labelled sanctioned AI sits at 13 percent of breaches, and an upper bar labelled shadow AI sits at 20 percent, with a note reading plus $670,000 per breach. A footer line reads: 73.8 percent of accounts were never yours to revoke.

$670,000 for the Tools You Did Not Buy.

One in five breaches now involves AI nobody approved. The account is the problem, not the model.

JAN 29, 2026 · GEORGE WILLIAMS

Diagram on a deep navy ground, headed "Flagged. And downloadable." Below it: around 100 models carried live payloads. A file labelled model weights passes a scanner marked unsafe, which does not stop it, and continues into a box labelled your machine, from which an arrow labelled reverse shell leaves for an outside server. A footer line reads: a model file is not data, it is a program that runs when you open it.

The Scanner Flagged It. The Download Worked.

Around 100 models on a public hub carried live payloads. Loading one opened a shell on the machine that loaded it.

JAN 15, 2026 · GEORGE WILLIAMS