"Yes, you can take a cut of your worker's tips."

That was New York City's official small business chatbot, in early 2024. That is illegal in New York.

The same bot said a restaurant could go cash-free, which the city had banned in 2020. Asked whether landlords have to accept tenants with housing vouchers, it said no. Ten people at The Markup asked separately. It told all ten of them no. That is also illegal.

The page said the bot was "trained to provide you official NYC Business information."

The city added a disclaimer and left it running for almost two more years.

The bot was doing what language models do, producing a plausible answer to whatever it is asked. Nobody had decided what should happen when the plausible answer was wrong.

The Question Nobody Writes Down

Every company deploying AI is making a version of that decision, usually without naming it. The model reads the documents, interprets the conversation, produces an answer. What almost nobody writes down is whether that answer gets to be the decision.

The distance between what AI can do and what it should be allowed to own is the Authority Gap. What you pay for crossing it without boundaries is the Token Tax.

Tokens are the visible part, the units consumed whenever a model processes or generates information. They are also the only part that arrives on an invoice. The governance, operational and trust costs never show up as a line item at all.

Two panels on a dark slate ground. On the left, headed "what gets billed", a model usage invoice showing a single line item, tokens, at one times, described as visible, metered, invoiced. An arrow points right to a panel headed "what the operating model absorbs", listing three numbered items: operations, meaning monitoring, remediation and rework; governance, meaning approval, traceability and ownership; and trust, meaning confidence lost when authority drifts. A footer line reads: the invoice shows only one cost, three more are absorbed.
Only the first of the four arrives as a line item.

This is not the ordinary cost of using AI where judgment creates value, and it does not mean AI should never act on its own. The tax arrives when the scope of that action, its escalation path and its accountable owner are undefined.

To see where it comes from, separate any AI-enabled process into three layers.

Intelligence. What is happening. AI interprets information, recognizes patterns and generates recommendations. This is where probabilistic systems earn their keep, working with uncertainty instead of requiring every case to be specified in advance.

Authority. What is allowed to happen. Business rules, permissions, approvals, policies and accountable ownership. This determines who, or what, may turn an interpretation into a business commitment.

Execution. How the authorized action happens reliably. Consistency, auditability, traceability and control. It turns a decision into action while preserving evidence of what happened and why.

Companies collapse all three into one and assume that if AI understands the situation, it should decide the outcome.

Not everyone does.

The Firm That Had Every Reason to Go Further

Morgan Stanley had more license to push than New York City did.

It was the first major Wall Street firm to put a bespoke GPT-4 tool in its employees' hands, with early access to the strongest models available and a research library of about 100,000 documents to point them at.

It built neither of its two tools to talk to a client.

The Assistant retrieves information from approved internal content for advisors. Debrief, with client consent, takes notes in client meetings, surfaces action items and drafts the follow-up email. Then it stops. Every draft goes to an advisor who reads it, edits it and decides whether it gets sent.

As of June 2024 the firm reported that 98% of advisor teams had adopted the Assistant. One of them, Don Whitehead, said Debrief saved him about half an hour per meeting on notetaking alone. Company-reported numbers, two years old, not an independent study. The pattern still holds. Boundaries did not stop the system from becoming useful.

Morgan Stanley bought the most capable AI it could get and pointed it at the advisor rather than the client. The difference with MyCity is not government versus finance. It is where authority sits.

Authority Gets Borrowed

In New York, nobody ever put it anywhere. That is the part of the story easiest to miss.

MyCity never had execution authority. It could not issue a permit, file a form or move a dollar. It only answered questions.

So nobody handed AI the power to act. Nobody handed it anything. Put a language model behind a government's logo, tell users it is trained on official information, and it acquires the institution's authority without a single person deciding to grant it.

Borrowed authority is harder to govern than the granted kind, because there is no approval to point at afterward.

A comptroller audit two years later found the bot still giving different answers to identical questions. Of the users who bothered to rate it, 71.4% were dissatisfied. The city was still reporting accuracy above 95%.

Two bars on a dark slate ground, labelled "New York City, July to August 2025", headed "The metrics disagree." A subtitle reads: more than 2,200 questions were asked, seventy people rated an answer. The first bar, reported accuracy, is nearly full at 95 percent or more. The second, negative feedback, reaches 71.4 percent. A footer line reads: 50 of 70 respondents submitted negative feedback, and 20 of 70 did not rate it negatively.
Both numbers are the city's own. They measure different things, and only one was reported upward.

The failure was not that generative AI produced uncertain answers. The failure was letting uncertain answers borrow the authority of government.

The Four Costs

The Token Tax rarely arrives as one dramatic failure. It compounds. Variable outputs require review; review creates workflows; workflows obscure ownership; unclear ownership slows everything behind it.

Financial. You pay for model reasoning where a validated rule would do, then pay again for review, remediation and rework.

Operational. You add monitoring, testing and exception handling to manage variability you created by putting AI in a role that never required interpretation.

Governance. A recommendation becomes an action, an action becomes a commitment, and ownership becomes hard to locate. Approval evidence comes back incomplete and auditors cannot reconstruct what happened.

Trust. Employees hesitate to rely on systems they cannot understand. Customers distrust systems that cannot explain themselves. Executives refuse to scale systems they cannot control.

Notice what all four have in common. None of them is a cost of intelligence. Every one is a cost of authority.

Authority Is Not the Enemy

None of this argues that AI should interpret and never act. Plenty of actions can be delegated safely, including routing a request, retrieving an approved record, preparing a draft or clearing a routine transaction inside defined limits.

The question is not whether AI acts. It is whether its decision rights match the consequences. A service agent might issue a refund below a threshold and escalate everything above it. A finance agent might assemble a payment package while release authority stays with an accountable approver. Autonomy can widen as reliability is demonstrated, as long as the limits stay explicit.

That is bounded authority. The system acts inside a defined envelope, and policy, permissions, monitoring and escalation determine where the envelope ends.

Watch the Envelope End

An invoice arrives and does not match the purchase order. The amounts are close but not equal, and the tolerance policy does not obviously cover the gap.

A theatrical system improvises. It picks the most plausible reading and posts.

A bounded system stops. It says what it found, what it cannot resolve, and what it needs. The question goes to the person who owns that tolerance, not to whoever happens to be nearest. They answer.

Then the part that matters. That answer can stay a one-time exception, become guidance the system applies to that process from now on, or become policy. Somebody decides which, and that decision is visible, attributable and reviewable later.

The exception is where authority stops being a diagram and becomes a name.

Two panels on a dark slate ground, labelled "when the system cannot decide", headed "The exception needs a decision record." On the left, an exception detected: invoice mismatch, policy does not resolve the gap, the system stops. An arrow points right to a decision record listing three fields: owner, a named accountable person; decision, approve, reject or make policy; and evidence, reason and time retained.
The record is what makes the boundary auditable rather than merely described.

The Objection Worth Taking Seriously

If a licensed professional reads every draft, what did you actually automate?

Fair question. The honest answer is that you automated the slow part, not the risky part. Reading a summary is faster than writing one. Checking a figure is faster than hunting for it. Morgan Stanley's 98% is an adoption number, not a productivity number, and adoption is easier to report than throughput.

The review bottleneck is real. It is also the price of being able to say who decided. Skip it and you do not avoid the cost. You meet it later, denominated in remediation instead of minutes.

The AI Monolith

Enterprise software already learned that packing too many responsibilities into one tightly coupled application creates fragility. Changes get harder to test, failures harder to isolate, ownership harder to define.

An AI monolith is not simply one large model. It is a system allowed to interpret, reach resources, decide and execute with no boundaries between those responsibilities. The trap springs when permissions, controls, escalation conditions and failure boundaries are not separated from one another.

The best AI architectures will not maximize autonomy. They will maximize useful autonomy.

You Have Probably Already Crossed It

Most organizations reading this have crossed the line somewhere. Not dramatically. A summarizer that started drafting the customer response. A classifier that began routing without review. An assistant that quietly started writing to the system of record, because writing to it was the obvious next step and nobody had written down that it should not.

The question is not whether it happened. It is whether anyone can tell you where.

Before assigning AI a role in any business process, ask: 1. Does this step genuinely require interpretation, or would a validated rule be sufficient? 2. Who has authority to turn the output into a business commitment? 3. What may the system do independently, and what conditions require escalation? 4. Which permissions and controls constrain those actions? 5. What evidence must be retained so the decision can be explained later? 6. What happens when confidence is low, information conflicts, or the system fails?

If the answers are unclear, you are not ready to give the system more responsibility. More autonomy will not resolve an undefined authority model. It will bury the ambiguity deeper.

What It Costs to Find Out Later

In February 2026 the MyCity chatbot's page went dark. The notice said the beta test had ended. Not that the answers had been wrong. That the test was over. The program around it had spent $100 million.

Nobody ever decided that chatbot should carry the authority of the city. That is exactly how it got it.

The same thing is happening inside your company right now, in smaller ways. The fix is not less AI. It is deciding, on purpose and in writing, what your systems are allowed to own.

Sources