AI Evaluation
How model behavior is measured, and the ways the measurement turns out to be shakier than the model.
5 pieces
The Exam Has Errors. The Grader Changes Its Mind.
Every AI decision you have approved rests on an evaluation. Three findings say the evaluation is shakier than the model.
APR 23, 2026
One Thousand Identical Questions. Eighty Different Answers.
Temperature zero exists to make a model repeatable. It does not, and the reason is not randomness.
APR 9, 2026
128,000 Tokens. Eleven of Thirteen Failed at 32,000.
The advertised context window is a capacity. What the model can actually use is a much smaller number, and nobody prints it.
MAR 26, 2026
1,600 Traces. 14 Ways to Fail. None of Them Is the Model.
Multi-agent systems break on specification, coordination and verification. That list is an org chart, not a research problem.
MAR 12, 2026
74 Percent of New Pages Contain AI. The Tails Go First.
Training on generated data does not produce nonsense. It produces a model that has forgotten the rare cases.
FEB 26, 2026