The moment that proved the model
Three AI builds found what review had missed
Until then, the spec had only been checked by reading it. So I tested it the way teams would use it, by building from it.
Three fresh sessions
Each AI session got only the pinned spec and built the same components from scratch.
Same failure, three times
All three failed the rendered gate on one button state. Every token was valid, so no static check could see it.
Rule fixed, process changed
I corrected the rule. Any release that changes a rule now gets the same test.
Recreation. Ratios are calculated live from the colours shown. AA needs 4.5:1.
A rule that broke its own standard
The button rule sent the label to a surface at 3.34:1. The same paragraph cited that ratio when it banned the pairing elsewhere. Reading missed it, but measuring the pixels didn't.
A rule that contradicted itself
A table rule asked for two things that can't both be true. The agent picked one and wrote clean code, and noted the conflict. So the release process now reads the agent's notes as well as the result.
Testing a spec by building from it finds defects that reviewing it never will.