unPaper lab
Measured, not imagined.
What we measure about how AI systems build slides, and what we changed because of it. Every number here is generated from the runs themselves.
- 1 August 2026
We told the AI it could stop early. It heard: do less
A safety net for long decks turned into a licence to write shorter ones — and the escape hatch it offered was never once used. One sentence fixed it.
- 1 August 2026
We deleted two of our own layouts, and the two zeroes meant different things
Two layouts had never been used by any model, in any run. Only one of those zeroes was evidence that nobody wanted it — the other was evidence about us.
- 1 August 2026
We expected a gap between cheap and frontier models. There isn't one
We believed cheap models leaned on our layout catalogue twice as hard as frontier ones. Measured properly — after buying the decks we were missing — every model tested uses it in four or five briefs out of five.
- 1 August 2026
What our rules cost an AI, and what they buy
The same models, the same briefs, twice — once working freely and once under unPaper's constraints. The rules did not cost structure; they bought conversion.
- 31 July 2026
We made our prompt bigger, then measured where bigger stops helping
Three sizes of the same prompt, four open-weight models, eighteen decks. Nothing broke — but past a point the extra example buys length, not better slides, so we built the largest version and did not ship it.
- 31 July 2026
Five structures we could not name, and one improvement we measured and threw away
A wider sample corrected two of our own decisions, so we shipped what the models actually build — and refused a converter change we could not prove was free.
- 31 July 2026
Seven AI models, one slide brief, and the diagram none of them had a name for
We asked seven families of model to build the same slides, measured what came back, and changed the template because of it — seven elements added, two removed.