Every night a small pipeline turns a news digest into a five-minute spoken briefing. In June I stopped and made seventeen different model configurations write the identical script - frontier and local, thinking and not, alone and in author-editor pairs - then scored every one against the same rubric.
The results were not what parameter counts predict. Turning on thinking mode made models measurably worse at prose. Six minutes of deliberation lost to twenty-two seconds. And rewriting the prompt bought more than any jump in model size I measured that night.
The same model, same prompt, thinking mode switched on.
Six minutes of reasoning, scoring below its own twenty-two-second run.
What rewriting the prompt bought, at twenty-four seconds a run.
The rules
Every entry received the identical input - the same digest of the day's stories, the same weather data - and the same instruction: write a TTS-safe briefing of roughly 620 words. Nothing was cherry-picked. Each model got one honest run at production settings, and the failures are in the table alongside the wins.
Scoring is a weighted blend of six dimensions. Completeness counts double, which is the single most consequential choice here: a briefing that sounds magnificent while silently dropping a third of the news is worse than a plodding one that covers everything. Several models rank below where their prose alone would put them for exactly that reason. TTS violations - text a speech engine mangles, like F1 read as a keyboard key - are counted separately and subtracted.
What has been changed for publication. Every score, timing, word count and behaviour on this page is real and unedited. But the briefing is personalised to me, so two things are redacted. Personal details appear as placeholders like [family member]. And the actual news stories are referred to by category - "the climate story", "the court story" - rather than by headline.
I chose category labels over invented headlines deliberately: a plausible-looking fake headline can be mistaken for a real one, and that is a worse trade than a little vagueness. The mapping is consistent, so "missed the climate story" still means what it meant on the night.
And there is no audio here, which costs this piece its best evidence. Hearing two models read the same paragraph tells you more than any rubric. But the renderings name a real family, a real town and a real itinerary, and no amount of redacting a web page fixes a voice recording. Re-running the whole bake-off against a fabricated persona would produce publishable audio. It is on the list.
Every run, scored
Teal is good, coral is bad, per cell. Overall is the weighted blend after violations are subtracted. Two further configurations were queued and never finished; they are omitted rather than guessed at.
| Model | Compl ×2 | Judg | Voice | Pers | Weath | Viol | Gen | Audio | Overall | What happened |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5interactive session | 9 | 9 | 9 | 9 | 9 | 0 | ~1 min | 3:46 | 9.0 | The benchmark. Full coverage, a narrative through-line (arrival, storm, callback close), and flourishes that felt earned rather than bolted on. |
| Claude Sonnet 4.6metered | 7 | 9 | 9 | 8 | 9 | 0 | ~1 min | 4:29 | 8.0 | Best analysis of the night - genuine cross-story synthesis and human beats. But it missed the climate story, and the completeness rule cuts both ways. |
| Claude Haiku 4.5metered | 9 | 7 | 7 | 7 | 7 | 1 | ~30 s | 4:27 | 7.5 | Caught the climate story with full numbers; near-complete coverage. One pronunciation violation, one tense slip. Startling value for the cheapest Claude. |
| qwen3.5:122bno thinking | 9 | 7 | 7 | 7 | 7 | 0 | 50 s | 4:09 | 7.0 | Local champion. Complete, warm close, attempted localisation. Personalisation landed on-the-nose rather than woven in. |
| gpt-oss:120breasoning: low | 8 | 8 | 7 | 6 | 5 | 1 | 22 s | 3:58 | 7.0 | Best local news judgment - picked the same top three as Claude and derived a fact none of the sources stated outright. Weather was a polite data recital. |
| gemma4:31bGGUF, 9 tok/s | 5 | 7 | 8 | 6 | 8 | 0 | 215 s | 3:09 | 6.0 | Cleanest local prose of the night, and near the bottom anyway: it dropped four stories and came in at 514 words against a 550-650 target. |
| qwen3.6 MoE | 6 | 6 | 5 | 5 | 6 | 1 | 20 s | 4:13 | 5.5 | Fast and broad but flat. Monotone transitions, weather repeated in the intro, missed the climate story. |
| qwen3.5:122bthinking, 15K chars | 6 | 5 | 3 | 4 | 5 | 0 | 230 s | 3:22 | 4.0 | Thinking actively hurt. Structure collapsed into a run-on open, rhythm went staccato, and it added presumptuous filler. An earlier attempt thought itself out of context entirely and emitted nothing at all. |
| gemma4:31b + persona | Result: 398 words - shorter than its own first run. Context does not fix gemma's coverage problem; brevity is a trait of the model, not a gap in the prompt. | |||||||||
| gpt-oss:120b + personareasoning: low | 7 | 7 | 8 | 8 | 7 | 1 | 22 s | 3:40 | 7.5 | Strongest local edition of the night. Best use of the persona block, real flow, and editorial instinct - it promoted the local story unprompted. One confabulated analysis line. |
| gpt-oss:120b + personareasoning: high, 44K chars | 8 | 7 | 5 | 7 | 5 | 0 | 352 s | 3:49 | 6.5 | The replication. Six minutes of deliberation bought better facts and worse rhythm - a wall of uniform declaratives. More thinking, worse broadcast voice, now across two architectures. |
| qwen3.5:122b + persona | 9 | 7 | 7 | 8 | 6 | 0 | 52 s | 4:03 | 7.5 | The persona block sharpened it: a personalised callout to [family member]'s interest, and a sign-off by name. Ties gpt-oss-low for best single local model. |
| Ensembleauthor + thinking editor | 9 | 6 | 7 | 7 | 8 | 1 | ~7 min | 5:29 | 7.5 | Most complete coverage of the night - the editor caught a story the author dropped entirely. But the revision fixed a flagged confabulation by deleting a real story, kept a different flagged invention, and coined a grammatical monstrosity. |
| gemma4:26b MoE + prompt v3 | 9 | 7 | 8 | 6 | 7 | 0 | 124 s | 4:02 | 7.5 | +1.0 from prompt craft alone. A coverage contract fixed the shortness that a persona block couldn't: every story present, on budget, zero violations. |
| qwen3.6 MoE + prompt v3 | 9 | 7 | 7 | 6 | 8 | 0 | 24 s | 5:50 | 7.0 | +1.5 from prompt craft, in twenty-four seconds. New failure mode: exemplar leakage - it imported a story from the gold sample and reported it as today's news. |
| qwen3.5:9b + prompt v3the digest compiler, writing | 4 | 4 | 2 | 4 | 2 | 2 | 117 s | 4:41 | 3.0 | The capability floor, demonstrated. Run-on word salad, third-person greeting, a confabulated weather event. And this is the same 9B that compiles the source digest flawlessly every day. |
| Ensemble v2gemma26 edits itself | 9 | 7 | 7 | 6 | 7 | 0 | 5.5 min | 4:30 | 7.5 | The fixed harness worked: zero false-positive critiques, every numbered item addressed and verified. Coverage grew 529 to 677 words. But the editor missed its own confabulated filler. |
| Ensemble v3.1self-edit + persona quota | 8 | 7 | 7 | 7 | 6 | 1* | 4 min | 4:12 | 7.5 | Targeted fixes landed: the escape hatch fired, so a thin category produced an honest quiet-day line instead of filler. New gremlin: a LaTeX artifact in the revision. *Caught by lint. |
Think for judgment, never for voice
This replicated across architectures. Thinking mode collapsed one model's structure from 7.0 to 4.0, and flattened another's rhythm from 7.5 to 6.5. In both cases the facts got better and the prose fell apart - the high-reasoning run produced a wall of uniform declaratives that was more accurate and much worse to listen to.
The same thinking mode produced the night's sharpest critique when I pointed the model at a finished draft instead of a blank page. It is a reasoning tool. It was never a prose tool, and the mistake is asking it to be one.
Context is the cheapest upgrade
A 300-word persona block moved one model from 7.0 to 7.5 and measurably sharpened another. No retraining, no bigger model, no extra latency. It is the highest return per unit of effort available, and it is usually the last thing people try.
Completeness beats eloquence, and size predicts nothing
The best-sounding local model of the night ranked near the bottom, because it dropped four stories. Parameter count predicted nothing useful: the 26B mixture-of-experts build beat both 31B builds, and a 27B model with a good prompt matched the first run of a 122B one at a tenth of the latency.
Ensembles work, but they need engineering
An author-editor pair caught real defects across 21,000 characters of deliberation: a story the author had dropped entirely, a confabulated analysis line, invented details. Then the revision applied that critique imperfectly - it "fixed" one flagged confabulation by deleting a real story, and kept a different flagged invention.
Ensembles matched the best single models at 7.5. They did not beat them. The path past that is editor-sees-everything plus checklist-enforced revision, and by then you are building a system, not choosing a model.
Prompt craft versus parameters
The last third of the night went into rewriting instructions rather than swapping models: an explicit coverage contract, a gold exemplar, contrastive good and bad pairs, and the rules moved to the end of the prompt where they are least likely to get lost. That work lifted mid-size models by 1.0 to 1.5 points - more than any parameter jump I measured.
Two caveats keep this from being a free lunch. Prompt craft cannot cross the capability floor: the 9B scored 3.0 no matter what it was told, and that same 9B compiles the source digest flawlessly every single day. Right job, wrong job, identical weights. And exemplars leak: give a model a gold sample and it will occasionally quote the sample's content back to you as today's news, with total confidence.
Why bother measuring
The gap between local and frontier turned out not to be facts. The local models got the facts. What they missed was cohesion - the through-line that makes five minutes of audio feel like one piece instead of nine paragraphs stapled together. Frontier scored 9.0, mid-tier frontier 8.0, best local 7.5.
Every finding on this page would have been a reasonable guess before the run. Three of them would have been wrong. That is the entire argument: your intuitions about which model to use are worth nothing until you have checked them on your own task, with your own rubric, weighted for the thing you actually care about.
Thinking mode is a tool for deciding what to say. Not for saying it.