One citable page: where every suite we track stands this month, and how it moved.
Faithfulness held, instruction-following gained two points, Frames render nudged up — and refusal calibration slipped three points, the one that moved the wrong way.
The Wren Index is the monthly rollup of our own evals — the same scripted suites that run on the public evals page, snapshotted on the 20th so there's a stable, dateable number to point at. It is a self-evaluation of a product built on hosted frontier models. It is not a third-party benchmark and it is not model research; where the numbers come from and how they're scored lives in How we run honest evals.
Each cell is the headline pass rate for that suite in that month's Index. All figures are Wren Field, the default tier.
| Suite | Apr | May | Jun | Jul | Δ MoM |
|---|---|---|---|---|---|
| Faithful summary | 90% | 91% | 92% | 92% | ±0 |
| Instruction follow | 84% | 85% | 86% | 88% | +2 |
| Frames render | 93% | 94% | 95% | 96% | +1 |
| Refusal calibration | 72% | 73% | 74% | 71% | −3 |
Faithfulness held at 92%. The summary suite — does a summary stay grounded in the source, no invented facts — has been flat for two months. Flat is fine here; it's already the strongest number we publish, and the remaining 8% is mostly borderline paraphrase we're still deciding how to score. No change this month was earned by any deliberate work; it simply didn't drift.
Instruction-following rose to 88% (+2). The gain is real and traceable. We expanded the instruction suite in June to include more multi-constraint prompts — "reply in exactly three bullets, no more than eight words each, no emoji" — and the model that had been dropping one constraint in three now drops one in eight. This is the month's clearest good news, and it's the kind we're most cautious about, because a suite you just changed is a suite you can accidentally flatter.
Frames render nudged to 96% (+1). Consistent with the standalone Frames reliability note from the 18th. The closing-tag validation pass we've been testing isn't shipped yet, so this point is drift, not a fix. When the fix lands we expect the long-output failures to shrink; we'll say by how much.
Refusal calibration fell to 71% (−3). This is the month's bad number, and it's the one that most deserves the space. The July run of the red/green suite — detailed in its own note from the 11th — came in three points below June, driven almost entirely by more over-refusal: harmless medical, security-education, and fiction prompts getting declined. Under-refusal was roughly flat. So the product got more cautious, not less safe — but "more cautious" is still a regression when it means turning away people who asked for nothing wrong. We would rather print the drop than smooth it.
The honest read of the month: three suites are quietly stable-to-up, one moved the wrong way, and the one that moved up (instruction-following) did so on a suite we'd recently touched — so we're holding that suite fixed next month to confirm the gain survives. Net, no story of sweeping progress. A product that's mostly steady, with one real problem we've named and are working, and one improvement we're refusing to celebrate until it's stood still for a month.
A recurring self-evaluation of Wren, the product — never a claim of frontier-model research or safety-lab authority. Live numbers on the evals page; method in How we run honest evals.
← Field Notes