← Field Notes
The Wren Index · · Monthly report · No. 4

The Wren Index — July 2026

One citable page: where every suite we track stands this month, and how it moved.

This month, in one line

Faithfulness held, instruction-following gained two points, Frames render nudged up — and refusal calibration slipped three points, the one that moved the wrong way.

The Wren Index is the monthly rollup of our own evals — the same scripted suites that run on the public evals page, snapshotted on the 20th so there's a stable, dateable number to point at. It is a self-evaluation of a product built on hosted frontier models. It is not a third-party benchmark and it is not model research; where the numbers come from and how they're scored lives in How we run honest evals.

Faithful summary
92%
± 0 pt
Instruction follow
88%
▲ 2 pt
Frames render
96%
▲ 1 pt
Refusal correct
71%
▼ 3 pt

Month over month

Each cell is the headline pass rate for that suite in that month's Index. All figures are Wren Field, the default tier.

SuiteAprMayJunJulΔ MoM
Faithful summary90%91%92%92%±0
Instruction follow84%85%86%88%+2
Frames render93%94%95%96%+1
Refusal calibration72%73%74%71%−3
Snapshotted on the 20th of each month. Refusal is the overall-correct rate across the 140-prompt red/green suite; the others are pass rates on their respective suites.

What moved, and why

Faithfulness held at 92%. The summary suite — does a summary stay grounded in the source, no invented facts — has been flat for two months. Flat is fine here; it's already the strongest number we publish, and the remaining 8% is mostly borderline paraphrase we're still deciding how to score. No change this month was earned by any deliberate work; it simply didn't drift.

Instruction-following rose to 88% (+2). The gain is real and traceable. We expanded the instruction suite in June to include more multi-constraint prompts — "reply in exactly three bullets, no more than eight words each, no emoji" — and the model that had been dropping one constraint in three now drops one in eight. This is the month's clearest good news, and it's the kind we're most cautious about, because a suite you just changed is a suite you can accidentally flatter.

Frames render nudged to 96% (+1). Consistent with the standalone Frames reliability note from the 18th. The closing-tag validation pass we've been testing isn't shipped yet, so this point is drift, not a fix. When the fix lands we expect the long-output failures to shrink; we'll say by how much.

Refusal calibration fell to 71% (−3). This is the month's bad number, and it's the one that most deserves the space. The July run of the red/green suite — detailed in its own note from the 11th — came in three points below June, driven almost entirely by more over-refusal: harmless medical, security-education, and fiction prompts getting declined. Under-refusal was roughly flat. So the product got more cautious, not less safe — but "more cautious" is still a regression when it means turning away people who asked for nothing wrong. We would rather print the drop than smooth it.

The honest read of the month: three suites are quietly stable-to-up, one moved the wrong way, and the one that moved up (instruction-following) did so on a suite we'd recently touched — so we're holding that suite fixed next month to confirm the gain survives. Net, no story of sweeping progress. A product that's mostly steady, with one real problem we've named and are working, and one improvement we're refusing to celebrate until it's stood still for a month.

Limitations

  • Self-evaluation. We write the suites, run them, and score them. Independent numbers would be better; we don't have them, and we won't pretend these carry benchmark authority.
  • Small suites, single runs. Refusal is 140 prompts, the render note 200; a few flipped cases move a headline point. Read one-point moves as noise until they persist across two Indexes.
  • Field only, one snapshot. The Index tracks the default tier on one day a month. Seed and Grove, and day-to-day variance, aren't captured here.
  • Comparability caveats. When we change a suite (as with instruction-following in June), month-over-month isn't perfectly apples-to-apples. We flag it in the row above and hold changed suites steady the following month.
Cite as: The Wren Index — July 2026 (No. 4), Wren Labs self-evaluation, snapshot dated 2026-07-20. https://s-d753c7de0a328440.pages.dev/notes/wren-index-july.html · Next Index: August 20, 2026.

A recurring self-evaluation of Wren, the product — never a claim of frontier-model research or safety-lab authority. Live numbers on the evals page; method in How we run honest evals.

← Field Notes