The Wren Index — July 2026 →
Month-over-month across every suite we track. Faithfulness held, instruction-following ticked up, refusal calibration slipped — and here's why.
Dated notes on how Wren behaves, measured. Every claim carries a number — and at least one number flatters nobody.
Ongoing notes on Wren's product behavior — how a specific surface actually performs under a repeatable test. Published when we have something worth measuring, not on a clock. Claim, method, data, limitations.
A recurring monthly report. One citable page that rolls up the month's eval trend across every suite we track — summary faithfulness, instruction-following, Frames render, refusal calibration — with month-over-month numbers and a plain narrative.
Month-over-month across every suite we track. Faithfulness held, instruction-following ticked up, refusal calibration slipped — and here's why.
200 build prompts, one metric that matters: does the thing show up on screen. 96% valid HTML — and a clear look at the 4% that break.
A small red/green suite of should-answer and should-refuse prompts. 71% correct — the least flattering number we publish, and the most useful.