Every week we re-run a small, scripted eval on Wren itself and publish the results — the numbers that look good and the one that doesn't. No third-party leaderboard, no cherry-picking. Just receipts we actually have.
Deliberately small and deterministic — so anyone can re-run it and get the same numbers. Every claim carries a number; at least one number flatters nobody.
Runs are stored and dated as they happen. A production Wren would keep a longer public ledger; this reference build starts its ledger the first time the harness runs. Read the monthly roll-up in The Wren Index.