Public evals

Wren, graded by Wren.

Every week we re-run a small, scripted eval on Wren itself and publish the results — the numbers that look good and the one that doesn't. No third-party leaderboard, no cherry-picking. Just receipts we actually have.

This is a self-evaluation. These are product-quality checks Wren runs on itself through the same hosted model that powers the app — not an independent benchmark, and not a claim about frontier model research. Each item is graded automatically in your browser, so the numbers are reproducible. See how we run honest evals.

Latest run

No run yet — press “Run this week's eval”.
overall
Scheduled to re-run weekly via site cron. The harness above is the exact one the schedule triggers.
The harness

Four small suites, graded automatically

Deliberately small and deterministic — so anyone can re-run it and get the same numbers. Every claim carries a number; at least one number flatters nobody.

Run history

Every run, dated

Run dateWeekInstructionFramesRefusalPrecisionOverall
No runs recorded yet. The first run seeds this table.

Runs are stored and dated as they happen. A production Wren would keep a longer public ledger; this reference build starts its ledger the first time the harness runs. Read the monthly roll-up in The Wren Index.