← Back to blog
Product

Introducing Wren Grove

Our most capable tier, built for the problems worth slowing down for. Grove thinks longer — and shows you the thinking.

Most of what you ask an AI to do doesn't need much deliberation. Tighten this paragraph. Rename these variables. Draft the email. For that, our faster tiers — Seed and Field — are the right tools, and they'll keep being the default. Speed is a feature.

But some problems deserve a pause. A proof that has to hold together across a dozen steps. A refactor where one wrong assumption early poisons everything after it. A research question where the honest answer is "it depends," and the value is in laying out exactly what it depends on. For those, we're releasing Wren Grove — our most capable tier, tuned for extended thinking, available today on Pro.

What Grove is

Grove is the third rung of the Wren ladder, after Seed (fastest) and Field (our balanced default). Like the others, it's served through a hosted frontier model — we're a product company, not a model lab, and we'll always say that plainly. What makes Grove Grove is how we've tuned and scaffolded it: longer reasoning budgets, a bias toward working problems out step by step, and a willingness to sit with ambiguity instead of rushing to the tidiest-sounding answer.

The visible difference is the thinking. Ask Grove something hard and you'll see a shimmer while it works, then a collapsed Thought for 8s line you can expand to read the whole chain. The reasoning isn't a byproduct we hide — it's part of the answer. If you don't trust a conclusion, you can check the steps that got there.

Extended thinking is worth it exactly when a wrong first instinct is expensive.

What extended thinking is good for

We've spent the last two months watching where Grove earns its keep. A rough map of where the extra deliberation pays off:

  • Multi-step reasoning — math, logic, and planning where an early mistake compounds. Grove catches its own contradictions more often because it has room to notice them.
  • Serious code work — a real refactor, a tricky bug, a design that has to account for edge cases. Grove pairs beautifully with Frames: it reasons through the approach, then renders the result live beside the chat.
  • Judgment calls — comparing options, weighing tradeoffs, pressure-testing a plan. It's better at saying "here's the case for each" than at pretending there's one obvious winner.
  • Careful reading — long documents where the answer is buried in a qualifier three paragraphs down.

Where it's not worth it: quick edits, casual back-and-forth, anything where you'd rather have a good answer now than a slightly better one in twelve seconds. Reach for Seed or Field there. Grove is a specialist, not an upgrade you should leave on for everything.

The honest numbers

We won't hand you a leaderboard we curated to make ourselves look good. What we'll do is show you our own weekly self-eval, run the same way every week, published whether the numbers flatter us or not. Here's how Grove compares to Field on the latest run:

SuiteWren FieldWren Grove
Multi-step reasoning74%89%
Instruction follow88%91%
Frames render (valid HTML)96%97%
Faithful summary (grounded)92%93%
Refusal calibration71%76%

Two things stand out, and we want to name both. Grove is meaningfully better at reasoning — that's the whole point, and it shows. But on summary and Frames rendering, it's barely ahead of Field, because those tasks don't reward deliberation much. Paying the latency tax there buys you almost nothing. And refusal calibration, at 76%, is still not where we want it. Grove sometimes over-refuses a benign request and occasionally under-refuses a sharp one. We publish that number because pretending otherwise would make this whole page worthless. The full methodology and every dated run live on our evals page.

Availability

Grove is included with Pro ($17/mo), alongside Seed and Field, higher limits, and unlimited Projects and Frames. Max and Team plans include it too, with more headroom. On the Free plan you get Seed and Field with fair daily usage — plenty for real work — but Grove stays behind Pro, because extended thinking is genuinely more expensive to run and we'd rather be straight about that than quietly ration it.

You switch tiers from the model picker in any chat. Start a conversation on Field, hit a problem that deserves more care, and promote it to Grove mid-thread — the context carries over.

Limits — the part we won't bury

Grove is slower. That's the deal you're making: seconds of thinking for a better answer. On easy questions that trade is a bad one, so don't default to it.

It can over-think. Given something simple, Grove occasionally talks itself into a more complicated answer than the moment needed. If a response feels ornate, drop to Field and compare.

It is not a knowledge oracle. Longer reasoning doesn't fix a wrong fact — it just reasons more confidently from it. For anything load-bearing, check the sources. Grove will show you its steps precisely so you can.

We built Grove for the moments where you'd want a careful colleague to think it through with you, out loud, where you can see the work. That's the whole product philosophy in one tier. If you have Pro, it's already in your model picker. Open a chat and give it the thing you've been chewing on.

← Back to blog See this week's evals →

Try Grove on a hard one.

Free to start; Grove unlocks with Pro. Bring the problem you've been avoiding.