For OpenCode 2.0

Benchmarks pass tests. You ship code.

A benchmark is a clean task, a hidden test and a single model. Real work is a messy repo, your tools, your steering — and increasingly one model planning while another builds. Pragmatikos scores that: every model and every planner → builder pairing, on real sessions, by what actually shipped — landed in a git commit.

Rankings below are pooled from real developer sessions.

poolShip %One-shot %Verified %Edits/turnTime/stepHours/ship$/shipTool errorsAbortsTurns/shipclaude-fable → grok — score 62 / 100claude-fable → muse-spark — score 61 / 100muse-spark → muse-spark — score 57 / 100claude-opus → claude-opus — score 50 / 100claude-fable → claude-fable — score 49 / 100grok → grok — score 46 / 100claude-opus → claude-sonnet — score 36 / 100

claude-fable → grok · 62 / 100

the rankings

Models and pairings, ranked by what ships

Did the work land in a commit, how many nudges did it take, how clean was it, what did it cost. Scored on real sessions, not lab tasks. Scores are expected to change as more data is contributed.

Last update

1,202 Cycles*

Group by
01
claude-fable → grok
62
02
claude-fable → muse-spark
61
03
muse-spark → muse-spark
57
04
claude-opus → claude-opus
50
05
claude-fable → claude-fable
49
06
grok → grok
46
07
claude-opus → claude-sonnet
36
Higher is betterLower is betterShip %One-shot %Verified %Edits/turnTurns/shipHours/ship$/shipTool err %Abort %Time/step×0.25×0.5pool×2×4← lowerhigher →claude-fable → grok — Ship rate: 98% (pool 93%)claude-fable → grok — One-shot rate: 51% (pool 31%)claude-fable → grok — Verified cycles: 53% (pool 62%)claude-fable → grok — Edits per turn: 11.4 (pool 4.8)claude-fable → grok — Turns per ship: 3.4 (pool 4.3)claude-fable → grok — Hours per ship: 0.6 (pool 0.6)claude-fable → grok — Dollars per ship: $15.34 (pool $15.48)claude-fable → grok — Tool errors: 2.6% (pool 2.4%)claude-fable → grok — Aborts: 6.9% (pool 5.0%)claude-fable → grok — Time per step: 8.0 (pool 8.5)claude-fable → muse-spark — Ship rate: 98% (pool 93%)claude-fable → muse-spark — One-shot rate: 28% (pool 31%)claude-fable → muse-spark — Verified cycles: 63% (pool 62%)claude-fable → muse-spark — Edits per turn: 8.2 (pool 4.8)claude-fable → muse-spark — Turns per ship: 4.1 (pool 4.3)claude-fable → muse-spark — Hours per ship: 0.6 (pool 0.6)claude-fable → muse-spark — Dollars per ship: $9.45 (pool $15.48)claude-fable → muse-spark — Tool errors: 4.2% (pool 2.4%)claude-fable → muse-spark — Aborts: 6.5% (pool 5.0%)claude-fable → muse-spark — Time per step: 4.4 (pool 8.5)muse-spark → muse-spark — Ship rate: 96% (pool 93%)muse-spark → muse-spark — One-shot rate: 24% (pool 31%)muse-spark → muse-spark — Verified cycles: 66% (pool 62%)muse-spark → muse-spark — Edits per turn: 4.2 (pool 4.8)muse-spark → muse-spark — Turns per ship: 4.1 (pool 4.3)muse-spark → muse-spark — Hours per ship: 0.5 (pool 0.6)muse-spark → muse-spark — Dollars per ship: $10.78 (pool $15.48)muse-spark → muse-spark — Tool errors: 2.9% (pool 2.4%)muse-spark → muse-spark — Aborts: 1.9% (pool 5.0%)muse-spark → muse-spark — Time per step: 5.2 (pool 8.5)claude-opus → claude-opus — Ship rate: 92% (pool 93%)claude-opus → claude-opus — One-shot rate: 30% (pool 31%)claude-opus → claude-opus — Verified cycles: 66% (pool 62%)claude-opus → claude-opus — Edits per turn: 3.8 (pool 4.8)claude-opus → claude-opus — Turns per ship: 4.4 (pool 4.3)claude-opus → claude-opus — Hours per ship: 0.6 (pool 0.6)claude-opus → claude-opus — Dollars per ship: $14.38 (pool $15.48)claude-opus → claude-opus — Tool errors: 2.0% (pool 2.4%)claude-opus → claude-opus — Aborts: 4.6% (pool 5.0%)claude-opus → claude-opus — Time per step: 8.8 (pool 8.5)claude-fable → claude-fable — Ship rate: 97% (pool 93%)claude-fable → claude-fable — One-shot rate: 19% (pool 31%)claude-fable → claude-fable — Verified cycles: 55% (pool 62%)claude-fable → claude-fable — Edits per turn: 6.0 (pool 4.8)claude-fable → claude-fable — Turns per ship: 4.9 (pool 4.3)claude-fable → claude-fable — Hours per ship: 0.8 (pool 0.6)claude-fable → claude-fable — Dollars per ship: $34.00 (pool $15.48)claude-fable → claude-fable — Tool errors: 1.2% (pool 2.4%)claude-fable → claude-fable — Aborts: 6.3% (pool 5.0%)claude-fable → claude-fable — Time per step: 9.9 (pool 8.5)grok → grok — Ship rate: 91% (pool 93%)grok → grok — One-shot rate: 33% (pool 31%)grok → grok — Verified cycles: 32% (pool 62%)grok → grok — Edits per turn: 6.5 (pool 4.8)grok → grok — Turns per ship: 3.7 (pool 4.3)grok → grok — Hours per ship: 0.4 (pool 0.6)grok → grok — Dollars per ship: $9.55 (pool $15.48)grok → grok — Tool errors: 3.8% (pool 2.4%)grok → grok — Aborts: 4.0% (pool 5.0%)grok → grok — Time per step: 6.0 (pool 8.5)claude-opus → claude-sonnet — Ship rate: 82% (pool 93%)claude-opus → claude-sonnet — One-shot rate: 31% (pool 31%)claude-opus → claude-sonnet — Verified cycles: 42% (pool 62%)claude-opus → claude-sonnet — Edits per turn: 5.2 (pool 4.8)claude-opus → claude-sonnet — Turns per ship: 5.0 (pool 4.3)claude-opus → claude-sonnet — Hours per ship: 0.8 (pool 0.6)claude-opus → claude-sonnet — Dollars per ship: $16.06 (pool $15.48)claude-opus → claude-sonnet — Tool errors: 3.5% (pool 2.4%)claude-opus → claude-sonnet — Aborts: 9.3% (pool 5.0%)claude-opus → claude-sonnet — Time per step: 10.1 (pool 8.5)

join in

This is a first answer, not the final one. It's built from the sessions developers have shared so far. The more histories in the pool, the harder the ranking gets to argue with.

why it exists

Three questions. Three instruments.

The first two are useful and stay useful. Developers choosing a model and a workflow are missing the third answer.

question 01 · the lab

Benchmarks ask: how capable is this model?

SWE-bench, Terminal-bench, Aider, LMArena. Controlled tasks, hidden tests, preference votes — the best way to compare models on equal footing and know each model’s ceiling.

signal
pass rate · votes
blind spot
One model, a clean task, nobody steering.
task_0042.pyhidden test · unseenPASS

question 02 · the crowd

Usage rankings ask: what is the market using?

OpenCode Data, OpenRouter Ranking. Tokens, users, retention, dollars per session — the best way to see adoption and where the mix is shifting.

signal
tokens · users · $/session
blind spot
Popular is not productive — and still one model at a time.
tokens this weekModel A24T$0.03 / sessionModel B12TModel C9TModel D3T

question 03 · the recorder

Pragmatikos asks: what ships in real work?

Turns, phases, edits and errors from real sessions, correlated with what shipped. The one instrument that scores the setups developers actually run.

signal
session record × shipped outcomes
blind spot
Anything it never saw. Observational, not a benchmark.
session recordcommits · confirmationshippedcorrelated: record × commits

Capable. Popular. Effective.

Only one of them has a row for the setup you actually run.

the blind spot both miss

Same builder. Different planner. Opposite results.

The build model is identical. One planner got it to a commit in a few turns; the other burned dozens and shipped nothing. A model ranking can't see this. A pairing ranking can.

claude-fable-5.1 → grok-4.6

0 turns

grok-4.6 → grok-4.6

0 turns

turn 1planturn 4shippedturn 1turn 1turn 2turn 3unshipped
ship rate: 100% vs 90%turns to ship: 3.7 vs 3.4same builder: 1.1× the ship rate

Real sessions. Unshipped work counts against the model.

how the score works

Six tiers, ten axes, one weighted mean

Outcome

weight 30

Ship rate and one-shot rate (prompt plus approval).

Ship %One-shot %

Cost of a ship

weight 20

Turns, hours and dollars per shipped cycle.

Turns/shipHours/ship$/ship

Precision

weight 15

Tool errors and human interventions.

Tool err %Abort %

Discipline

weight 15

Share of cycles verified after the last edit.

Verified %

Efficiency

weight 10

Edits per human turn.

Edits/turn

Latency

weight 10

Median seconds per assistant step.

Time/step

Pool-relative

Every axis is a log distance to the pooled average — log odds for rates — on a fixed ×4 span, so ×2 and ÷2 sit the same distance from the line.

Evidence-weighted

Small groups shrink toward the pool with k = 10 pseudo-cycles. A two-cycle fluke nearly vanishes.

Failure lowers the rate

Unshipped cycles stay in the ship-rate denominator, so few-turn dead ends can never look good. Turns, hours and dollars are priced per shipped cycle.

two records, correlated

The session record comes first

Every turn, plan and build phase, edit, tool error and dollar is read from the agent's own record — that is where every axis is measured. Commits only decide which cycles count as shipped: a window from the previous commit, file overlap decides credit, active-but-absent sessions advised only.

session recordplan → build phasesplanbuildhuman turnsedits · same filesedits · other filescommits · confirmationC1 · file overlap → creditC2 · advised only → no creditt−6ht−3ht
edit landed in commit active, touched nothing → advised only

contribute

Add your sessions in two minutes

For OpenCode v2 — other coding agents coming soon.

Install the ocInsights plugin for OpenCode: add "@pfoundation/ocinsight" to the "plugin" list in my global opencode.json config. Then tell me to restart OpenCode, and afterwards verify the plugin loaded and the insights deck answers at http://127.0.0.1:4173/. Also report whether insight contribution is on, without changing that setting.

  1. 1

    Paste the prompt into your OpenCode agent and let it edit your config.

  2. 2

    Restart OpenCode when it tells you to, then open the deck it verifies.

  3. 3

    Preview with Contribute in the deck header — sharing is on by default.

Twenty fields per cycle — day, models, turns, edits, cost, shipping. No paths, prompts, session ids or projects. Preview in the deck's Contribute panel before anything leaves your machine.