Two weeks ago we published a same-model A/B where the harness decided everything: 24 of 25 tasks solved in octomind, 19 of 25 in opencode, at double the cost. We just re-ran that experiment with DeepSeek V4 Flash, on a stricter benchmark, twice the tasks — and the model refused to make it dramatic. 45 versus 43.

On 50 real pull requests harvested from merged fixes across C++, JavaScript, PHP, Python and Rust, deepseek-v4-flash solved 45/50 in octomind and 43/50 in opencode — same model, same system prompt, sealed network, about three cents a task either way. The model is strong enough to be harness-neutral. What's left for a harness to prove is how often it finishes.

What we ran

octobench cases aren't synthetic puzzles. Each one is a real bug or feature that a maintainer actually merged, replayed at the parent commit, and graded by the project's own held-out tests plus a judge panel scoring 0–100. Fifty cases, ten per language.

Since the glm-5.2 run, we rebuilt the methodology, because the first version had a hole you could drive a truck through: the bench is built from merged PRs, and nothing stopped an agent from going and reading the merged PR. In early rounds opencode fetched the upstream fix from raw.githubusercontent.com in 13 of 49 cases and copied it. octomind's stock role was subtler but worse in spirit — it ordered the agent to websearch the upstream issue and "mirror its approach and API contracts EXACTLY." On a benchmark made of merged PRs, that's an instruction to read the answer key.

So this campaign runs clean:

  • One shared system prompt for both harnesses, distilled from octomind's role with the client-specific parts stripped. A sync script fails the build if they drift.
  • No route to the answer. Web search and web fetch disabled in both clients; GitHub unreachable from the agent's container once setup finishes.
  • Same model, same endpoint — official deepseek-v4-flash on api.deepseek.com, deepseek:deepseek-v4-flash in octomind, deepseek/deepseek-v4-flash in opencode.

One measurement note that changed our read of things: we started on a third-party token-plan endpoint and switched to the official API mid-campaign. The official endpoint is roughly 3× faster per request. A chunk of what we'd previously logged as "slow client" turned out to be slow provider.

The scoreboard

octomindopencode
solved45/50 (90.0%)43/50 (86.0%)
judge average88.5985.59
total cost≈$1.59 ($0.032/case)≈$1.53 ($0.031/case)
agent runtime10.4h (12.5m/case)8.7h (10.4m/case)
tokens (in / out / thinking)4.4M (1.9M / 1.5M / 1.0M)4.0M (2.1M / 573K / 1.3M)
cache read224.2M254.3M

Two tasks and three judge points apart. For calibration: our measured judge noise on this suite is about 1.7 points, so the gap is real but roughly double the noise floor — not the five-task blowout the glm run produced. Costs are a coin flip.

And because aggregate numbers hide the interesting stuff, here is every case. ✅/❌ is the verdict from the project's own held-out tests; j = judge score 0–100; then cost, wall time, steps, tokens (in/out/reasoning) and cache read.

Full per-case results — all 50 tasks, both harnesses
caseoctomind (deepseek-v4-flash)opencode (deepseek-v4-flash)
dev/oneshot/cpp/catch2_textflow_utf8✅ j=94.67 $0.02 7m · 35 steps · 87K tok (34K/29K/23K) · 1.4M cache✅ j=88.33 $0.03 9m · 72 steps · 90K tok (37K/13K/40K) · 3.8M cache
dev/oneshot/cpp/cpphttplib_query_verbatim✅ j=91.0 $0.02 15m · 73 steps · 60K tok (35K/16K/9K) · 2.2M cache✅ j=85.67 $0.01 21m · 65 steps · 46K tok (29K/7K/10K) · 1.9M cache
dev/oneshot/cpp/fmt_printf_oob_read✅ j=94.33 $0.01 6m · 47 steps · 47K tok (31K/11K/6K) · 1.3M cache✅ j=93.33 $0.01 4m · 49 steps · 47K tok (31K/6K/10K) · 1.2M cache
dev/oneshot/cpp/libgit2_revwalk_pathspec_root✅ j=94.33 $0.04 9m · 74 steps · 118K tok (56K/35K/27K) · 5.2M cache✅ j=94.33 $0.04 11m · 135 steps · 90K tok (44K/17K/29K) · 5.8M cache
dev/oneshot/cpp/opencv_bool_boundingrect✅ j=94.33 $0.01 3m · 33 steps · 26K tok (15K/7K/4K) · 639K cache✅ j=92.67 $0.01 4m · 53 steps · 31K tok (19K/6K/7K) · 867K cache
dev/oneshot/cpp/redis_acl_effective_keys✅ j=90.67 $0.06 25m · 95 steps · 200K tok (63K/74K/63K) · 6.3M cache✅ j=92.67 $0.05 39m · 132 steps · 133K tok (70K/21K/42K) · 9.5M cache
dev/oneshot/cpp/simdjson_bigint_boundary✅ j=95.0 $0.01 3m · 41 steps · 33K tok (17K/10K/5K) · 920K cache✅ j=93.0 $0.03 33m · 88 steps · 80K tok (57K/10K/12K) · 4.0M cache
dev/oneshot/cpp/spdlog_srcloc_lifetime✅ j=92.67 $0.02 7m · 48 steps · 52K tok (18K/21K/12K) · 1.3M cache✅ j=94.33 $0.01 8m · 48 steps · 43K tok (18K/7K/18K) · 1.1M cache
dev/oneshot/cpp/yamlcpp_binary_emit_styles❌ j=49.33 $0.02 7m · 73 steps · 68K tok (33K/22K/13K) · 2.9M cache❌ j=36.33 $0.02 5m · 64 steps · 57K tok (28K/8K/21K) · 1.6M cache
dev/oneshot/cpp/yamlcpp_octal_scalars✅ j=94.33 $0.01 5m · 29 steps · 32K tok (17K/10K/6K) · 720K cache✅ j=92.67 $0.01 4m · 30 steps · 35K tok (21K/3K/10K) · 612K cache
dev/oneshot/js/axios_stack_decoration✅ j=93.0 $0.01 6m · 25 steps · 22K tok (10K/8K/4K) · 520K cache✅ j=93.33 $0.01 15m · 30 steps · 34K tok (18K/4K/12K) · 640K cache
dev/oneshot/js/eslint_unreachable_loop_crash✅ j=89.0 $0.01 2m · 39 steps · 36K tok (25K/8K/3K) · 953K cache✅ j=94.33 $0.01 1m · 33 steps · 20K tok (14K/4K/2K) · 497K cache
dev/oneshot/js/fastify_query_method✅ j=92.33 $0.06 19m · 223 steps · 176K tok (110K/43K/22K) · 8.4M cache✅ j=93.33 $0.05 21m · 137 steps · 140K tok (99K/18K/23K) · 9.9M cache
dev/oneshot/js/hono_csp_dual_policy✅ j=94.33 $0.01 3m · 15 steps · 30K tok (17K/8K/5K) · 358K cache✅ j=94.33 $0.01 2m · 23 steps · 30K tok (17K/4K/8K) · 429K cache
dev/oneshot/js/pino_single_target_level✅ j=95.0 $0.01 4m · 28 steps · 31K tok (21K/7K/3K) · 729K cache✅ j=93.67 $0.02 9m · 76 steps · 65K tok (33K/11K/20K) · 3.0M cache
dev/oneshot/js/pinopretty_strip_controls❌ j=41.67 $0.05 17m · 143 steps · 159K tok (63K/56K/40K) · 4.3M cache❌ j=39.0 $0.04 10m · 134 steps · 113K tok (62K/20K/30K) · 7.7M cache
dev/oneshot/js/react_hidden_hydration_hang❌ j=41.67 $0.32 271m · 1322 steps · 419K tok (107K/183K/129K) · 79.3M cache❌ j=36.67 $0.32 63m · 394 steps · 441K tok (187K/72K/182K) · 79.4M cache
dev/oneshot/js/undici_async_mock_reply✅ j=95.67 $0.05 11m · 87 steps · 134K tok (52K/47K/35K) · 6.1M cache✅ j=94.33 $0.05 14m · 108 steps · 117K tok (47K/21K/49K) · 7.9M cache
dev/oneshot/js/vite_hmr_restart_stale✅ j=86.33 $0.02 6m · 58 steps · 82K tok (37K/26K/19K) · 2.1M cache✅ j=86.67 $0.05 17m · 136 steps · 118K tok (52K/21K/45K) · 8.8M cache
dev/oneshot/js/webpack_lazy_backend_shutdown✅ j=93.67 $0.04 14m · 81 steps · 122K tok (39K/47K/35K) · 5.2M cache✅ j=94.67 $0.04 12m · 74 steps · 111K tok (52K/11K/48K) · 4.6M cache
dev/oneshot/php/cakephp_rate_limit_ip✅ j=92.67 $0.01 4m · 41 steps · 52K tok (24K/18K/10K) · 1.2M cache✅ j=92.67 $0.02 5m · 85 steps · 64K tok (35K/11K/18K) · 3.0M cache
dev/oneshot/php/carbon_period_end_sync✅ j=95.33 $0.15 38m · 229 steps · 376K tok (82K/165K/130K) · 20.9M cache✅ j=93.33 $0.04 12m · 111 steps · 111K tok (49K/19K/43K) · 6.7M cache
dev/oneshot/php/commonmark_fence_tabs❌ j=38.33 $0.02 5m · 60 steps · 74K tok (31K/24K/19K) · 2.1M cache❌ j=41.67 $0.05 11m · 125 steps · 117K tok (60K/15K/42K) · 7.9M cache
dev/oneshot/php/composer_policy_source✅ j=94.67 $0.02 5m · 95 steps · 81K tok (51K/19K/11K) · 2.7M cache✅ j=94.33 $0.02 8m · 84 steps · 68K tok (50K/10K/9K) · 2.9M cache
dev/oneshot/php/dbal_sqlite_alter_preserves_constraints✅ j=93.0 $0.05 8m · 100 steps · 143K tok (85K/36K/23K) · 6.8M cache✅ j=89.33 $0.04 7m · 101 steps · 114K tok (72K/17K/26K) · 6.2M cache
dev/oneshot/php/flysystem_mimetype_null✅ j=95.0 $0.00 1m · 18 steps · 17K tok (13K/4K/1K) · 289K cache✅ j=94.0 $0.01 4m · 46 steps · 30K tok (20K/4K/6K) · 720K cache
dev/oneshot/php/guzzle_cookie_prefixes✅ j=94.0 $0.02 4m · 42 steps · 66K tok (33K/19K/13K) · 1.4M cache❌ j=45.0 $0.02 17m · 63 steps · 86K tok (49K/8K/28K) · 2.7M cache
dev/oneshot/php/monolog_max_trace_length✅ j=94.67 $0.01 2m · 24 steps · 35K tok (20K/10K/6K) · 568K cache❌ j=36.0 $0.01 3m · 34 steps · 33K tok (18K/5K/11K) · 657K cache
dev/oneshot/php/symfony_jsonpath_singular✅ j=94.33 $0.07 11m · 80 steps · 229K tok (102K/69K/58K) · 8.2M cache✅ j=91.33 $0.04 10m · 59 steps · 123K tok (56K/11K/57K) · 4.3M cache
dev/oneshot/php/twig_stringable_array_key✅ j=94.33 $0.02 5m · 66 steps · 76K tok (37K/23K/16K) · 2.3M cache✅ j=94.33 $0.02 6m · 90 steps · 70K tok (42K/11K/17K) · 3.3M cache
dev/oneshot/python/anyio_cancel_spin✅ j=94.0 $0.01 3m · 34 steps · 43K tok (21K/13K/9K) · 854K cache✅ j=95.33 $0.01 6m · 41 steps · 41K tok (23K/5K/13K) · 897K cache
dev/oneshot/python/anyio_tls_idna2008✅ j=91.0 $0.03 7m · 75 steps · 111K tok (58K/31K/23K) · 2.6M cache✅ j=93.67 $0.02 5m · 51 steps · 54K tok (26K/6K/22K) · 1.3M cache
dev/oneshot/python/celery_flush_concurrent_append✅ j=95.0 $0.01 1m · 18 steps · 23K tok (14K/5K/3K) · 318K cache✅ j=94.33 $0.01 3m · 25 steps · 37K tok (16K/4K/17K) · 570K cache
dev/oneshot/python/click_powershell_completion✅ j=95.67 $0.02 7m · 54 steps · 79K tok (42K/23K/15K) · 1.9M cache✅ j=93.33 $0.02 6m · 51 steps · 69K tok (31K/9K/30K) · 2.5M cache
dev/oneshot/python/fastapi_sse_line_splitting✅ j=94.67 $0.01 3m · 32 steps · 39K tok (19K/12K/8K) · 757K cache✅ j=94.33 $0.01 2m · 27 steps · 25K tok (13K/4K/8K) · 490K cache
dev/oneshot/python/poetry_show_outdated_explicit_source✅ j=91.0 $0.03 8m · 107 steps · 91K tok (48K/27K/16K) · 4.9M cache✅ j=89.0 $0.04 9m · 147 steps · 104K tok (62K/14K/28K) · 6.9M cache
dev/oneshot/python/pydantic_pipeline_constraints✅ j=94.33 $0.01 3m · 28 steps · 58K tok (28K/18K/12K) · 911K cache✅ j=93.33 $0.01 4m · 45 steps · 46K tok (22K/7K/17K) · 1.2M cache
dev/oneshot/python/pytest_no_summary_hook_scope✅ j=95.33 $0.01 2m · 32 steps · 22K tok (12K/7K/4K) · 435K cache✅ j=94.33 $0.01 3m · 62 steps · 31K tok (15K/6K/10K) · 1.1M cache
dev/oneshot/python/tornado_zero_timeout✅ j=96.67 $0.00 1m · 18 steps · 6K tok (4K/2K/0K) · 223K cache✅ j=94.0 $0.00 3m · 22 steps · 21K tok (13K/3K/5K) · 309K cache
dev/oneshot/python/werkzeug_float_url_notation✅ j=92.67 $0.01 9m · 39 steps · 46K tok (21K/15K/10K) · 1.2M cache✅ j=92.33 $0.01 30m · 59 steps · 35K tok (16K/8K/11K) · 1.3M cache
dev/oneshot/rust/bytes_truncate_release✅ j=95.33 $0.01 2m · 25 steps · 48K tok (22K/15K/11K) · 610K cache✅ j=94.33 $0.01 2m · 34 steps · 28K tok (17K/4K/7K) · 598K cache
dev/oneshot/rust/chrono_iter_reverse✅ j=94.33 $0.03 6m · 34 steps · 92K tok (25K/37K/31K) · 1.1M cache✅ j=93.33 $0.03 11m · 80 steps · 97K tok (29K/19K/49K) · 3.7M cache
dev/oneshot/rust/image_ico_mask_transparency✅ j=96.0 $0.00 1m · 10 steps · 11K tok (9K/2K/1K) · 140K cache✅ j=94.33 $0.01 2m · 42 steps · 33K tok (23K/5K/5K) · 879K cache
dev/oneshot/rust/rayon_par_array_windows✅ j=94.33 $0.01 2m · 23 steps · 27K tok (18K/6K/3K) · 452K cache✅ j=95.0 $0.01 2m · 36 steps · 42K tok (31K/5K/6K) · 1.0M cache
dev/oneshot/rust/ripgrep_maxdepth_ignore_skip✅ j=92.67 $0.05 9m · 77 steps · 158K tok (78K/45K/34K) · 5.9M cache✅ j=94.33 $0.04 8m · 100 steps · 116K tok (73K/14K/29K) · 6.8M cache
dev/oneshot/rust/rustls_misplaced_extensions❌ j=41.0 $0.11 18m · 164 steps · 250K tok (115K/79K/56K) · 18.9M cache❌ j=46.0 $0.11 18m · 145 steps · 239K tok (141K/26K/72K) · 22.3M cache
dev/oneshot/rust/serdejson_enum_key_string✅ j=94.67 $0.02 5m · 66 steps · 70K tok (28K/24K/17K) · 2.2M cache✅ j=94.33 $0.02 5m · 66 steps · 50K tok (23K/8K/19K) · 2.2M cache
dev/oneshot/rust/tokio_alt_timer_cancel_race✅ j=93.67 $0.02 6m · 46 steps · 55K tok (28K/16K/11K) · 1.3M cache✅ j=90.0 $0.03 7m · 71 steps · 86K tok (52K/10K/24K) · 3.5M cache
dev/oneshot/rust/toml_datetime_value✅ j=94.0 $0.01 3m · 42 steps · 49K tok (30K/12K/7K) · 1.4M cache✅ j=94.33 $0.02 5m · 72 steps · 73K tok (46K/9K/18K) · 3.0M cache
dev/oneshot/rust/uuid_parse_panic✅ j=93.33 $0.01 4m · 23 steps · 45K tok (20K/14K/11K) · 619K cache✅ j=93.67 $0.02 7m · 45 steps · 72K tok (23K/10K/40K) · 2.0M cache
  • octomind: 45/50 PASS (90.0%) · 5 FAIL · jAvg 88.59 · ≈$1.59 total ($0.032/case) · 10.4h agent (12.5m/case) · 4.4M tok (88K/case; 1.9M in / 1.5M out / 1.0M reas) · 224.2M cache read (4.5M/case)
  • opencode: 43/50 PASS (86.0%) · 7 FAIL · jAvg 85.59 · ≈$1.53 total ($0.031/case) · 8.7h agent (10.4m/case) · 4.0M tok (79K/case; 2.1M in / 573K out / 1.3M reas) · 254.3M cache read (5.1M/case)

Every case has full artifacts — agent traces, validation output, judge verdicts — in the per-case table pinned at the exact commit of this run. Reproduction instructions are in the same repo.

The new Flash in numbers

The official V4 Flash went live July 31, and the release is the real thing: DeepSeek's own GA numbers have it beating the V4 Pro preview on agentic benchmarks — Terminal Bench 2.1 at 82.7, DeepSWE at 54.4, Toolathlon verified at 70.3 — from a Mixture-of-Experts with 284B total and 13B activated parameters and a 1M-token context window (specs via its OpenRouter listing).

What that translates to on real-PR work: twelve minutes and three cents per landed fix, a ~94 judge average on cases it passes, in either harness. No babysitting, no mystery timeouts. For the routine 90% of maintainer work, this model at flash pricing is simply a solved problem.

And here's the interesting part — it's harness-neutral. The glm-5.2 run showed a five-task gap between the same model in two harnesses. V4 Flash lands two tasks apart. Two. Part of that is the model carrying more of its own discipline: fewer wasted reads, fewer early victory laps, less need for a supervisor to keep it honest. Part of it is that this benchmark is stricter with both clients — we sealed the answer key, and that cut against octomind's stock role as much as anyone's web access. We made the test harder for ourselves and the gap narrowed anyway. Both things are true.

One ceiling, five walls

The five cases octomind fails are the same five opencode fails: yaml-cpp's binary emit styles, pino-pretty's control stripping, commonmark's fence tabs, rustls's misplaced extensions, and React's hidden hydration hang.

Identical walls across harnesses is what a model blind spot looks like — not a harness artifact. These are the composition-heavy fixes where the model keeps missing a sibling path or a per-message rule no matter how it's prompted. Four instrumented reruns of pino-pretty all missed the same sibling. That's the model, not the steering.

The honest read of the 86–90% band: the last 10% is where model quality ends and everything else begins.

What the harness still decides is the finish

octomind has zero unique failures. opencode's two extra losses are both finishes that never landed: Guzzle's cookie prefixes, where opencode spent 17 minutes and failed while octomind passed in 4; Monolog's trace length, where opencode bailed in 3 minutes with a judge score of 36 while octomind landed a 94.67 in 2.

The React case is the same story at maximum volume. It's a hang-by-design bug — every wrong attempt blocks instead of failing — on a giant repo. opencode stopped at 63 minutes and 394 steps, scoring 36.67. octomind ground for 271 minutes and 1,322 steps — for the same $0.32 — to a fix that passes 66 of 67 hidden tests. We record it as a failure, because it is one, but the judge scored the near-miss at 41.67. (One run died at minute 213 on a provider quota error and was resumed from a restored session to finish at all.)

That's finalization: not better luck, more refusal to stop early. Strip the React outlier and octomind also averages about 7 minutes a case to opencode's 9.

The tax is real, though. octomind wrote 2.6× the output tokens — 1.5M versus 573K. It talks more. And on the React-sized repos its fine-grained tool loop against a fat per-call context is a genuine speed ceiling; coarser steps and context slimming are next on our own list. A harness that finishes everything will still cost you in words.

The model sets the ceiling. The harness sets the floor.

After the glm run we wrote that model quality is table stakes and the harness is where the leverage is. V4 Flash tightens that into something more precise: a model this good raises the ceiling for everyone — pick the model you like, the ceiling barely moves between harnesses. What the harness decides is how often you actually reach it, and whether the last five percent lands or gets declared done at minute 63.

glm-5.3 just dropped. Same bench, next question.

Yesterday Zhipu released GLM-5.3: same base model as 5.2, with the gains coming from post-training alone. Zhipu claims a 50% coding improvement over 5.2 and calls it the strongest open-weights coding model. It's live through their coding plan now. Open weights are promised in about two weeks, behind a safety review.

This matters to us more than most releases, because glm-5.2 is the model that started this story. In the first run, glm-5.2 in octomind beat Claude Code running claude-opus-5 (24/25 solved against 23/25, $63 against $82, 3.6 hours against 6.7) — while paying full freight: that run went through an endpoint with no prompt caching, so every turn re-bought its whole context at list price.

So the next campaign writes itself. glm-5.3 versus glm-5.2 on the same sealed 50-case bench, in both harnesses, answer key locked. A claimed 50% post-training jump is exactly the kind of number vendor benchmarks love and real pull requests interrogate. And there's a sharper question underneath: V4 Flash just showed that a strong model narrows the harness gap to two tasks. If 5.3 really is that much better than a model already beating Opus, does the floor rise with the ceiling — or does the last 10% stay exactly right where it was? Worth finding out.

Reproduction steps and the raw per-case artifacts are pinned at the commit this post describes. GLM-5.3 versus 5.2 is next on the calendar, with Claude Code and Codex columns on the same sealed bench right behind it.