Research investigation R0920 / benchmark analysis

Vibe Code Bench’s Score Reset: Model Regression or a Different Test?

The apparent collapse is not interpretable as a model-performance reversal from the public record alone: the earlier benchmark and VCB 1–100 change the task, scoring, budget, state, and evaluation structure.

Current public editionv1Sep 20, 2026
Verified observations
13

11 measured fields

Supported claims
7

7 material findings

Cited sources
6

5 primary or authoritative

Research score
87

Automated topic and evidence score

Interactive figureVibe Code Bench’s Score Reset: Model Regression...
CSV JSON
Data status100 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesLast updated Sep 20, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v1 / latestSep 20, 202613 records / 6 sources

    Initial public snapshot with 13 records and 6 cited sources.

Coverage note

The evidence matrix supports material changes in task structure, scoring, budget, state semantics, and test content. Public records do not provide a complete version-and-harness crosswalk or matched reruns that isolate model movement.

Dataset ID
spd:vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565
Stable URL
/research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565
Version
v1
Coverage
2026-09-20
Records
13
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
Vibe Code Bench’s Score Reset: Model Regression or a Different Test? data records
EntityMetricValueUnitObservedSourceTransform
Vibe Code Bench 1–100application arcs100arcs2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Vibe Code Benchapplication specifications100applications2026-05-13https://arxiv.org/html/2603.04601
Vibe Code Bench 1–100authored workflows11,299workflows2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Vibe Code Benchbrowser workflows964workflows2026-05-13https://arxiv.org/html/2603.04601
Gemini 3.1 Pro Previewearlier Vibe Code Bench accuracy32.03%percent2026-05-13https://arxiv.org/html/2603.04601
GPT-5.3-Codexearlier Vibe Code Bench accuracy61.77%percent2026-05-13https://arxiv.org/html/2603.04601
Vibe Code Benchgeneration wall-clock budget per application5 hourshours2026-05-13https://arxiv.org/html/2603.04601
Vibe Code Bench 1–100maximum sequential changes per arc10iterations2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Vibe Code Bench 1–100ordered change requests939requests2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Vibe Code Bench 1–100regression workflow share33.16%percent2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Vibe Code Bench 1–100shared generation budget per arc10 hourshours2026-09-20https://www.vals.ai/benchmarks/vcb-1-100
Claude Opus 5VCB 1–100 score28.53%percent2026-09-16https://www.vals.ai/benchmarks/vcb-1-100
Gemini 3.1 Pro Preview (02/26)VCB 1–100 score6.69%percent2026-09-16https://www.vals.ai/benchmarks/vcb-1-100

Measurement technique

How to read this report

  1. 01Evidence matrix: compared the earlier Vibe Code Bench paper, VCB 1–100 methodology and leaderboard, earlier scaffold and run-artifact repositories, a grading-version issue, and a secondary leaderboard mirror.
  2. 02Classified each comparison field as materially changed, partially matched, or not publicly crosswalked: task structure, test content, scoring, budget, state/reset behavior, model labels, harness, and artifacts.
  3. 03Used only records available through the September 20, 2026 UTC cutoff. Recovering saved public evidence was not a new collection or experiment.
  4. 04Did not infer a model-only effect where benchmark and system conditions changed together.
Next report / 01AI Model Economics Index All research reports
YOUR READING SPACE

Notifications