Research investigation R0920 / benchmark analysis
Vibe Code Bench’s Score Reset: Model Regression or a Different Test?
The apparent collapse is not interpretable as a model-performance reversal from the public record alone: the earlier benchmark and VCB 1–100 change the task, scoring, budget, state, and evaluation structure.
Snapshot only. There is not enough history to claim a trend yet.
Version ledger
Frozen public editions
Each edition preserves the records, method, sources, and downloads available at publication time.
The evidence matrix supports material changes in task structure, scoring, budget, state semantics, and test content. Public records do not provide a complete version-and-harness crosswalk or matched reruns that isolate model movement.
- Dataset ID
- spd:vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565
- Stable URL
- /research/vibe-code-bench-s-score-reset-model-regression-or-a-different-test-ecb46565
- Version
- v1
- Coverage
- 2026-09-20
- Records
- 13
- Fields
- 7
- Updated
Measurement technique
How to read this report
- 01Evidence matrix: compared the earlier Vibe Code Bench paper, VCB 1–100 methodology and leaderboard, earlier scaffold and run-artifact repositories, a grading-version issue, and a secondary leaderboard mirror.
- 02Classified each comparison field as materially changed, partially matched, or not publicly crosswalked: task structure, test content, scoring, budget, state/reset behavior, model labels, harness, and artifacts.
- 03Used only records available through the September 20, 2026 UTC cutoff. Recovering saved public evidence was not a new collection or experiment.
- 04Did not infer a model-only effect where benchmark and system conditions changed together.
Sources
Evidence
4 publishers supporting 13 records. Expand a publisher to inspect its cited pages.