AI21 Proposes a Task-Based Alternative to Simply Scaling Up AI Agents
Company-reported experiments favor checking candidates for factual search, merging reports for deep research, and dividing coding work by what can be verified.
Listen to this story
The audio brief
Story brief
3 key pointsAI21 Labs’ October 6 research lays out a task-dependent design rule for agents: verify candidate answers when outcomes are checkable, but merge diverse attempts when completeness is hard to assess. The approach shifts effort away from simply adding model capacity or retries and toward the stage most likely to bottleneck a task. AI21 applies the distinction to search, deep research, retrieval, and coding, while noting that some evidence comes from small verifier samples and that research benchmarks use expert-authored rubrics that may not fit every request.
- 01
On SWE-Bench Pro, AI21 reports an 80.8% resolve rate for a pipeline using parallel MiniMax-M3 exploration, GPT-5.2 context distillation, and one frontier-model patch-writing pass.
- 02
Adding candidate patches to context search cut coding issues where agents missed every relevant file from 9.3% to 2.7%; resolve rate rose 3.2 percentage points.
- 03
For short-answer search, AI21’s oracle test found correct answers often appeared among attempts; an independent verifier can reject bad candidates before voting.
A stronger model or more attempts can improve an AI agent, but AI21 Labs argues that scaling can also leave architectural waste untouched. In research published October 6, the company proposes a different starting point: determine whether the task’s outcome can be checked, then spend on either verifying candidates or combining varied attempts.
Search needs a check, not just a vote
For a search question seeking one person or another short answer, explicit conditions make verification relatively narrow. AI21 first tested selection using known correct answers—an “oracle” experiment unavailable to a working agent. Correct answers frequently appeared among multiple attempts, pointing to selection rather than generation as the bottleneck.
Majority voting cannot rescue a correct answer when a wrong one dominates the pool. An independent verifier instead researches each candidate and rejects failures before voting. AI21’s earlier verifier experiments provide useful limits: both benchmarks used 100-question samples, training labels inherited automated-grader noise, and verification added latency.
Research coverage resists a single winner
Deep research presents the opposite problem. Individual claims can be checked, but knowing whether a report covers everything relevant requires doing the research again, AI21 argues. Even an oracle choosing the best single report fell below the previous state of the art. Merging reports offered a higher ceiling.
AI21 says merging reports from seven low-ranked agents produced the top result on DeepResearch Bench II. The benchmark’s original paper describes 132 tasks across 22 domains, evaluated against 9,430 yes-or-no criteria covering information recall, analysis and presentation. Those criteria come from expert-written investigative articles, not a ready-made checklist for every real-world research request.
The same coverage problem appears when indexing documents for retrieval. A question may need one line or several sections. AI21 recommends storing multiple chunk sizes, accepting extra storage and retrieval work rather than committing before the question arrives.
Coding changes the answer midway
Coding combines both cases. Tests provide an end-of-run check, but finding all relevant files during exploration remains a coverage problem. AI21 says feeding candidate patches into context search, alongside the issue description, cut the share of issues where the agent missed every relevant file from 9.3% to 2.7%. Resolve rate rose 3.2 percentage points.
Its pipeline assigns open-source MiniMax-M3 models to explore the repository in parallel. GPT-5.2 distills their attempts into the context needed for a fix; a frontier model then writes the patch once. The division confines the strongest model’s patch-writing work to the stage with a test-suite check, rather than using it throughout exploration.
Reader comments
Newest comments first. Replies stay oldest first.