xAI Releases Grok 4.7 With Longer Reasoning at Grok 4.6 Prices

The new model is available through xAI’s API, Cursor and Grok Build, but its performance case depends on company results that do not fully line up with a published independent summary.

By 4 min read
xAI Releases Grok 4.7 With Longer Reasoning at Grok 4.6 Prices
xAI Releases Grok 4.7 With Longer Reasoning at Grok 4.6 Prices

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
xAI has released Grok 4.7, and the headline commercial change is that the standard model keeps Grok 4.6’s price: two dollars per million input tokens and six dollars per million output tokens. It’s available through the xAI API, Cursor, Grok Build, and third-party coding tools and cloud platforms. The new model is built for work that unfolds over many steps, sometimes over several hours. xAI says it uses a larger base model and a longer reinforcement-learning run, aimed at making Grok check its work, manage long context, and avoid stopping before a task is actually complete. That could help with coding and knowledge work, but longer reasoning may also consume more tokens, so the real cost is per successful job—not just per-token pricing. A faster version costs twice as much. The evidence is promising, but not cleanly comparable. xAI reports 46.3 percent on CursorBench 4.0, 71 percent on DeepSWE v1.1 at high effort, and 38 percent on Terminal-Bench 4.0. The Decoder, citing Artificial Analysis, reported 26 percent on that last benchmark under a configuration that isn’t established as equivalent. Developers should treat both the benchmark setup and their own workload as important. xAI also reports improved safeguards, including 3.3 percent risky cyber-prompt compliance and a 62.4 percent score on LatchBio’s biosafety benchmark. The key question now is whether longer runs deliver more completed work, with fewer retries, at an acceptable total cost.

Story brief

3 key points

xAI has made Grok 4.7 available through its API, Cursor, Grok Build and third-party platforms, keeping standard pricing at $2 per million input tokens and $6 per million output tokens. The model is designed for longer, multi-step coding and knowledge-work tasks, but extended reasoning may increase total spend. Evidence remains mixed: xAI reports 38.0% on Terminal-Bench 4.0, while an Artificial Analysis result cited...

  1. 01

    Grok 4.7 uses a larger base model and longer reinforcement-learning run focused on tasks lasting many hours.

  2. 02

    xAI reports 46.3% on CursorBench 4.0 and 71.0% on DeepSWE v1.1 at high effort.

  3. 03

    A faster variant costs twice the standard input and output rates.

xAI has released Grok 4.7, a successor to Grok 4.6 that it says will spend longer on difficult work, check its own output more carefully and manage extended context better. The immediate commercial hook is simple: the standard model keeps Grok 4.6’s $2-per-million input and $6-per-million output token prices.

The release follows xAI’s effort to make its models more useful for coding and knowledge work that unfolds over many steps rather than one prompt. xAI says Grok 4.7 uses a larger base model than Grok 4.6 and received a longer reinforcement-learning run, a post-training process that rewards selected behavior. The training mix was weighted toward tasks that can take many hours to finish, according to the company.

xAI’s stated target is a familiar failure mode for autonomous AI systems: stopping before the work is actually done. It says the new model is better at self-verification and long-context management, and that it understands the company’s Grok Bot harness natively. Those are company claims, not a substitute for testing in a developer’s own tools, data and task setup.

One benchmark, two published figures
38.0%xAI’s listed Terminal-Bench 4.0 result

xAI lists a 38.0% Terminal-Bench 4.0 score for Grok 4.7 xHigh in its launch table.

26%Figure reported by The Decoder

The Decoder reported that Artificial Analysis measured Grok 4.7 at 26% on Terminal-Bench 4.0, alongside higher results for GPT-6 Astra and Claude Fable 5.1.

xAI publishes a broad table of results across coding, engineering, office work, legal work and clinical reasoning. Its table puts Grok 4.7 xHigh at 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1 at high effort, and 38.0% on Terminal-Bench 4.0. It also compares the model with Grok 4.6 and several rivals, including Fable 5.1 and GPT-5.6 Sol.

But The Decoder’s account of the Artificial Analysis Intelligence Index places Grok 4.7 at 46 overall, versus 53 for Claude Fable 5.1 and GPT-6. Its Terminal-Bench figure is also lower than xAI’s published result. The two figures should not be treated as interchangeable: xAI labels its result as an xHigh setting, while the published independent summary does not establish a matching configuration.

For customers already using Grok, the launch lowers the barrier to trying the new version. Grok 4.7 is available now in Cursor, Grok Build and the Grok API, as well as through third-party coding harnesses, model routers and cloud platforms, according to xAI. The standard rate matches Grok 4.6: $2 per million input tokens and $6 per million output tokens.

That does not necessarily mean a difficult job costs the same. A model that reasons for longer can consume more tokens before it produces its final answer. xAI also offers a faster variant at twice the standard input and output prices. The meaningful operational measure is therefore not token price alone, but whether longer runs complete more tasks successfully with fewer retries.

xAI says Grok 4.7 has an entirely new safeguard stack and is its strongest model yet on refusals and resistance to jailbreaks. It reports that the model allowed 3.3% of risky dual-use cyber prompts through on HackerBench v0.3 and scored 62.4% on LatchBio’s biosafety benchmark. Those figures describe xAI’s reported evaluations; they do not independently establish how the safeguards will behave in every deployment.

The next move belongs to developers and evaluators. xAI has made Grok 4.7 easy to access and held its headline price steady while promising stronger performance on long-running work. Whether that becomes a durable advantage will depend on comparable tests that show how much extra reasoning improves completed work—and how much it costs.

Editorial analysis

Our Read

The most important part of this release may be the practical bargain xAI is trying to establish: more sustained model work without a higher token rate. That proposition is especially relevant after Grok 4.6 spread through several cloud and developer channels. But the performance story needs a cleaner basis before price alone can settle a buying decision. The next evidence to watch is reproducible, like-for-like testing on long-running coding and tool-use tasks, including the configuration behind the conflicting Terminal-Bench figures.

Sources

  1. x.aiIntroducing Grok 4.7
  2. the-decoder.comxAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6

Loading discussion...

YOUR READING SPACE

Notifications

xAI Releases Grok 4.7 With Longer Reasoning at Grok 4.6 Prices | Superpower Daily