Grok 4.7 Review: Same Price, Trails Fable 5.1 and GPT-6
xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing on a larger base model, but Artificial Analysis scores it 46 versus 53 for Fable 5.1 and GPT-6.
xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing on a larger base model, but Artificial Analysis scores it 46 versus 53 for Fable 5.1 and GPT-6.
Introduction
xAI released Grok 4.7 on September 21, 2026, again publishing the announcement under the "SpaceXAI" brand it has used for recent launches. According to the official page, Grok 4.7 is "our most capable model for coding and knowledge work," built to work longer on difficult tasks, check its own work more carefully, and ship with what xAI calls its best-calibrated safeguards to date. The company says the model is served "at the same price and speed as Grok 4.6" — notable given that xAI also describes Grok 4.7 as running on a new, larger base model than its predecessor. That combination of unchanged pricing and a bigger underlying model is the central story of this release, alongside a wide gap between xAI's own benchmark table and the independent numbers Artificial Analysis published and The Decoder reported.
Feature Overview
xAI attributes Grok 4.7's gains to three changes layered on the new base model: a longer reinforcement-learning run "on a harder mix of tasks, weighted toward problems that take many hours to complete"; stronger self-verification, with the model checking its own work more carefully; and native training to operate inside the Grok Bot harness. On CursorBench 4.0, xAI says Grok 4.7 is "at the frontier in price-performance" — a claim specifically about price-performance, not the raw benchmark score.
Safety is framed as a headline feature. xAI describes an "entirely new safeguard stack" and, per its own testing, calls Grok 4.7 the strongest model it has tested on refusals and jailbreak resistance. The company also reports the model tops LatchBio's biosafety benchmark at 62.4%, and on HackerBench v0.3 — xAI's own benchmark for risky or malicious cyber tasks — only 3.3% of risky dual-use prompts get through. xAI says it has also started giving select cybersecurity partners invite-only access to the model's red-team capabilities for defense research.
On pricing, xAI's documentation lists standard global rates of $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for prompts under 200,000 tokens — identical to Grok 4.6's rates. Once a request's prompt reaches 200,000 tokens, the entire request bills at the higher tier: $4.00/$1.00/$12.00 per million. A separate "Grok 4.7 Fast" variant runs the same model on faster infrastructure at double the standard rate, but per xAI's documentation it is available only through Cursor and Grok Build, not the public API and not through Grok Build's free tier. The US regional endpoint runs 1.1x the global rate.
Usability Analysis
Grok 4.7 is available the same day as the announcement in Cursor and Grok Build, with xAI inviting developers to "try it in Grok Build for free." It also ships to the Grok API, third-party coding harnesses, model routers, and cloud platforms, continuing xAI's pattern of broad same-day distribution. Docs.x.ai lists the model id as grok-4.7, a 500,000-token context window, configurable reasoning effort, and a May 2026 knowledge cutoff. One integration detail worth noting: on the Responses API, Grok 4.7 always returns reasoning.encrypted_content, even when a request's include parameter doesn't list it — a behavior developers building custom tooling around reasoning traces will need to account for.
Pros and Cons
Pros:
- Pricing remains at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, despite xAI describing Grok 4.7 as built on a new, larger base model.
- A longer reinforcement-learning run and native training for the Grok Bot harness target longer, harder multi-step tasks, according to xAI.
- xAI reports its strongest self-tested safety results yet, including a 3.3% pass-through rate for risky dual-use prompts on its own HackerBench v0.3 benchmark.
- In xAI's own four-model comparison table, Grok 4.7 posts the highest score on EEBench (64.0%) and the Harvey Legal Agent Benchmark (19.6%).
- Broad same-day availability spans Cursor, Grok Build, the xAI API, third-party harnesses, and cloud platforms.
Cons:
- Independent benchmarking from Artificial Analysis places Grok 4.7 at 46 on its Intelligence Index, well below the 53 each that Claude Fable 5.1 and GPT-6 score, per The Decoder.
- The parameter count behind Grok 4.7 is not disclosed on either xAI's announcement page or its documentation.
- The Terminal-Bench 4.0 score xAI reports (38.0%) is well above the 26% Artificial Analysis measured independently, and the two figures are not reconciled.
- xAI's own comparison table benchmarks Grok 4.7 against GPT-5.6 Sol Max rather than OpenAI's newer GPT-6 Astra flagship.
Comparison
xAI's own benchmark table compares Grok 4.7 xHigh against Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max:
| Metric | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input $/M | $2 | $2 | $4 | $10 |
| Output $/M | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench (electrical eng.) | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
*The DeepSWE v1.1 figure for Grok 4.7 is a high-effort score; every other Grok 4.7 result in the table is xHigh.
Reading the table row by row, Grok 4.7 posts the top score of the four only on EEBench and the Harvey Legal Agent Benchmark. It trails Fable 5.1 Max on CursorBench, AA Briefcase, Terminal-Bench (38.0% vs 57.9%), and HealthBench Professional, and trails GPT-5.6 Sol Max on DeepSWE (71.0% vs 72.7%) and HealthBench Professional (56.7% vs 60.5%). xAI's text commentary on a separate chart — which compares Grok 4.7 with Grok 4.6, Fable 5.1, and GPT-6 Astra on GDPval and AA Briefcase — states only that Grok 4.7 "improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models"; the underlying numbers for that chart are not given as text.
Artificial Analysis' independent measurement, cited by The Decoder, tells a different story on two fronts. Its Intelligence Index v4.3.2 places Grok 4.7 at 46, which The Decoder characterizes as "mid-pack," against 53 each for Claude Fable 5.1 and GPT-6. Its own Terminal-Bench 4.0 run measured Grok 4.7 at 26% — well below the 38.0% in xAI's table, and also below GPT-6 Astra's 60%, Claude Fable 5.1's 55%, and DeepSeek V4.1 Flash's 27% in that same independent run. The two Terminal-Bench figures come from separate evaluations — xAI's own and Artificial Analysis' — and neither source explains the difference, so they should not be merged into a single number.
Outlook
Holding pricing flat on what xAI describes as a larger base model keeps Grok positioned as the low-cost option: at $2/$6 per million tokens, its list price is half of GPT-5.6 Sol Max's input rate and under a third of its output rate, and a fifth of Fable 5.1 Max's input rate. Whether that bet pays off for developers depends heavily on which numbers they trust: xAI's own table shows a model that leads on a couple of specialized benchmarks and undercuts GPT-5.6 Sol Max and Fable 5.1 Max sharply on price, while Artificial Analysis' independent testing places it behind Claude Fable 5.1 and GPT-6 on its Intelligence Index, and behind GPT-6 Astra and Claude Fable 5.1 on its own Terminal-Bench 4.0 run. Buyers evaluating Grok 4.7 for coding-heavy or safety-sensitive workloads should weigh both sets of numbers rather than either alone.
Conclusion
Grok 4.7 keeps xAI's pricing unchanged from Grok 4.6 while adding a larger base model, longer reinforcement learning, and a new safeguard stack the company positions as its strongest yet. It leads xAI's own comparison table on two specialized benchmarks and undercuts GPT-5.6 Sol Max and Fable 5.1 Max on price, but Artificial Analysis' independent testing places it mid-pack on its Intelligence Index, and its 26% Terminal-Bench 4.0 measurement sits well below the 38.0% xAI reports. Teams already using Grok Build or Cursor, and cost-sensitive developers comfortable weighing vendor and independent benchmarks separately, are best positioned to evaluate it.
Editor's Verdict
Grok 4.7 Review: Same Price, Trails Fable 5.1 and GPT-6 is a workable proposition that fills a clear gap, even if it doesn't fundamentally change the landscape.
The strongest case for paying attention: pricing remains at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, despite xAI describing Grok 4.7 as built on a new, larger base model. That alone raises the bar for what readers should expect in this space. Reinforcing that, a longer reinforcement-learning run and native training for the Grok Bot harness target longer, harder multi-step tasks, according to xAI — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the 500,000-token context window matches Grok 4.6's, and once a prompt reaches 200,000 tokens every token in that request bills at double the standard rate. On the other side of the ledger, one constraint is real rather than a marketing footnote: independent benchmarking from Artificial Analysis places Grok 4.7 at 46 on its Intelligence Index, well below the 53 each that Claude Fable 5.1 and GPT-6 score, per The Decoder. It should factor into any serious decision. Layered on top of that, the parameter count behind Grok 4.7 is not disclosed on either xAI's announcement page or its documentation — which narrows the set of teams for whom this is an obvious yes.
For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, the smart move is to track its trajectory and revisit once the rough edges are filed down. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- Pricing remains at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, despite xAI describing Grok 4.7 as built on a new, larger base model.
- A longer reinforcement-learning run and native training for the Grok Bot harness target longer, harder multi-step tasks, according to xAI.
- xAI reports its strongest self-tested safety results yet, including a 3.3% pass-through rate for risky dual-use prompts on its own HackerBench v0.3 benchmark.
- In xAI's own four-model comparison table, Grok 4.7 posts the highest score on EEBench (64.0%) and the Harvey Legal Agent Benchmark (19.6%).
- Broad same-day availability spans Cursor, Grok Build, the xAI API, third-party harnesses, and cloud platforms.
Cons
- Independent benchmarking from Artificial Analysis places Grok 4.7 at 46 on its Intelligence Index, well below the 53 each that Claude Fable 5.1 and GPT-6 score, per The Decoder.
- The parameter count behind Grok 4.7 is not disclosed on either xAI's announcement page or its documentation.
- The Terminal-Bench 4.0 score xAI reports (38.0%) is well above the 26% Artificial Analysis measured independently, and the two figures are not reconciled.
- xAI's own comparison table benchmarks Grok 4.7 against GPT-5.6 Sol Max rather than OpenAI's newer GPT-6 Astra flagship.
References
Comments0
Key Features
1. New, larger base model with a longer reinforcement-learning run targeting harder, longer tasks. 2. Same $2/$6 per-million-token pricing as Grok 4.6. 3. Entirely new safeguard stack; xAI's own testing calls it the strongest yet on refusals and jailbreak resistance. 4. Available same-day in Cursor, Grok Build, the xAI API, and third-party platforms. 5. 500K token context; Responses API always returns encrypted reasoning content.
Key Insights
- The 500,000-token context window matches Grok 4.6's, and once a prompt reaches 200,000 tokens every token in that request bills at double the standard rate.
- xAI credits Grok 4.7's gains to a longer reinforcement-learning run weighted toward tasks that take many hours to complete, plus native training for the Grok Bot harness.
- On CursorBench 4.0, xAI's own table shows Grok 4.7 xHigh at 46.3%, ahead of Grok 4.6 High (40.4%) and GPT-5.6 Sol Max (41.7%), but behind Fable 5.1 Max (51.8%).
- On DeepSWE v1.1, Grok 4.7's high-effort score of 71.0% trails GPT-5.6 Sol Max's 72.7% but edges past Fable 5.1 Max's 70.0%.
- Artificial Analysis measured Grok 4.7 at 26% on Terminal-Bench 4.0, versus the 38.0% xAI reports in its own table — the figures come from separate evaluations and neither source reconciles them.
- Artificial Analysis' Intelligence Index v4.3.2 places Grok 4.7 at 46, describing it as 'mid-pack,' compared with 53 each for Claude Fable 5.1 and GPT-6, per The Decoder.
- xAI's comparison table benchmarks Grok 4.7 against GPT-5.6 Sol Max, not OpenAI's newer GPT-6 Astra flagship.
- On the Responses API, Grok 4.7 always returns encrypted reasoning content, and a 'Fast' variant at double the per-token price is available only through Cursor and Grok Build, not the public API.
Was this review helpful?
Share
Related AI Reviews
Meta Expands Muse Agent With Video Chat, Glasses, Charm
Meta expanded its Muse AI agent at Connect 2026 with video chat, smart glasses, Mac control, and a Muse Charm wearable shipping in December.
Qwen3.8-Omni-Flash Undercuts Gemini Flash on Price
Alibaba's new omnimodal model prices at $0.15/$0.47 per million tokens, far below Gemini 3.8 Flash, with benchmarks that are close but mixed.
TypeSafe AI Launches Jev, a New System One Model
Ex-OpenAI researcher Diogo Almeida's TypeSafe AI released Jev in early access, a structured-decision model priced at $0.042 per million input tokens.
PrismML's Bonsai 2 27B Shrinks Qwen3.8 27B by More Than 9x
PrismML's Bonsai 2 27B compresses Qwen3.8 27B to 5.9GB, over 9x smaller, while retaining 98.2% of its full-precision benchmark performance.
