SpaceXAI Releases Grok 4.6 To Match OpenAI GPT-5.6 Sol Benchmark Scores
SpaceXAI launched its Grok 4.6 model on August 12, 2026, posting performance claims that match OpenAI’s GPT-5.6 Sol on primary evaluation benchmarks. The release arrives five weeks after SpaceXAI deployed Grok 4.5.

Five weeks ago, Grok 4.5 lagged behind top models from OpenAI and Anthropic by double-digit benchmark gaps. SpaceXAI closed that distance through continuous training iterations, pushing Grok 4.6 into direct competition with established industry software.
The update targets long-running software agents, interactive product development, and structured office knowledge tasks. The model reached immediate availability inside the Cursor code editor, the Grok Build platform, and the company's enterprise API.
SpaceXAI posted an overall composite score of 61 on the Artificial Analysis Intelligence Index for Grok 4.6. That number equals GPT-5.6 Sol Max and trails Anthropic’s Claude Fable 5 Max by a single point.
The official product statement from SpaceXAI describes the design goal of the update:
"Grok 4.6 turns broad product ideas into working first versions. It can establish an application's structure and visual language, build the core interactions, and refine the result through feedback."
Benchmark breakdowns show varied results across specific technical categories. On the GDPVal-AA version 2 test for knowledge work, Grok 4.6 scored an Elo rating of 1753. GPT-5.6 Sol scored 1728 on the same test, and Claude Fable 5 scored 1741.
![]() |
| Credit: SpaceXAI |
On software development benchmarks, Grok 4.6 scored 69.9% on CursorBench version 3.2. GPT-5.6 Sol scored 67.2%, and Claude Fable 5 scored 70.5% on that evaluation.
The DeepSWE version 1.1 repository coding test showed separation between the systems.
Grok 4.6 reached 65.9%, up from 54% in Grok 4.5. GPT-5.6 Sol Max leads that test at 73%, and Claude Fable 5 Max registered 70%.
On command-line software engineering evaluated under Terminal-Bench version 3.0, Grok 4.6 achieved 26%. GPT-5.6 Sol Max and Claude Fable 5 Max both recorded 34%.
SpaceXAI published details on the training process used to build the software:
"A longer supplemental training run used curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe."
SpaceXAI maintained token pricing at $2 per million input tokens and $6 per million output tokens. Prompt caching fees changed, rising from $0.30 to $0.50 per million cached tokens.
Competitor pricing sits higher for equivalent benchmark scores.
OpenAI sells GPT-5.6 Sol at $5 per million input tokens and $30 per million output tokens. Anthropic sells Claude Opus 5 at $5 per million input tokens and $25 per million output tokens.
Third-party evaluation firm Artificial Analysis tracked turn efficiency during testing runs. Grok 4.6 finished complex agent tasks in an average of 53 turns using roughly 500 million input tokens.
Claude Opus 5 Max needed an average of 103 turns and 2 billion input tokens to complete the same benchmark suite. Artificial Analysis recorded an average execution cost of $0.84 per task for Grok 4.6.
On multi-turn customer interaction testing through tau-cubed banking benchmarks, Grok 4.6 hit 50.7% accuracy. Qwen3.8 Max holds the top score in that category at 51.3%.
For long-form analytical document processing on the AA-Briefcase test, Grok 4.6 registered an Elo score of 1577, matching performance tiers seen in Claude Fable 5.
SpaceXAI trained fine-tuning data trajectories using Grok 4.5 across reasoning tasks, agent harnesses, STEM fields, and software engineering. Reinforcement learning focused on kernel optimization, web development, computer-aided design, and general coding environments.
The model retains a 500,000-token context window. Subscribers on Cursor individual and team tiers received access to Grok 4.6 upon release.
The deployment follows SpaceXAI's recent acquisition of Cursor and the release of the Grok Bot autonomous software agent.
