• 3 min read
xAI releases Grok 4.6 for coding agents and visual work
xAI releases Grok 4.6 for long-running agents, coding, and visual work, with API pricing from $2 per million input tokens.

Source: X
xAI has released Grok 4.6, a model aimed at longer-running agents, complex coding tasks, and interactive visual projects. According to X, the model is available through Cursor, Grok Build, its API, and partners including OpenRouter, Vercel, and Cloudflare.
Grok 4.6 matches GPT-5.6 Sol with a score of 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. Fable 5 Max leads the chart at 62, while Grok 4.5 scores 56. xAI says competitor results come from the developers' published system cards or public benchmark leaderboards.
The company’s published evaluation table shows Grok 4.6 ahead of Grok 4.5 on most listed tests, although it does not lead every benchmark:
- GDPVal-AA v2: 1,753, compared with 1,526 for Grok 4.5, 1,728 for GPT-5.6 Sol Max, and 1,741 for Fable 5 Max
- CursorBench v3.2: 69.9%, versus 66.7%, 67.2%, and 70.5%, respectively
- DeepSWE v1.1: 65.9%, versus 54%, 73%, and 70%
- FrontierCode v1.1 Extended: 61.3%, versus 56.6%, 60.6%, and 63.6%
- APEX-Agents: 57.5%, versus 47.1%, 56.7%, and 59.2%
- Terminal-Bench v3.0: 26%, versus 15.7%, 34.6%, and 34.1%
- APEX-SWE: 56.4%, versus 53.6%, no reported GPT-5.6 Sol score, and 58.8%
- AA-Briefcase: 1,577, versus 1,313, 1,502, and 1,574
- Harvey LAB (Vals): 15.8%, versus 12.9%, 2.5%, and 11.3%
A longer training run for agentic work
Grok 4.6 received a longer supplemental training run than Grok 4.5. xAI used curated model-generated reasoning data, advanced technical concepts, engineering data, and a revised optimizer and training recipe before supervised fine-tuning and reinforcement learning.
The company also used Grok 4.5 to regenerate supervised fine-tuning trajectories across different reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work. Model-based checks filtered problematic traces.

Recommended reading
Nvidia releases 30B Nemotron model with 1M context
For reinforcement learning, xAI trained Grok 4.6 on knowledge work, general coding, kernel optimization, web development, computer-aided design, and other domain-specific environments.
xAI says the model is particularly capable of turning a broad product idea into a working first version. Its stated workflow is to research unfamiliar subject matter, structure an application, implement core interactions, and refine the result through multiple feedback rounds. On longer tasks, the model also showed more self-testing and verification, according to the company.
The release emphasizes visual and interactive projects as well. Given a concrete product idea, xAI says Grok 4.6 can establish an application’s structure and visual language in a single pass, creating a substantial starting point for further iteration.
API pricing and safety testing
API pricing starts at $2 per million input tokens and $6 per million output tokens. A faster variant costs twice as much. xAI is also offering 2x included usage in Grok Build and Cursor during the initial launch promotion.
The company says Grok 4.6's safeguards were recalibrated for its expanded capabilities. Its testing covered legitimate uses such as vulnerability patching, engineering design, and AI research, with what xAI describes as its broadest pre-deployment evaluation suite to date, alongside post-deployment and third-party testing.
Grok 4.6 can be tried in Grok Build, integrated through the API, or accessed through the listed partner platforms. xAI did not provide a calendar date for the end of the launch usage promotion.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.


