xAI has introduced Grok 4.7, positioning it as the company's most capable model for coding and knowledge work with availability through Cursor, the Grok API, and other software tools. The model succeeds Grok 4.6 and is designed to work longer on difficult tasks, verify its output more carefully, and handle longer context windows. The announcement underscores xAI's effort to compete directly in professional and technical workloads while maintaining a focus on safety.
Model Improvements
Grok 4.7 uses a new, larger base model compared with its predecessor. It was trained with a longer reinforcement learning run on a harder mix of tasks, with more weight given to problems that can take many hours to complete. The model also natively understands the Grok Bot harness, which improves conversational tasks and general knowledge work.
Software and Engineering Benchmarks
Grok 4.7 scored 46.3 percent on CursorBench 4.0, exceeding Grok 4.6 at 40.4 percent and GPT-5.6 Sol at 41.7 percent, though Fable 5.1 Max led at 51.8 percent. In DeepSWE v1.1, the new model reached 71.0 percent in a high-effort setting, slightly behind GPT-5.6 Sol at 72.7 percent but ahead of Grok 4.6 and Fable 5.1 Max. It also recorded 64.0 percent on EEBench, a clear improvement over Grok 4.6, GPT-5.6 Sol, and Fable 5.1 Max.
Professional and Terminal Work
On the Harvey Legal Agent Benchmark, Grok 4.7 achieved 19.6 percent, compared with 15.8 percent for Grok 4.6, 2.5 percent for GPT-5.6 Sol, and 6.7 percent for Fable 5.1 Max. The model scored 1,657 on AA Briefcase v1.1, ahead of Grok 4.6 and GPT-5.6 Sol, while Fable 5.1 Max posted 1,678. For Terminal-Bench 4.0, Grok 4.7 reached 38.0 percent, above Grok 4.6 at 20.3 percent and GPT-5.6 Sol at 37.3 percent, but below Fable 5.1 Max at 57.9 percent.
Clinical Reasoning
On HealthBench Professional, Grok 4.7 scored 56.7 percent, an improvement over Grok 4.6 at 48.5 percent. The result placed the model below GPT-5.6 Sol at 60.5 percent and Fable 5.1 Max at 62.1 percent. This mixed outcome reflects the broader competitive landscape in which vendors balance performance, cost, and safety.
Safety and Cybersecurity
xAI said Grok 4.7 includes a new safeguard stack and delivered stronger internal results on refusals and jailbreak resistance than earlier models. The company reported that the model reached 62.4 percent on LatchBio's biosafety benchmark and allowed only 3.3 percent of risky dual-use prompts through on HackerBench v0.3. Select cybersecurity partners are receiving invite-only access to red-team functions for defense research.
Pricing and Availability
Grok 4.7 is priced from $2 per million input tokens and $6 per million output tokens. A faster variant is available at $4 per million input tokens and $12 per million output tokens and is set as the default. Distribution includes Cursor on desktop, web, iOS, command-line tools, software development kits, the Grok API, and third-party model routers and cloud platforms.
The launch highlights how AI providers are competing more aggressively on coding and professional task performance while facing closer scrutiny over safety controls. Grok 4.7 outperformed its predecessor across several disclosed benchmarks, although it trailed at least one rival in categories such as clinical reasoning and terminal work. The immediate test for xAI will be whether developers adopt the model in enough volume to validate its pricing and its expansion into day-to-day technical and professional work.