→ Back to Home
Generative AI

xAI Releases Grok 4.7 with Strong SWE-Bench Scores and High-Throughput Modes

On September 22, 2026, xAI rolled out Grok 4.7, its latest generative model specialized for coding and complex knowledge work. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens, accompanied by a faster execution variant at double the price designed to deliver twice the generation speed. Benchmarks released by xAI highlight strong agentic engineering performance, including 71.0% on DeepSWE v1.1, 46.3% on CursorBench 4.0, and 64.0% on EEBench. This release matters directly to software engineering teams, platform architects, and DevOps leads evaluating autonomous coding agents. The competitive DeepSWE score positions Grok 4.7 as a viable backbone for automated refactoring, continuous integration triage, and developer copilots. Offering a high-speed variant at a deterministic price multiple provides enterprise teams with predictable options to trade execution latency against inference costs depending on whether the workload is synchronous (such as developer IDE interactions) or asynchronous (such as overnight test migrations). The rollout fits into the broader 2026 acceleration toward specialized, agent-ready foundation models. The industry has decisively moved beyond simple conversational generation toward multi-step problem solving in complex codebases. As frontier labs optimize their architectures for agentic tool use, long-context reasoning, and specialized domain tasks, engineering organizations are seeing a diversification of viable providers for mission-critical developer tooling beyond existing market incumbents. In practice, engineering teams should benchmark Grok 4.7 against their internal evaluation harnesses and existing coding agents. While published benchmarks show compelling capabilities on software engineering tasks, real-world utility depends on how effectively the model adheres to organization-specific linters, tool-calling schemas, and repo-level context windows. Platform teams should also assess whether the doubled cost of the high-speed tier delivers enough latency reduction in interactive agent workflows to justify the operational expense over standard inference.
#grok#xai#generative-ai#coding-agents#llm-benchmarks
Read original source