The long term is becoming the domain of Grok 4.7.
SpaceXAI positionne Grok 4.7 sur le code et les missions professionnelles de longue durée, avec 500 000 tokens de contexte, un raisonnement réglable et des tarifs inchangés.
An AI assignment no longer has to end after a few lines of text or a single code fix. Development systems are increasingly expected to explore entire repositories, use several tools, run tests, recover from failures, and keep working toward a usable result. Elsewhere, the same systems may spend hours preparing a presentation, reviewing a legal file, or assembling a financial document.
That is the territory SpaceXAI has chosen for Grok 4.7. The model uses a larger foundation than Grok 4.6 and underwent a longer reinforcement-learning process built around harder tasks, including assignments designed to take several hours. The company says it has improved both self-verification and the model’s ability to manage long contexts. Built to stay on the job
The main change is not a new interface or an additional media generator. Grok 4.7 is primarily aimed at agentic tasks, where a model must plan a sequence of actions, call tools, inspect the results, and adjust its approach.
This applies directly to software development. Instead of answering an isolated technical question, the system can inspect a project, edit multiple files, run tests, review errors, and continue until it produces a workable result. Similar behavior can support document creation, presentations, and research assignments involving several sources.
SpaceXAI also says the model was trained to understand the Grok Bot harness natively. That may improve conversations and tool use within the company’s own environment. It also complicates comparisons, since an agentic model’s performance depends on its surrounding software, available tools, and operating instructions. Clear gains without an across-the-board lead
The published results show a meaningful improvement over Grok 4.6. On CursorBench 4.0, which focuses on longer-running coding assignments, Grok 4.7 scores 46.3%, compared with 40.4% for its predecessor. It also finishes ahead of GPT-5.6 Sol at 41.7%, but behind Fable 5.1 at 51.8%.
The ranking changes from one test to another. Grok 4.7 reaches 71% on DeepSWE v1.1, slightly below GPT-5.6 Sol at 72.7% and just above Fable 5.1 at 70%. On Terminal-Bench 4.0, its 38% result narrowly exceeds GPT-5.6 Sol’s 37.3%, while Fable 5.1 remains well ahead at 57.9%.
The new model leads SpaceXAI’s comparison table on EEBench, an electrical-engineering test, with 64%. It also posts the highest listed result on the Harvey Legal Agent Benchmark at 19.6%. Its 56.7% HealthBench Professional score, however, trails GPT-5.6 Sol at 60.5% and Fable 5.1 at 62.1%.
Those mixed outcomes make it difficult to describe Grok 4.7 as universally superior. Its position depends on the field, execution environment, and selected reasoning level. SpaceXAI evaluates Grok 4.7 in xHigh mode in its main table, while competitors use their own maximum settings. These configurations do not necessarily produce the same latency or token consumption.
Cost-per-task claims deserve similar caution. The final bill depends on context length, the number of steps, tool calls, failed attempts, and output volume. A lower per-token rate does not guarantee that every completed assignment will cost less. Moving beyond software development
SpaceXAI is also emphasizing document and presentation creation. On AA Briefcase v1.1, which measures office assignments extending over several hours, Grok 4.7 receives a score of 1,657. That compares with 1,546 for Grok 4.6 and 1,678 for Fable 5.1.
On GDPval, which covers work associated with multiple professions, the model receives an Elo score of 1,695. It finishes ahead of Grok 4.6 at 1,605 and GPT-6 Astra at 1,542, but behind Fable 5.1 at 1,735.
These tests broaden the product’s intended role. Coding remains its most prominent use case, but SpaceXAI is also targeting lawyers, analysts, health professionals, and teams producing structured business materials. Benchmark scenarios cannot substitute for testing against an organization’s own documents, policies, and software. A 500,000-token context window
The technical documentation lists a 500,000-token context window. Grok 4.7 accepts text and images as input and returns text. It also supports function calling and structured outputs, both of which are important when connecting a model to an agent or automated workflow.
A context window of that size can accommodate a substantial repository, a large document collection, or the accumulated history of a long-running assignment. It does not guarantee that every detail will be recalled or applied with equal accuracy. File selection, context organization, and the preservation of important decisions still matter.
Developers can select low, medium, high, or xHigh reasoning effort. High is the default setting, and reasoning cannot be fully disabled. The xHigh option allocates more computation to difficult problems, with correspondingly higher latency and consumption. Stable base pricing, with qualifications
Base API pricing remains $2 per million input tokens and $6 per million output tokens. Cached input costs $0.50 per million tokens. The model documentation notes that different pricing applies when a request exceeds 200,000 context tokens, so the listed base rate does not necessarily cover the entire 500,000-token window on the same terms.
A fast variant is also available. SpaceXAI says it doubles output speed while doubling the price. That claim concerns token generation speed, not necessarily the total completion time of an assignment involving network requests, external tools, or extended reasoning.
Grok 4.7 is available in Cursor and Grok Build, as well as through the company’s API, third-party coding harnesses, model routers, and cloud platforms. SpaceXAI has not released the model weights. Access is provided as a proprietary service through the company’s infrastructure and partner platforms. A new safety layer that remains difficult to compare
The company also describes an entirely new safeguard stack. It says Grok 4.7 is more resistant to jailbreak attempts and better at