Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8.


Benchmark
Senior SWE-Bench evaluates agents as compared to a senior engineer. Instructions are under-specified to be more like a message from a colleague rather than a requirements document. Bug tasks are sourced from real PRs that needed actual runtime investigation to resolve. And scoring rewards code quality, not just correctness. For a tasteful solve in Senior SWE-Bench, agents must produce code that meets behavioral requirements and passes conservative thresholds on code quality measurements.
Tasks are drawn from real PRs across 50 public and 50 private problems spanning different libraries and multi-service applications. The benchmark is hard by construction: the strongest frontier models fail to produce a tasteful solve on more than 65% of tasks.
Results are reported as pass and tasteful solve rates at k=1 and k=3, alongside average output token cost per trial. Grok 4.7 results below are at xhigh effort.
Results
Grok 4.7 improves on Grok 4.6 on both the tasteful pass@1 and pass@3. Tasteful pass@1 rises from 26.3% to 27.4% and tasteful pass@3 rises from 38.9% to 40.0%. The pass@3 result places it sixth overall, level with Opus 4.8 at 40.0%.


The cost is notably lower compared to other leading models. Grok 4.7 and Opus 4.8 post identical tasteful pass@3 rates, but Grok 4.7 reaches that at $0.24 per trial against $3.36, using roughly a third the output tokens. Against Fable 5.1 and GPT-5.6 Sol the gap is pricing rather than efficiency: Grok 4.7 spends a comparable number of output tokens as both — 39.4K against 37.9K and 32.7K — at a quarter to an eighth the cost. It also matches GPT-5.6 Sol’s pass@3 rate at 63.2% and is only three points behind on tasteful pass@3.
Grok 4.7 also solves a slightly smaller share of tasks than Grok 4.6, at a pass@3 of 63.2% against 65.3%, but converts more of those solves into tasteful ones: 63.3% against 59.6%. The leader group converts between 64% and 72%.
Reliability is the weaker axis. Grok 4.7 posts a tasteful pass^3 of 7.4%, against 22.1% for Fable 5.1 and 16.8% for Grok 4.6. Its capability shows up across repeated attempts more than on any given one, and by this measure it is less consistent than its predecessor.
Conclusions
Senior-level engineering remains an open frontier as even the most performant models still struggle on many of the tasks. With that in mind, Grok 4.7 is a true improvement over 4.6. It sits just outside the leader group on the benchmark’s headline capability measure while costing between a quarter and a 1/14th as much per trial as every model ranked above it. While it still has limitations on first attempts, the performance is an improvement, and at a fraction of the cost. For teams running agents in loops, where retries are not a major concern, that combination may matter more than first-attempt performance alone.
You can see the full results on the leaderboard at Senior SWE-Bench.


Jonathan Schlosser is an AI Advocate at Snorkel.ai. He has previously worked as an AI Engineer, an Educator, and Founder. He has taught graduate-level AI and Data Science courses at UNC Chapel Hill for the last few years, and has taught hundreds of data science professionals through programs with Correlation One, Amazon, Google, and others. With 10+ years across the data and AI spectrum, from data science roles to building AI-driven applications, he writes about deep learning, LLMs, and AI agents with a direct and simplified instructional approach.
Recommended articles









