Research

Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial

September 22, 2026
3 min read

Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8.

Benchmark

Senior SWE-Bench evaluates agents as compared to a senior engineer. Instructions are under-specified to be more like a message from a colleague rather than a requirements document. Bug tasks are sourced from real PRs that needed actual runtime investigation to resolve. And scoring rewards code quality, not just correctness. For a tasteful solve in Senior SWE-Bench, agents must produce code that meets behavioral requirements and passes conservative thresholds on code quality measurements.

Tasks are drawn from real PRs across 50 public and 50 private problems spanning different libraries and multi-service applications. The benchmark is hard by construction: the strongest frontier models fail to produce a tasteful solve on more than 65% of tasks.

Results are reported as pass and tasteful solve rates at k=1 and k=3, alongside average output token cost per trial. Grok 4.7 results below are at xhigh effort.

Join our newsletter
Get the scoop on new benchmarks, research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.

Results

Grok 4.7 improves on Grok 4.6 on both the tasteful pass@1 and pass@3. Tasteful pass@1 rises from 26.3% to 27.4% and tasteful pass@3 rises from 38.9% to 40.0%. The pass@3 result places it sixth overall, level with Opus 4.8 at 40.0%.

The cost is notably lower compared to other leading models. Grok 4.7 and Opus 4.8 post identical tasteful pass@3 rates, but Grok 4.7 reaches that at $0.24 per trial against $3.36, using roughly a third the output tokens. Against Fable 5.1 and GPT-5.6 Sol the gap is pricing rather than efficiency: Grok 4.7 spends a comparable number of output tokens as both — 39.4K against 37.9K and 32.7K — at a quarter to an eighth the cost. It also matches GPT-5.6 Sol’s pass@3 rate at 63.2% and is only three points behind on tasteful pass@3.

Grok 4.7 also solves a slightly smaller share of tasks than Grok 4.6, at a pass@3 of 63.2% against 65.3%, but converts more of those solves into tasteful ones: 63.3% against 59.6%. The leader group converts between 64% and 72%.

Reliability is the weaker axis. Grok 4.7 posts a tasteful pass^3 of 7.4%, against 22.1% for Fable 5.1 and 16.8% for Grok 4.6. Its capability shows up across repeated attempts more than on any given one, and by this measure it is less consistent than its predecessor.

Conclusions

Senior-level engineering remains an open frontier as even the most performant models still struggle on many of the tasks. With that in mind, Grok 4.7 is a true improvement over 4.6. It sits just outside the leader group on the benchmark’s headline capability measure while costing between a quarter and a 1/14th as much per trial as every model ranked above it. While it still has limitations on first attempts, the performance is an improvement, and at a fraction of the cost. For teams running agents in loops, where retries are not a major concern, that combination may matter more than first-attempt performance alone.

You can see the full results on the leaderboard at Senior SWE-Bench.

Share this article

Jonathan Schlosser is an AI Advocate at Snorkel.ai. He has previously worked as an AI Engineer, an Educator, and Founder. He has taught graduate-level AI and Data Science courses at UNC Chapel Hill for the last few years, and has taught hundreds of data science professionals through programs with Correlation One, Amazon, Google, and others. With 10+ years across the data and AI spectrum, from data science roles to building AI-driven applications, he writes about deep learning, LLMs, and AI agents with a direct and simplified instructional approach.

Recommended articles

View all articles
Image
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for
September 23, 2026
Ankit Aich
,
Jonathan Schlosser
series-e-blog
Data 2.0 and the research era of AI data
Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.
September 22, 2026
Alex Ratner
Image
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later, usually in ways that are harder to trace back and fix.
September 15, 2026
Jonathan Schlosser
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.