author

Srikar Kodati

Research Engineer, Benchmarks
,
Snorkel AI

Srikar Kodati is a Research Engineer at Snorkel AI, working on benchmarks and evaluations and for the Open Benchmark Grants program. Srikar previously developed Bank of America’s AML model and  large scale recommendation models at OTG Management.

The latest from Srikar Kodati

Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives
Blog
NEW
Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives

Terminal-Bench 3.0 (formerly Frontier-Bench) recently launched, built to track what AI agents can and can’t do across real computer work. Terminal-Bench 2.1 has been saturating, with top agents reaching 84%; on Terminal-Bench 3.0, the best model, Claude Opus 5, achieves just 43.5%. Terminal-Bench 3.0 raises the bar with 74 authentic, verifiable tasks across 7 domains, designed to expose meaningful gaps…

Aug 24, 2026
Learn more about Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives
 Illution Back
Illution Front

For models that need to be right. Not just good enough.