author

Ankit Aich

Senior Research Scientist
,
Snorkel AI

Ankit Aich is a senior research scientist at Snorkel AI. He focuses on frontier model benchmarking, data curation, failure point analysis, and research.

He finished his PhD and MS in computer science from UIC-CS. His research aims to improve patient health outcomes by working towards collaborative-AI technology. He has worked at multiple places including the National Institutes of Health, the University of Pennsylvania, Scale AI, and FINRA. Language modeling stands as a cornerstone of my work, serving as a powerful tool to decode and comprehend the complexities of our world. He believes language offers a unique lens through which we can observe and understand our surroundings.

The latest from Ankit Aich

Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Blog
NEW
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results

We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for…

Sep 23, 2026 •
Learn more about Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Blog
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes

We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat…

Sep 01, 2026 •
Learn more about Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks
Blog
Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks

Anthropic’s Claude Opus 5 recently debuted as the second model overall on the current Senior SWE-bench leaderboard, behind Fable 5. It also achieves the highest score of any evaluated model on the benchmark’s Bug & Performance Investigation category, reinforcing the rapid progress frontier coding models continue to make on increasingly realistic software engineering tasks. Just as notable, Opus 5 reaches…

Jul 27, 2026 •
Learn more about Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks
 Illution Back
Illution Front

For models that need to be right. Not just good enough.