# Snorkel AI > Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine platform technology with research-driven data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. ## Company - [Snorkel AI](https://snorkel.ai/): Official homepage and overview. - [About Snorkel AI](https://snorkel.ai/company/): Company mission, history, founders, and current positioning. - [Frequently Asked Questions](https://snorkel.ai/frequently-asked-questions/): Canonical answers about what Snorkel is, what it delivers, and how teams work with the company. - [How Snorkel Works](https://snorkel.ai/how-it-works/): Snorkel's platform, research-validated methodology, expert workflows, and Evaluate → Curate → Refine development loop. ## Data and Environments - [Data Development](https://snorkel.ai/data-development/): Snorkel Data Series and custom datasets, benchmarks, evaluations, and environments for frontier AI. - [Use Cases](https://snorkel.ai/use-cases/): Applications across coding, agentic reasoning, computer use, multimodal AI, and specialized domains. ## Specialized Agents - [Specialized Agents](https://snorkel.ai/specialized-agents/): Custom agents grounded in enterprise-specific data, knowledge, workflows, and operating environments. - [Enterprise Stories](https://snorkel.ai/enterprise-stories/): Examples of real-world enterprise deployment and measurable outcomes. ## Research and Evaluation - [Research](https://snorkel.ai/research/): Research in data development, evaluation, training systems, benchmarks, and frontier AI. - [Benchmark Leaderboards](https://snorkel.ai/leaderboard/): Model and agent results across Snorkel's benchmark portfolio. - [Open Benchmarks Grants](https://benchmarks.snorkel.ai/): Snorkel's program supporting open, community-driven benchmarks and evaluation research for agentic AI. - [Snorkel AI Reading Group](https://snorkel.ai/reading-group/): A recurring forum examining emerging work in AI benchmarking, evaluation, data development, and agentic systems. ## Snorkel-Developed Benchmarks - [Senior SWE-Bench](https://snorkel.ai/leaderboard/senior-swe-bench/): Evaluates coding agents on realistic senior-level software engineering work. - [Senior SWE-Bench Leaderboard (Live)](https://senior-swe-bench.snorkel.ai/): The live, continuously updated Senior SWE-Bench site with the full task taxonomy, agent leaderboard, dataset, and citation. ## Open Benchmarks Grants-Supported Projects - [Terminal-Bench 3.0](https://snorkel.ai/leaderboard/terminal-bench-3-0/): A harder, more domain-diverse successor to Terminal-Bench 2.1, assembled in the open under continuous adversarial review, built with Harbor and the Laude Institute. - [Agents' Last Exam](https://snorkel.ai/leaderboard/agents-last-exam/): Evaluates agents on long-horizon, economically valuable professional workflows with verifiable outcomes. - [OSWorld 2.0](https://snorkel.ai/leaderboard/os-world-2-0/): Evaluates computer-use agents on long-horizon workflows across web and desktop environments. - [Terminal-Bench 2.1](https://snorkel.ai/leaderboard/terminal-bench-2-1/): Evaluates agents on challenging work in terminal environments. - [Continual Learning Bench](https://snorkel.ai/leaderboard/continual-learning-bench/): Measures whether agents improve across sequential, stateful tasks. - [SlopCode Bench](https://snorkel.ai/leaderboard/slopcode-bench/): Measures how code quality changes as coding agents repeatedly modify and extend their solutions. ## Collaborative Benchmarks and Research Initiatives - [CUA-Bench](https://snorkel.ai/cua-bench-computer-use-agent-benchmark/): Evaluates computer-use agents on professional software. - [BigLaw Bench: Research](https://snorkel.ai/biglaw-bench-research/): A Harvey benchmark built with Snorkel AI as the data partner for hard agentic legal research problems. - [Terminal-Bench Science](https://snorkel.ai/terminal-bench-science/): Extends terminal-based agent evaluation to computational scientific workflows. ## Research Papers - [Learning from Less: Effectiveness of RLVR in Low Data and Compute Regimes](https://snorkel.ai/research-paper/learning-from-less-rlvr-low-data-compute-effectiveness/): Accepted to MLSys 2026. - [RIFT: A Rubric Failure Mode Taxonomy and Automated Diagnostics](https://snorkel.ai/research-paper/rift-a-rubric-failure-mode-taxonomy-and-automated-diagnostics/): Accepted to an ICLR 2026 workshop. - [Benchmarking Agents in Insurance Underwriting Environments](https://snorkel.ai/research-paper/benchmarking-agents-in-insurance-underwriting-environments/): Research on agentic reasoning and tool use in commercial underwriting. ## Technical Articles and Research Discussions - [Milestone-Based Evaluation and Training for Long-Horizon AI Agents](https://snorkel.ai/blog/long-horizon-ai-agent-evaluation-training/): Milestone-based evaluation and training help long-horizon AI agents move beyond pass/fail scores with phase-level diagnostics, partial credit, and better RL reward signals. - [Enterprise Environments and Training AI Agents for Real-World Workflows](https://snorkel.ai/blog/enterprise-environments-ai-agents/): Enterprise environments give AI agents realistic tools, data, and policies to train and evaluate them on real workflows, like insurance underwriting. - [Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks](https://snorkel.ai/blog/opus-5-swe-bench-error-analysis/): Opus 5 ranks #2 on Senior SWE-Bench and leads on bug investigation; a trajectory-level analysis of 195 runs shows why, and where it still fails. - [Senior SWE-Bench: Evaluating Coding Agents Like Senior Engineers](https://snorkel.ai/blog/senior-swe-bench-evaluating-coding-agents-like-senior-engineers/): A reading-group deep dive into Senior SWE-Bench, a benchmark for evaluating coding agents on realistic, long-horizon software engineering tasks and code quality. - [Data Quality and Rubrics: How to Build Trust in Your Models](https://snorkel.ai/blog/data-quality-and-rubrics-how-to-build-trust-in-your-models/): How rubric-based evaluation transforms data annotation, enabling scalable, high-quality judgment for GenAI models across complex, open-ended tasks. - [The Science of Rubric Design](https://snorkel.ai/blog/the-science-of-rubric-design/): How to design and iterate rubrics by treating them like models, maximizing goal alignment and inter-rater agreement. - [Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning](https://snorkel.ai/blog/introducing-snorkelspatial/): A procedurally generated, programmatically verified benchmark for evaluating spatial reasoning capabilities in LLMs. - [4B FinQA Model Outperforms 235B Model with the Right Data](https://snorkel.ai/blog/building-finqa-an-open-rl-environment-for-financial-reasoning-agents/): Technical research co-authored with Berkeley. - [Why Coding Agents Need Better Data, Evals, and Environments](https://snorkel.ai/blog/coding-agents-evals/): An overview of the data, evaluations, and environments behind coding agents. - [Continual Learning for AI Agents](https://snorkel.ai/blog/continual-learning-ai-agents-explained/): A research explainer on continual learning for AI agents. - [GRPO](https://snorkel.ai/grpo/): An educational resource on Group Relative Policy Optimization. - [OLMIX: Data Mixing Throughout Language Model Development](https://snorkel.ai/blog/olmix-data-mixing-reading-group/): A research reading on data mixing throughout language model development. - [Code World Models and AutoHarness for LLM Agents](https://snorkel.ai/blog/code-world-models-and-autoharness-for-llm-agents/): A research reading on code world models and agent environments. - [Collaborative Gym](https://snorkel.ai/blog/collaborative-gym-a-framework-for-enabling-and-evaluating-human-agent-collaboration/): A research reading on enabling and evaluating human-agent collaboration. - [JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment](https://snorkel.ai/blog/judgment-bench-comparing-rubric-and-preference-evaluation-for-quality-assessment-legal-ai/): A research reading on quality assessment for legal AI. ## Additional Resources - [Snorkel Blog](https://snorkel.ai/blog/): Research findings, benchmark results, technical perspectives, and company announcements. - [Snorkel Press](https://snorkel.ai/press/): Official announcements and media resources.