We develop methods, benchmarks, and training systems that turn expert data into frontier AI

building benchmarks and collaborating with

Image
Image
Image
Image
Image
Image
Image
Image
Image
agent-le-logo
rdi-foundation
Cua Logo
Image
Image
key research areas

Vision and impact

We help labs advance frontier models by working with domain experts to design and build complex, realistic datasets that drive model performance.

initiatives

Community and open science

Open benchmarks, conversations, and research for real-world AI performance.

Image

Open Benchmarks Grants

Backed by a $30M commitment, Open Benchmarks Grants funds new benchmarks, red-teams existing ones to find weaknesses, and supports researchers through a fellowship.

Image

Benchtalks

Our podcast series at the intersection of AI evaluation, data quality, and real-world impact.
Image

Reading Group

A recurring forum for researchers and practitioners to explore the latest frontier developments in AI while building meaningful connections within the community.

DEEP RESEARCH Expertise

Technical advisors and distinguished affiliates

Stephen Bach headshot

Stephen Bach

Brown University
Eliot Horowitz Assistant Professor, Computer Science Department
Jason Fries headshot

Jason Fries

Stanford University
Assistant Professor of Biomedical Data Science and of Medicine
Jared Dunnmon headshot

Jared Dunnmon

Co-Founder & Chief Scientist, Stealth Startup
Prev. Dir. of AI at DIU
Fred Sala headshot

Fred Sala

Chief Scientist
,
Snorkel AI
Assistant Professor @ University of Wisconsin-Madison
Chris Ré headshot

Chris Ré

Co-Founder
,
Snorkel AI
Professor @ Stanford University
Ludwig Schmidt headshot

Ludwig Schmidt

Stanford University · LAION
Stanford researcher and LAION collaborator
Karthik Narasimhan headshot

Karthik Narasimhan

Princeton University
Professor of Computer Science
Yu Su headshot

Yu Su

Ohio State University
Associate Professor of Computer Science and Engineering
Lewis Tunstall headshot

Lewis Tunstall

Hugging Face
Machine Learning Engineer
PUBLICATIONS

Browse research blogs
and academic papers

Type: All Types
Sort: Newest
Long context models in the enterprise: benchmarks and beyond
Blog
Long context models in the enterprise: benchmarks and beyond

Snorkel researchers devised a new way to evaluate long context models and address their “lost-in-the-middle” challenges with mediod voting.

Jun 06, 2024 •
Learn more about Long context models in the enterprise: benchmarks and beyond
How ROBOSHOT boosts zero-shot foundation model performance
Blog
How ROBOSHOT boosts zero-shot foundation model performance

ROBOSHOT acts like a lens on foundation models and improves their zero-shot performance without additional fine-tuning.

Apr 30, 2024 •
Learn more about How ROBOSHOT boosts zero-shot foundation model performance
Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC
Blog
Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC

Microsoft infrastructure facilitates Snorkel AI research experiments, including our recent high rank on the AlpacaEval 2.0 LLM leaderboard.

Learn more about Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC
How Skill-it! enables faster, better LLM training
Blog
How Skill-it! enables faster, better LLM training

Humans learn tasks better when taught in a logical order. So do LLMs. Researchers developed a way to exploit this tendency called “Skill-it!”

Mar 12, 2024 •
Learn more about How Skill-it! enables faster, better LLM training
Large language model training: three phases that shape LLM training
Blog
Large language model training: three phases that shape LLM training

Training large language models is a multi-layered stack of processes, each with its unique role and contribution to the model’s performance.

Feb 27, 2024 •
Learn more about Large language model training: three phases that shape LLM training
LoRA: Low-Rank Adaptation for LLMs
Blog
LoRA: Low-Rank Adaptation for LLMs

Low-rank adaptation (LoRA) lets data scientists customize GenAI models like LLMs faster than traditional full fine-tuning methods.

Feb 21, 2024 •
Learn more about LoRA: Low-Rank Adaptation for LLMs
New benchmark results demonstrate value of Snorkel AI approach to LLM alignment
Blog
New benchmark results demonstrate value of Snorkel AI approach to LLM alignment

Snorkel researchers’ state-of-the-art methods created a 7B LLM that ranked 2nd, behind only GPT-4 Turbo, on AlpacaEval 2.0 leaderboard.

Jan 24, 2024 •
Learn more about New benchmark results demonstrate value of Snorkel AI approach to LLM alignment
Retrieval augmented generation (RAG): a conversation with its creator
Blog
Retrieval augmented generation (RAG): a conversation with its creator

Snorkel CEO Alex Ratner spoke with Douwe Keila, an author of the original paper about retrieval augmented generation (RAG).

Jan 16, 2024 •
Learn more about Retrieval augmented generation (RAG): a conversation with its creator
Characterizing the Impacts of Semi-supervised Learning for Weak Supervision
Labeling training data is a critical and expensive step in producing high accuracy ML models, whether training from scratch or fine-tuning. To make labeling more efficient, two major approaches are programmatic weak supervision (WS) and semi-supervised learning (SSL). More recent works have either explicitly or implicitly used techniques at their intersection, but in various complex and ad hoc ways. In this work, we define a simple, modular design space to study the use of SSL techniques for WS more systematically. Surprisingly, we find that fairly simple methods from our design space match the performance of more complex state-of-the-art methods, averaging...
Research Paper
Characterizing the Impacts of Semi-supervised Learning for Weak Supervision

Labeling training data is a critical and expensive step in producing high accuracy ML models, whether training from scratch or fine-tuning. To make labeling more efficient, two major approaches are programmatic weak supervision (WS) and semi-supervised learning (SSL). More recent works have either explicitly or implicitly used techniques at their intersection, but in various complex and ad hoc ways. In…

Jan 16, 2024 •

Jeffrey Li, Jieyu Zhang, Ludwig Schmidt & Alexander Ratner

Learn more about Characterizing the Impacts of Semi-supervised Learning for Weak Supervision
1 … 13 14 15 … 41

October 8, 2026 | San francisco

ImageImage

A one-day, invite-only summit providing a first look at the benchmarks and research that will shape the frontier.

 Illution Back
Illution Front

Let’s research together

Join our team of leading researchers and help shape the future of AI.