Announcing Snorkel’s $350M fundraise to build the frontier lab for AI data.
Introduction
Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.
Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M. We’re honored to now partner with the leading frontier labs, hyperscalars, neolabs, vertical AI leaders, enterprises, and U.S. government agencies who view data as one of the most important ingredients for safe and effective AI.
We started Snorkel as a research project a decade ago at Stanford. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one.
Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier and keep humans in the driver’s seat of AI progress for decades to come.
At Snorkel, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier.
With this round of funding, we are also doubling down on our commitments to support open data development for benchmarking and evaluation; an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data.
Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at Snorkel.
Data 2.0 – the research lab era of AI data


All modern automation approaches – e.g. digital LLMs, self-driving, physical AI – follow an asymptotic improvement curve where the first mile is driven by volume of simpler data, and the last mile is driven by quality of more complex data. Intuitively: when a model knows very little, almost any data contains additive bits of information. When a model is expert level, only the right data at the right level of quality and complexity will move the needle.
The first mile phase – “Data 1.0” – is all about volume, and is largely a staffing and logistics problem.
The enduring last mile phase – “Data 2.0” – is all about the right curriculum of extremely complex, precisely targeted, high quality data – and is primarily a research and technology problem.
As a concrete example, consider coding data. Several years ago, LLMs could barely do code auto-complete, and the focus was on large volumes of raw coding data for pre-training, and then, large volumes of simple coding problems and code preference labels for scaled SFT and RLHF.
Today: LLMs are approaching superhuman capabilities in many areas of coding. Frontier data is more valuable than ever given the economic impacts, but harder than ever to produce. Good coding datasets and environments for RL must approximate software development problems that a senior engineer might struggle with over days or weeks; must capture nuanced goals and reward signals over complex product-scale outputs; must distributionally target nuanced model error modes, like a finely calibrated curriculum for an advanced student; must be robust to these advanced models’ attempts to hack, cheat, or circumvent; and often must pass several hundred other quality control checks to be viable and aligned with increasingly precise training objectives.
The same pattern holds across an exploding surface area of domains, from legal to finance to biology and beyond. Where once getting a sufficient volume of relevant experts to develop basic chatbot Q&A tasks or preference labels was enough, now frontier data points and environments represent advanced expert tasks that might take humans days or weeks to complete, have nuanced failure modes or routes for AI cheating, require complex environments simulating e.g. entire companies, and are beyond the ability of even the smartest human experts to develop alone at scale.
Frontier data is no longer solvable by the optimized supply of human hours and bodies. The frontier ahead will be driven by research-driven approaches combining the best of human expertise and AI acceleration.
Data as a human and machine endeavor
While humans alone will increasingly struggle to develop frontier data, humans must remain at the center of the AI data development loop – accelerated and improved by AI.
Frontier data must be human-centric, first of all, because frontier data is valuable only insomuch as it targets something that a model does not already know. Purely synthetic data will naturally be highly correlated with what an AI model already knows – not what it still needs to learn. The basic circularity of AI creating its own data, and the resulting mode collapse of training on such pure synthetic data (other than for distillation), mandates that human-in-the-loop data will be the most valuable kind.
Second, from a normative standpoint: data is how we measure AI and align it to human values and judgement. Data developed without any humans in the loop is tantamount to abdicating human oversight and alignment entirely. Therefore, humans must be involved in the development of data for core reasons of oversight, alignment, and safety.
However: the new frontier of data, as already noted, has become far too complex for humans alone. Modern RL environments and datasets now simulate tasks that might take a human days or weeks to accomplish; require hundreds of quality control checks to avoid subtle errors leading to misaligned AI; and must stump the latest advanced models.
The key challenge, then, is to build AI systems that accelerate and improve human expertise – and in doing so, enable and empower humans to stay at the center of AI data evaluation and development as model capabilities exponentially advance in the decades ahead.
This is a deep technical project with years of rich work ahead covering many distinct angles. We’ve studied this problem academically for a decade, and built our internal Agentic Data Platform at Snorkel to support it, but there are years of exciting research and development problems ahead, advancing the ways in which specialized AI models and agents can:
- Accelerate human development effort by expanding seeds, constraints, and/or sketches from experts into properly constructed environments and data instances
- Guide human efforts towards distributional targets and model error modes.
- Give live feedback and do post-submission quality control, review, and revision
- Route data/environment subcomponents to the right human experts for targeted review
And so on, across a rich and growing surface area of human-computer interaction.
The end goal being data that is actually additive to the frontier of AI – and a process of building it that keeps humans in the driver’s seat even as that frontier rapidly accelerates.
Building the RSI engine for data development


The most exciting consequence of a human-AI system is its ability to drive a recursive self improvement loop for data, with specialized AI models improving human outcomes, and human supervision in turn improving these models – creating a powerful compounding loop to keep pace with an accelerating RSI model frontier.
Much of our work over the last decade of research and development has focused on this critical loop, and our Agentic Data Platform is designed primarily to propel and harvest this core dynamic. Human experts are supported by a stable of specialized AI models and agents – sometimes hundreds per task type – which accelerate and improve human outcomes to keep pace with the advancing frontier. In turn, human feedback at scale is used as weak supervision to continually measure and improve these agents.
As one example: when building coding agent environments and data, we use hundreds of specialized agents to do quality control in addition to human expert review, which currently accelerates our QC efficiency by 50%+ and improves accuracy of review by 15+ accuracy points compared to a human + off-the-shelf-only LLM review baseline. Scaled human review and feedback is then used as weak supervision to improve these agents, resulting in a current 2x+ accuracy improvement over a non-specialized frontier LLM baseline. AI improves human experts, and human experts improve AI – and statistics like these continuously improve as the flywheel continues.
The RSI loop for data development is what has powered our growth to date. This will compound and accelerate in the years ahead, allowing human-in-the-loop data to keep pace with the RSI-fueled acceleration of frontier AI.
Guiding the way with open benchmarks


We believe that it is more important than ever that AI development is guided and measured by an ecosystem of open, independent, and robust benchmarks – all driven by data and environment development.
Benchmarks are tests for AI – essentially, collections of datasets and environments that measure progress in key capability areas, and serve as both leaderboards and guideposts for AI progress. In the history of AI, benchmarks have always played an outsized role in directing research and development efforts. And at their best, they also provide critical transparency and insight into model performance, promote safer usage and deployment, and form the backbone of empirical computer science.
While benchmarks sometimes get pushback as becoming targets for gamification and overfitting, we believe that the best mitigation for this “benchmaxxing” failure mode is an ecosystem of more benchmarks, both public and private, that are robust, diverse, continuous, and independently created. Goodhart’s law – the famous adage that any metric which becomes a target ceases to be a good metric – did not imply that we should drop all metrics, but rather make them more diverse, robust, and less gameable; similarly with benchmarks, the way forward is more, not less.
At Snorkel, we believe that a majority of public benchmarks should be developed in open, independent ways, in order to support neutral and diverse evaluation of AI capabilities and gaps. To forward this, we will be significantly expanding our Open Benchmarks Grants program, which provides funding, research, and data development support to open, independent benchmarks, and has supported projects like Terminal Bench, OSWorld 2.0, Agents’ Last Exam, and many more to date. More news here very soon!
Supporting a diverse ecosystem of intelligence
We believe that the AI ecosystem will ultimately evolve to become a rich one, containing both massive generalist models at the frontier, and a myriad of specialized models tuned for every enterprise, organization, and perhaps even every individual.
At every level of specialization, different data and environments are needed. Where generalist frontier models win via maximal coverage of an area, specialized models win via deep focus on specific workflows, use cases, environments, and existing user data or signals that serve as starting points and anchors for data development.
We believe that an increasingly large portion of the world’s data development will focus on supporting specialized AI over the next several years, leading to an increasingly diverse ecosystem of intelligence – and plan to invest heavily in support of this at Snorkel.
Supporting safe intelligence
Last but not at all least: AI safety is one of the most critical considerations in the days and years ahead – and we believe that safe intelligence fundamentally begins with robust, high quality data.
Datasets and environments not only drive how we benchmark and evaluate models for safety; they define in a very literal way the objectives and penalties that shape AI model learning and alignment during training. Strong alignment comes from high-quality, well-designed data and environments that are realistic; comprehensive and diverse in their coverage of real world scenarios; and robust to increasingly advanced AI capabilities for cheating or “reward hacking” from task definition and environment construction to rubric design and evaluator development. Just as human alignment starts with lessons learned as children at home, so too must AI alignment and safety start with robust design of training datasets and environments.
A significant portion of our research and development efforts in the years ahead will focus on advancing the science and technology of data and environment development for robust alignment and safe AI – and we are excited to make a significant impact here.
The new frontier of AI research
As AI capabilities rapidly improve, we believe that advancing the science and technology of data and environment development will be one of the most interesting and impactful areas of AI research. The challenges ahead – defining the shape of new Data 2.0 workloads; accelerating the critical collaboration of human experts and AI, and shaping this interaction into an RSI data engine; defining the open benchmarks that guide the field; diversifying to support an expanding ecosystem of intelligence; and supporting fundamentally safe and aligned intelligence – will be some of the most central to AI progress.
We are incredibly excited for the next decade of research ahead at Snorkel.


Alex Ratner is the co-founder and CEO at Snorkel AI, and an affiliate assistant professor of computer science at the University of Washington. Prior to Snorkel AI and UW, he completed his Ph.D. in computer science advised by Christopher Ré at Stanford, where he started and led the Snorkel open source project. His research focused on data-centric AI, applying data management and statistical learning techniques to AI data development and curation.
Recommended articles









