Research

Accuracy top concern for Foundation Model adoption—Poll

January 31, 2023
3 min read

Most poll respondents at Snorkel AI’s recent Foundation Model Virtual Summit named questionable accuracy as the biggest barrier preventing them from getting organizational value from Foundation Models.

The January 17 Summit brought together 12 presenters and more than 600 attendees at 10 virtual sessions. Between presentations, we polled attendees on how they expect Foundation Models to fit into their organizations. The results point to broad enthusiasm for Foundation Models (FMs) and Large Language Models (LLMs). But hesitations endure.

Attendees ranked concerns about governance and cost closely behind quality. Notably, just 8% of respondents said they expected “poor use case fit” to block them from using FMs, which suggests that almost all respondents expect to make use of FMs at some point, and tracks closely with other poll results from the same event.

Poll results: top uses and challenges of FMs

Foundation Models will likely provide their greatest value through classification and extraction tasks, according to nearly half of the respondents to our poll—though realizing that value may have to wait. Enterprises generally require a high rate of accuracy to deploy new classification or extraction tools into production, and the results of our first question show that respondents aren’t sure FMs are ready to clear that bar.

We also asked attendees to estimate the timeline by which they expected their organization to deploy its first production use of FMs or LLMs. Of those who responded, only 6% said they were unlikely to use Foundation Models in their business. That mirrors the 8% who named “poor use-case fit” as the reason FMs may not yield value for their organization, but there is some nuance between these questions.

Deploying a new resource doesn’t necessarily mean that the organization will get value out of it. Many experiments fail, but it appears that most of our attendees are eager to experiment. Nearly two-thirds of respondents expected their organization to launch its first production use of FMs in the next year—if it hasn’t already.

Final thoughts

The above results are subject to meaningful error due to our small sample size and significant sampling bias; the respondents all chose to attend a half-day educational session about Foundation Models and answer our optional questions.

However, given that our attendees spanned many industries, including technology, healthcare, and financial services, we think these results—flawed as they are—indicate substantial interest in this new era of artificial intelligence and machine learning.

Like our attendees, we at Snorkel are excited about Foundation Models. Our researchers have published papers on how to get the greatest value from FMs, and our engineers have integrated FM-enabled features into the Snorkel Flow platform. 

To see our new Foundation Model features in action, register now for a Snorkel Flow live demo on February 16, 2023.

Share this article
Image
Matt Casey
Data Science Content Lead

Matt Casey leads content production at Snorkel AI. In prior roles, Matt built machine learning models and data pipelines as a data scientist. As a journalist, he produced written and audio content for outlets including The Boston Globe and NPR affiliates.

Recommended articles

View all articles
os-world-reading-group
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each.
September 3, 2026
Snorkel Team
Image
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat
September 1, 2026
Ankit Aich
,
Jonathan Schlosser
Image
Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
The speed of new frontier model releases keeps accelerating. Meanwhile benchmarks struggle to keep up and saturate quickly, often being left in the dust. Most benchmarks are static datasets with no active maintenance, causing them to lose value fast. Some benchmarks are looking to change this by becoming Continuous Benchmarks. Terminal-Bench is one of the most widely reported benchmarks on
August 28, 2026
Justin Bauer
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.