Tag

Evaluation

AI evaluation systematically measures a model’s performance on tasks. Classically, this applied metrics like accuracy or precision to clear and discrete numerical or categorical targets. Moden evaluation also assesses the output of generative models to ensure they create content within an organization’s standards and guidelines.

All articles on Evaluation

Image
LLM-as-a-judge for enterprises: evaluate model alignment at scale
Discover how enterprises can leverage LLM-as-Judge systems to evaluate generative AI outputs at scale, improve model alignment, reduce costs, and tackle challenges like bias and interpretability.
March 26, 2025
Matt Casey
Image
Why GenAI evaluation requires SME-in-the-loop for validation and trust
It’s critical enterprises can trust and rely on GenAI evaluation results, and for that, SME-in-the-loop workflows are needed. In my first blog post on enterprise GenAI evaluation, I discussed the importance of specialized evaluators as a scalable proxy for SMEs. It simply isn’t practical to task SMEs with performing manual evaluations – it can take weeks if not longer, unnecessarily
March 20, 2025
Shane Johnson
Image
Why enterprise GenAI evaluation requires fine-grained metrics to be insightful
GenAI needs fine-grained evaluation for AI teams to gain actionable insights.
March 18, 2025
Shane Johnson
Image
What is specialized GenAI evaluation, and why is it so critical to enterprise AI?
Specialized GenAI evaluation ensures AI assistants meet business requirements, SME expertise, and industry regulations—critical for production-ready AI.
March 5, 2025
Shane Johnson
Image
How LLM evaluation drives better models in Snorkel Flow
Discover how Snorkel AI’s methodical workflow can simplify the evaluation of LLM systems. Achieve better model performance in less time.
December 17, 2024
Rebekah Westerlind
Image
LLM evaluation in enterprise applications: a new era in ML
Learn about the obstacles faced by data scientists in LLM evaluation and discover effective strategies for overcoming them.
November 25, 2024
Matt Casey
Image
Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows
Snorkel AI has made building production-ready, high-value enterprise AI applications faster and easier than ever. The 2024.R3 update to our Snorkel Flow AI data development platform streamlines data-centric workflows, from easier-than-ever generative AI evaluation to multi-schema annotation.
October 9, 2024
Matt Casey
Image
How data slices transform enterprise LLM evaluation
Enterprises must evaluate LLM performance for production deployment. Custom, automated eval + data slices present the best path to production.
August 1, 2024
Vincent Sunn Chen
Image
Data-centric AI with Snorkel and MinIO
High-performing AI systems require more than a well-designed model. They also require properly constructed training and testing data.
July 12, 2024
Keith Pijanowski (Guest blogger)
Image
Weak supervision for non-categorical applications + superalignment
We need more labeled data than ever, so we have explored weak supervision for non-categorical applications—with notable results.
July 2, 2024
Changho Shin
Image
Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm
To tackle generative AI use cases, Snorkel AI + AWS launched an accelerator program to address the biggest blocker: unstructured data.
June 27, 2024
Team Snorkel
Image3
How does the Snorkel Flow label model work?
The Snorkel Flow label model plays an instrumental role in driving the enterprise value we create. Here’s a peek at how it works.
June 18, 2024
Chris Glaze
Image
Vision language models: how LLMs boost image classification
Vision language models demonstrate impressive image classification capabilities, but LLMs can help improve their performance. Learn how.
June 12, 2024
Reza Esfandiarpoor
Image1
Long context models in the enterprise: benchmarks and beyond
Snorkel researchers devised a new way to evaluate long context models and address their “lost-in-the-middle” challenges with mediod voting.
June 6, 2024
Amanda Dsouza
Image
How to build production-grade RAG retrieval with Snorkel Flow
See a walkthrough of how Snorkel Flow users build applications with production-grade RAG retrieval components.
June 4, 2024
Marty Moesta
Image
How Bonito helps fine-tune specialized LLMs faster than ever
Fine-tuning specialized LLMs demands a lot of time and cost We developed Bonito to make this process faster, cheaper, and easier.
May 28, 2024
Nihal Nayak
Image
How ROBOSHOT boosts zero-shot foundation model performance
ROBOSHOT acts like a lens on foundation models and improves their zero-shot performance without additional fine-tuning.
April 30, 2024
Dyah Adila
Image1
How Snorkel topped the AlpacaEval leaderboard (and why we’re not there anymore)
Snorkel AI placed a model at the top of the AlpacaEval leaderboard. Here’s how we built it, and how it changed AlpacaEval’s metrics.
April 9, 2024
Hoang Tran
Image1
CRFM’s HELM and enterprise LLM evaluation beyond accuracy
As Snorkel AI prepares to build better enterprise LLM evaluations, we spoke with Yifan Mail from Stanford’s CRFM HELM project.
April 3, 2024
Vivek Krishnamurthy
Image
Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days
Snorkel AI helped a client solve the challenge of social media content filtering quickly and sustainably. Here’s how.
March 26, 2024
Gabe Smith
Snorkel teams with Microsoft to showcase new AI research at NVIDIA image
Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC
Microsoft infrastructure facilitates Snorkel AI research experiments, including our recent high rank on the AlpacaEval 2.0 LLM leaderboard.
March 18, 2024
Snorkel Team
Image
How Skill-it! enables faster, better LLM training
Humans learn tasks better when taught in a logical order. So do LLMs. Researchers developed a way to exploit this tendency called “Skill-it!”
March 12, 2024
Fred Sala
Image2
Enterprise GenAI to surge in 2024: survey results
Enterprise GenAI 2024: applications will likely surge toward production, according to Snorkel AI Enterprise LLM Summit survey results .
February 29, 2024
Matt Casey
Image
Retrieval augmented generation (RAG): a conversation with its creator
Snorkel CEO Alex Ratner spoke with Douwe Keila, an author of the original paper about retrieval augmented generation (RAG).
January 16, 2024
Team Snorkel
Image
Stanford professor discusses exciting advances in foundation model evaluation
Snorkel CEO Alex Ratner chatted with Stanford Professor Percy Liang about evaluation in machine learning and in AI generally.
January 2, 2024
Team Snorkel
Image
How predictive AI + generative AI build amazing document understanding
A proof-of-concept project that combines predictive AI + generative AI to minimize LLM’s risks while keeping their advantages.
December 5, 2023
Shahebaz Mohammad
Image
How to fine-tune Llama 2 in Snorkel Flow
Data scientists can fine-tune Llama 2 to adapt it to specific tasks. The Snorkel Flow data development platform makes it easy to do so.
November 28, 2023
Hoang Tran
Image
How AI-powered claims processing creates new efficiencies in insurance
Insurance claims processing has long required a lot of tedious and expensive human labor, but artificial intelligence (AI) can help.
October 18, 2023
Team Snorkel
Image1
Bloomberg’s Gideon Mann on the power of domain specialist LLMs
Gideon Mann, head of ML Product and Research at Bloomberg LP, chatted with Snorkel CEO Alex Ratner about building BloombergGPT.
October 17, 2023
Team Snorkel
Image
How to fine-tune GPT-3.5 Turbo in Snorkel Flow
Snorkel Flow makes it easy to fine tune LLMs like GPT-3.5 Turbo to work better for specific domain and enterprise requirements.
October 13, 2023
Hoang Tran