Snorkel Logo

On-demand agenda

Evaluating LLM Systems

LLM evaluation is critical for generative AI in the enterprise, but measuring how well an LLM answers questions or performs tasks is difficult. Thus, LLM evaluations must go beyond standard measures of “correctness” to include a more nuanced and granular view of quality.

In practice, enterprise LLM evaluations (e.g., OSS benchmarks) often come up short because they’re slow, expensive, subjective, and incomplete. They leave AI initiatives blocked because there is no clear path to production quality.

In this session, Venkatesh Rao, Staff Product Manager at Snorkel AI, and Rebekah Westerlind, Software Engineer at Snorkel AI, will discuss the importance of LLM evaluation, highlight common challenges and approaches, and explain the core concepts behind Snorkel AI’s approach to data-centric LLM evaluation.

Join us to learn more about:

  • Understanding the nuances of LLM evaluation
  • Evaluating LLM response performance at scale
  • Identifying where additional LLM fine-tuning is needed