Research

Forager: Rapid Data Exploration for Rapid Model Development

October 14, 2021
•
3 min read
•

Machine Learning Whiteboard (MLW) Open-source Series

We started our machine learning whiteboard (MLW) series earlier this year as an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in learning about machine learning.In this episode, Fait Poms, a Ph.D. student at Stanford working research and has worked with us on our data exploration interfaces, focuses on the question of “how quickly an ML practitioner, graduate student, or ML engineer can train a reasonable machine learning model for a new computer vision task?” while also diving into three exciting papers:

Paper Abstracts:

Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare Categories

For machine learning models trained with limited labeled training data, validation stands to become the main bottleneck to reducing overall data annotation costs. We propose a statistical validation algorithm that accurately estimates the F-score of binary classifiers for rare categories, where finding relevant examples to evaluate is particularly challenging. Our key insight is that simultaneous calibration and importance sampling enable accurate estimates even in the low-sample regime (< 300 samples). Critically, we also derive an accurate single-trial estimator of the variance of our method and demonstrate that this estimator is empirically accurate at low sample counts, enabling a practitioner to know how well they can trust a given low-sample estimate. When validating state-of-the-art semi-supervised models on ImageNet and iNaturalist2017, our method achieves the same estimates of model performance with up to 10x fewer labels than competing approaches. In particular, we can estimate model F1 scores with a variance of 0.005 using as few as 100 labels.

Mandoline: Model Evaluation under Distribution Shift

Many real-world ML deployments face the challenge of training a rare category model with a small labeling budget. In these settings, there is often access to large amounts of unlabeled data. Therefore it is attractive to consider semi-supervised or active learning approaches to reduce human labeling effort. However, prior approaches make two assumptions that do not often hold in practice; (a) one has access to a modest amount of labeled data to bootstrap learning, and (b) every image belongs to a common category of interest. In this paper, we consider the scenario where we start with as-little-as five labeled positives of a rare category and a large amount of unlabeled data, of which 99.9% of it is negatives. We propose an active semi-supervised method for building accurate models in this challenging setting. Our method leverages two key ideas: (a) Utilize human and machine effort where they are most effective; human labels are used to identify “needle-in-a-haystack” positives, while machine-generated pseudo-labels are used to identify negatives. (b) Adapt recently proposed representation learning techniques for handling extremely imbalanced human-labeled data to train models with noisy machine labeled data iteratively. We compare our approach with prior active learning and semi-supervised approaches, demonstrating significant improvements in accuracy per unit labeling effort, particularly on a tight labeling budget

Background Splitting: Finding Rare Classes in a Sea of Background

Share this article

Recommended articles

View all articles
Image
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for
September 23, 2026
•
Ankit Aich
,
Jonathan Schlosser
series-e-blog
Data 2.0 and the research era of AI data
Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.
September 22, 2026
•
Alex Ratner
Image
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer.
September 22, 2026
•
Jonathan Schlosser
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.