In the news
Image

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

September 4, 2026

Fortune featured Vincent Sunn Chen’s perspective on why benchmark scores can shift shortly before a model launch: results depend on the exact model checkpoint, compute budget, harness, and evaluation configuration, all of which may still be changing in the final days. Vincent called for clearer industry norms around disclosing those changes so researchers and users can interpret reported performance with greater confidence.

Share this article

Recommended press articles

View all press articles
Logo for Alex Ratner, Snorkel AI CEO on Data-Centric AI and AI Development
In the news
Alex Ratner, Snorkel AI CEO on Data-Centric AI and AI Development
September 24, 2026
Logo for Alex Ratner on Data 2.0 and Why AI Still Needs Human Expertise
In the news
Alex Ratner on Data 2.0 and Why AI Still Needs Human Expertise
September 24, 2026
Logo for Alex Ratner on Snorkel’s $350M Raise and the Shift to Data 2.0
In the news
Alex Ratner on Snorkel’s $350M Raise and the Shift to Data 2.0
September 24, 2026
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.