In the news
September 4, 2026
Fortune featured Vincent Sunn Chen’s perspective on why benchmark scores can shift shortly before a model launch: results depend on the exact model checkpoint, compute budget, harness, and evaluation configuration, all of which may still be changing in the final days. Vincent called for clearer industry norms around disclosing those changes so researchers and users can interpret reported performance with greater confidence.
Recommended press articles
View all press articles






