OpenAI Revises Astra's Performance Metrics Post-Launch, Raising Questions on Benchmark Integrity

OpenAI altered several evaluation benchmarks for its GPT-6 Astra model after the initial blog post publication, with some adjustments improving Astra's scores and lowering those of rival Anthropic's models. The post experienced a delayed rollout and was retracted before being republished with revised figures. OpenAI stated the changes were made to ensure the numbers represent the best estimate of model performance.
Related stories
OpenAI Restricts Advanced Cyber Tools in Astra Launch to Prevent Misuse · Corporate earnings
This summary is AI-generated and original to Mobble; the linked article is the authoritative source.
Original headline: “OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch.” Browse more stories.