Dynamic Observation Beating AI News: Since OpenAI released GPT-6 Astra on September 3, various model benchmark data has been continuously adjusted. Some modifications have improved Astra's performance, while scores of some competitor models have declined, leading to external scrutiny of AI "leaderboard manipulation" and evaluation transparency.
Notably, Astra's hallucination rate was initially lowered from 4.2% to 2%, GPT-5.6 Sol dropped from 12.2% to 9.4%, and then both rebounded to 4.2% and 12.2%, respectively. In mathematical evaluations, Anthropic's Fable 5.1 score temporarily decreased from 87.8% to 78%, but has now risen to 83%; GPT-5.6 Sol decreased from 83% to 80.5%, then recovered to 83%.
Furthermore, Astra's performance in the ARC-AGI-3 evaluation increased from the pre-release draft's 98.6% to a final page score of 99.99%, while its programming evaluation score saw a slight uptick from 57.7% to 57.9%. OpenAI stated that evaluation results can be influenced by factors such as model versions, tool configurations, reasoning levels, and test runs. This adjustment aims to ensure that the data more accurately reflects the model's optimal performance.
However, Stanford University researchers suggest that frequent re-running of evaluations may involve a practice known as "Benchmaxxing," where test conditions are adjusted to maximize benchmark scores. Industry insiders point out that as competition among AI models intensifies, evaluation data has become a crucial tool for measuring model capabilities and market competitiveness. Enhancing the transparency and reproducibility of benchmark testing is increasingly receiving more attention.

