According to DolphinBeat monitoring, the large-scale model benchmarking platform Arena has launched AutoEval, which uses a reward model (a rating model learning from human preferences) to simulate user voting and quickly generate leaderboard scores for new models. With the release of a new model, instead of waiting for days to accumulate human votes, it can receive an estimated ranking in less than 1 hour. This score is labeled as AutoEval and will later be corrected by human voting results.
This system is trained on data from millions of real user preference sets. During evaluation, AutoEval compares the answer quality of different models and calculates scores according to Arena's original leaderboard rules. Currently, it supports evaluation of text, vision, audio, and code models.
Arena's backtesting shows that the correlation of AutoEval with subsequent human rankings is over 0.98. When the true score difference between two models exceeds 10 points, its accuracy in determining the winner exceeds 90%; when the difference exceeds 15 points, it reaches 100%. However, when the model gap is less than 5 points, it is still difficult to consistently differentiate.
