header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Arena Launches AutoEval: AI First Simulates Human Scoring, New Model Can Generate Scores in Just 1 Hour

According to DolphinBeat monitoring, the large-scale model benchmarking platform Arena has launched AutoEval, which uses a reward model (a rating model learning from human preferences) to simulate user voting and quickly generate leaderboard scores for new models. With the release of a new model, instead of waiting for days to accumulate human votes, it can receive an estimated ranking in less than 1 hour. This score is labeled as AutoEval and will later be corrected by human voting results.

This system is trained on data from millions of real user preference sets. During evaluation, AutoEval compares the answer quality of different models and calculates scores according to Arena's original leaderboard rules. Currently, it supports evaluation of text, vision, audio, and code models.

Arena's backtesting shows that the correlation of AutoEval with subsequent human rankings is over 0.98. When the true score difference between two models exceeds 10 points, its accuracy in determining the winner exceeds 90%; when the difference exceeds 15 points, it reaches 100%. However, when the model gap is less than 5 points, it is still difficult to consistently differentiate.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish