header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Information keeps getting revised, and large models get confused: even the best-performing GPT-5.5 can only answer 59.3% correctly.

Beating AI News Flash: Teams including Tencent Hunyuan have released EvolveScaler, specifically designed to test whether large models can keep up with constantly changing information. It continuously injects modifications, retractions, additions, and invalidations into long texts, then asks the model to answer questions based on the latest state.


For example, the model is first shown 40 days of game records, then asked: "If you didn't fight that mini-boss on day 7, could you still win in the end?" This would cascade into changes in subsequent equipment, HP, and battle outcomes. The model has to re-derive dozens of days starting from day 7, and the answer is not readily available in the original text.


The team designed 117 task types and 159 question types, with a maximum of approximately 1,200 events. The 14 model configurations include GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Preview Pro, GLM-5.2, and Qwen3.5 Plus. In the hardest tier, the median score was only 11.3%; the best-performing GPT-5.5-xhigh scored just 59.3%.


This dataset can also be used to train models. After an internal A3B model was trained on 6,000 samples of this type of data, it improved across all 8 external tests, with an average gain of 5.25 points.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish