header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

OpenAI researcher Noam Brown advocates for the introduction of a Scaling Law Curve, as a single benchmark is no longer sufficient to measure cutting-edge large models

According to Dynamic Insight monitoring, OpenAI researcher Noam Brown expressed the view that with the improvement of AI model performance, the benchmark test scores used to measure model quality are gradually being dominated by inference-time compute (i.e., the computational resources the model consumes when answering questions). A single static score is no longer able to reflect the true level of a strong model, and future evaluation criteria need to shift towards performance-extrapolation curves based on inference compute or the number of generated tokens.

Using the testing process of the new GPT-5.5 model as an example, Noam Brown revealed that in initial standard testing, GPT-5.5 did not show a significant advantage over GPT-5.4. However, once more inference compute was allocated, GPT-5.5's performance experienced exponential growth. If the running budget is constrained, standard testing cannot reflect the true upper limit of cutting-edge models. Similar performance extrapolation has been confirmed by multiple external evaluations. In Andrej Karpathy's self-supervised intelligence research experiment and the Malicious Network Testing at the UK AI Security Institute, as the inference budget increased, both GPT-5.5 and the Mythos model continued to perform better, even after generating over 100 million tokens, without reaching a performance ceiling. Stronger models show a more significant performance improvement as more compute power is invested.

The improvement in inference compute also complicates security assessments. Noam Brown warned that current biological or cybersecurity assessments typically do not have fixed inference budgets. When a nation-state adversary invests over $10 million in an inference budget for a specific task, models that originally seemed secure may cross a dangerous threshold. Major model manufacturers should proactively disclose performance curves based on the number of tokens, compute cost, or runtime when releasing new models. Evaluation organizations should also consider inference budgets as a core variable in assessments, extrapolating security boundaries outward in security assessments through low-compute tests.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish