header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

这个信息看起来像是某个AI模型评测榜单(可能叫Hermes或类似名称)的更新,但我需要提醒几点: **我无法验证这个具体榜单的真实性**,因为: - 我的知识截止日期是2025年5月,之后发生的事情我无法确认 - "Opus 5.5"和"Astra"这些具体版本号/名称,我无法核实是否真实存在 - 这类榜单可能来自不同的评测机构,标准各异 **如果你是在分享一条新闻**,我可以帮你分析这类榜单通常意味着什么: 1. **排名第一 ≠ 综合最强**:很多榜单只测特定维度(如推理、代码、长文本),Opus 5.5排第一可能只是在某个基准上领先 2. **"第二却最贵"的解读**: - 定价通常反映的是**推理成本**(模型规模、算力消耗),不完全等于能力 - 如果Astra第二但最贵,可能说明它参数量大、推理开销高,或者厂商定位高端 - 也可能是"性价比"角度:花更多钱但没拿到第一,用户会觉得不划算 3. **榜单的局限性**: - 评测集可能被训练数据污染 - 实际使用体验和跑分差距很大 - 厂商可能针对榜单优化 **如果你能补充**:这个

动察 Beating AI News Flash: Nous Research has launched the Hermes Index, a model leaderboard for Hermes Agent, mainly used to compare the performance of different models within Hermes. All models use the same Hermes Agent operating framework, and those that support adjustable reasoning intensity are uniformly set to high. The leaderboard takes the average score of four tests—Hermes Bench, Terminal-Bench 4, Terminal-Bench Science, and SkillsBench—and also tracks the average cost per task.


Among them, Hermes Bench is a new test created by Nous Research itself, with a total of 150 tasks and 25 scenario categories, covering Skills, research, charts and creation, memory, tool calling, and security. The Agent directly handles the workspace and real files, and scoring is ultimately based on the generated files, task status, and tool usage records.


The first version tested 14 models in total. Claude Opus 5.5 ranked first with 63.31 points, GPT 6 Astra ranked second with 56.25 points, and Claude Sonnet 5.5 ranked third with 53.14 points. Astra cost an average of $11.61 per task, the highest on the entire leaderboard. Opus 5.5 was $4.99, and Sonnet 5.5 was $2.82; some of their data is still marked by the officials as provisional.


Among lower-priced models, DeepSeek V4.1 Flash scored 36.91 points, with an average cost of $0.259 per task; GPT 6 Luna scored 33.89 points, at $0.141 per task. Both entered the Pareto frontier of cost and performance, meaning there is no model that is cheaper and also scores higher.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish