动察 Beating AI News Flash: Nous Research has launched the Hermes Index, a model leaderboard for Hermes Agent, mainly used to compare the performance of different models within Hermes. All models use the same Hermes Agent operating framework, and those that support adjustable reasoning intensity are uniformly set to high. The leaderboard takes the average score of four tests—Hermes Bench, Terminal-Bench 4, Terminal-Bench Science, and SkillsBench—and also tracks the average cost per task.
Among them, Hermes Bench is a new test created by Nous Research itself, with a total of 150 tasks and 25 scenario categories, covering Skills, research, charts and creation, memory, tool calling, and security. The Agent directly handles the workspace and real files, and scoring is ultimately based on the generated files, task status, and tool usage records.
The first version tested 14 models in total. Claude Opus 5.5 ranked first with 63.31 points, GPT 6 Astra ranked second with 56.25 points, and Claude Sonnet 5.5 ranked third with 53.14 points. Astra cost an average of $11.61 per task, the highest on the entire leaderboard. Opus 5.5 was $4.99, and Sonnet 5.5 was $2.82; some of their data is still marked by the officials as provisional.
Among lower-priced models, DeepSeek V4.1 Flash scored 36.91 points, with an average cost of $0.259 per task; GPT 6 Luna scored 33.89 points, at $0.141 per task. Both entered the Pareto frontier of cost and performance, meaning there is no model that is cheaper and also scores higher.

