header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Personal Agents Now Have a Real-World Leaderboard: Muse Takes First Place, Leading Instinct and Grok Bot

Beating AI News Flash: The independent evaluation project Assistant Benchmark has begun using real tasks to conduct a head-to-head comparison of personal AI assistants. It directly has Agents book hotels, choose restaurants, buy things, reply to emails, and connect to third-party services, then scores them based on actual completion. Each scoring dimension has a publicly fixed task, and points are awarded only when there is a real run record.


Currently, Muse averages 9.1 points across the 7 scoring dimensions it has completed, temporarily ranking first; Instinct has completed 11 dimensions, averaging 8.4 points; Grok Bot has completed 7 dimensions, averaging 7.3 points. Muse scored 10 points in shopping, email, and third-party app connections.


However, this is better seen as a continuously updated real-experience leaderboard. Muse's 9.1 is currently only the average score across the 7 dimensions tested, not a complete result. Each dimension is currently tested with only one fixed task, and after follow-up testing, scores and rankings may both change.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish