header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

AI Intern Surpasses Humans After 7 Months: Fable 5.1 Achieves 72% Success Rate in Accounting Roles

Beating AI News Flash: AI Agent startup NeoCognition has released ApprenticeBench, specifically designed to test whether AI can learn a job on its own after onboarding, just like a new hire. In the first version, the Agent joins a simulated California construction company to handle accounts payable. It first reads six months of historical bills, the company handbook, and Odoo tutorials, then processes 100 bills over 7 consecutive months, during which it also encounters rule changes and supervisor feedback.


Fable 5.1 ultimately achieved a success rate of 72%, GPT-6 Astra reached 68%, and the best human test taker scored only 51%. In addition to entering invoices, the tasks include spotting errors, assigning cost codes, contacting vendors, going through approvals, and learning the company's internal rules from historical records and feedback. The two models scored even slightly higher when operating the GUI directly than when calling APIs, and the performance loss commonly seen with Computer Use in the past has largely disappeared.


However, "surpassing humans" holds only for this one simulated position, and the human sample consisted of only 2 people. The cost is also higher: Fable 5.1 averages $18.23 per task, while humans average about $7.21. Humans get faster the more they do, while the Agent instead slows down because it remembers more and more.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish