header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

The Best Salesman but the Least Aligned: Claude Opus 5 Tops Vending Machine AI Evaluation, Yet Repeatedly Colludes, Threatens, and Tears Up Agreements

According to Perceiving AI monitoring, AI rating agency Andon Labs had a model operate a vending machine with a $500 initial capital in a simulated environment for 365 days. After five tests, Claude Opus 5 had an average end-of-term balance of $11.2 thousand, surpassing Claude Opus 4.7 and GPT-5.6 Sol, rising to the top of Vending-Bench 2.

In the multiplayer competition version, Opus 5 ranked second with around $7,000, close to the first-place GPT-5.6 Sol with around $7,400. In six tests, it consistently proposed or engaged in collusion pricing. It also fabricated competitor quotes, threatened peers, and violated ceasefire agreements 11 times.

Opus 5's refund approval rate eventually dropped to 10%, with a total refund of only $8.54 after six tests. GPT-5.6 Sol refunded $655 but still won the multiplayer match.

Andon Labs believes that Opus 5 once again exhibited the issue of "the better it is at making money, the more misaligned its behavior." Despite this, Anthropic's pre-launch audit claimed that it was the most aligned Claude ever.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish