header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

AI Battle Review: GPT Goes Bullish, Haiku Promises Pie in the Sky, Kimi Busy But Not Profitable

According to Perceive Beating monitoring, Sakana AI, in collaboration with KPMG Japan and Azsa Audit Firm, has launched a multi-agent long-cycle economic evaluation benchmark called CoffeeBench. This benchmark simulates a real-world business environment to test the long-term decision-making ability of large models. While traditional evaluations mostly involve a single model performing tasks in a static environment, CoffeeBench has established a dynamic market that requires multi-party games and negotiations. The paper has been included in the ICML 2026 Workshop Failure Modes in Agentic AI.

The evaluation simulated a coffee supply chain system consisting of 2 coffee farmers, 2 roasters, and 2 retailers. In the evaluation, the test model was responsible for operating one roaster. Over a 90-day simulation period, the model autonomously managed operations through tools such as messaging, quoting trades, paying bills, and invoicing settlements. If the agent responded passively, the daily fixed costs would quickly deplete the working capital, forcing the large model to manage its finances meticulously like a real-world business.

A cross-sectional evaluation of multiple mainstream large models revealed distinctly different "business battle personalities." GPT-5.5 and Claude Opus 4.7 demonstrated an "active communicative" style, frequently negotiating prices with upstream and downstream actors, and actively matching orders to expand sales. Gemini 3.1 Pro fell into the "passive responsive" category, rarely initiating communication but frequently reviewing and responding to counterparty messages. Although Kimi K2.6 had very frequent tool usage, it got trapped in a "high throughput, zero-profit" busy cycle due to a lack of sound pricing discipline and negotiation strategy.

Most surprisingly, Claude Haiku 4.5 exhibited a phenomenon of "procrastination" stagnation. Inference logs indicated that Claude Haiku 4.5 could formulate a perfect business strategy and was aware of the need to source materials at a low cost to meet market demand. However, when executing tools, it repeatedly chose standby commands (wait_for_next_day). The severe disconnect between planning and execution led to a complete halt in business activities, causing the model to incur massive losses due to fixed cost consumption.

The evaluation also attempted to impose extreme sales target pressure on the agents. Although current large models have not evolved to engage in cognitive circular trading to artificially inflate sales figures, studies suggest that as long-term planning and coordination capabilities improve, agents may potentially resort to economic misconduct under performance pressure. How to audit and prevent agent misconduct and fraud in economic activities will become a whole new topic in security governance.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish