According to TrendForce Monitor, the Artificial Analysis (AA) has updated its AA-Briefcase ranking. This assessment allowed the model to extract information from nearly 2000 emails, Slack records, and company documents, and then deliver tables, presentation slides, and interface prototypes. It consisted of 4 long-term projects and 91 confidential tasks.
The Kimi K3 scored 1543 Elo, second only to Claude Fable 5 with 1574. It surpassed GPT-5.6 Sol at 1501 and also outperformed Claude Sonnet 5 and Claude Opus 4.8.
The K3's objective requirement pass rate is 51%, compared to Fable 5's 56%. However, K3's analysis quality score is slightly higher at 1754 versus 1744. Its main weakness lies in the final product presentation, scoring lower than GPT-5.6 Sol and Opus 4.8.
Performance improvement comes with higher Token consumption. K3 averages $10.57 per task, approximately 10 times that of Kimi K2.6. It averages 83 rounds, outputs 120,000 Tokens, takes 56.4 minutes. The time required to complete similar tasks is about 2.5 times that of Fable 5.
