According to MotionBeat Monitoring, the AI Agent infrastructure company Composio integrated the same Kimi K3 with Kimi Code, Hermes, and Claude Code, completing 28 identical tasks. The success count of the three execution frameworks only differed by 2 tasks, but the token usage and costs varied significantly.
Kimi Code completed 22 tasks, Hermes completed 21 tasks, and Claude Code completed 20 tasks. The median token usage per task was 61,000, 67,000, and 340,000 respectively. Composio estimated that the average cost per task was around $0.22, $0.28, and $2. Individual task token usage varied by up to 30 times.
Hermes exhibited the fastest speed, with a median task completion time of 179 seconds. Kimi Code took 297 seconds, and Claude Code took 348 seconds.
Another study by Writer yielded similar results. Researchers replaced only the execution framework in 22 enterprise tasks across 6 models, leading to a 38% reduction in token usage, a 41% reduction in individual task costs, a 44% decrease in time taken, while maintaining a consistent level of completion quality.
