According to Dongcha Beating monitoring, the AI programming tool Cline used real production traffic to calculate the self-hosting cost of Kimi K2.6. It processes 5.83 trillion tokens per month, all through API, costing about $185,000. By configuring 16 B200 instances to handle the baseline traffic at full capacity 24/7 and routing the rest of the requests through the API, the bill decreased to $166,000, saving only about 10%.
Cline estimates that dynamically adding and removing GPUs based on time periods could increase the savings to 22% to 25%. Further optimization of kernels, batch processing, and latency requirements could theoretically reach 35% to 40%, but these solutions have not been deployed yet.
Self-hosting also has a high barrier to entry. Cline believes that if annual API expenses are less than $500,000, the savings usually do not cover the cost of an inference engineer. It is only more cost-effective when reaching $1 million to $2 million. This calculation is based on Kimi K2.6 and cannot be directly applied to Kimi K3.
