Dynamic Insights: Beating AI News Flash: Wafer, an AI inference startup with only 8 team members, has completed a $40 million Series A funding round, valuing the company at over $2 billion. The company had previously turned down acquisition offers from several cloud providers and inference service companies.
Wafer's inference business reached $8 million in ARR in just 3 months. Instead of designing chips, Wafer focuses on helping models run on existing chips faster and more cost-effectively. Wafer enables the AI agent to automatically adjust the model, inference engine, kernel, cache, quantization, and scheduling, and then find the optimal solution for different hardware such as NVIDIA and AMD. Previously, much of this work had to be done manually by inference performance engineers, but Wafer aims to hand this part over to AI.
In Wafer's self-tests, GLM 5.2 ran on AMD MI355X with a throughput of about 80% of NVIDIA B200, at a cost of less than half. Current customers include Vercel and Inworld AI. Vercel has also independently released test results, stating that Wafer's throughput for running GLM 5.2 is approximately twice that of other serverless providers.
After the funding round, Wafer immediately began to expand its team. They have opened positions in four categories: technology, growth, CEO's office, and go-to-market strategy, all requiring presence in San Francisco five days a week. The technology positions offer a base salary of $250,000 plus equity.

