header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Prime Intellect inference platform goes live: internally already running nearly 1 trillion tokens per day.

动察 Beating AI News Flash: Prime Intellect releases Prime Inference, an open-source model inference service.


Users can call models directly on demand; if a task requires long-term stable operation, they can also lock in a portion of inference compute in advance. The interface is compatible with the OpenAI API, and the first publicly deployed model is GLM-5.3.


This system has already been used internally at Prime Intellect. The company says it processes nearly 1 trillion tokens per day, used for reinforcement learning, synthetic data, model evaluation, and Coding Agent. Since January of this year, Prime Intellect has also been running large-scale deployments for customers.


Prime Inference focuses on optimizing long-context workloads for Agents. It separates prompt processing and response generation onto different GPUs, reducing p90 inter-token latency by nearly 40% in tests. Prime Intellect has also compressed the KV cache, increasing the number of tokens each decoder can cache from 1.09 million to 1.63 million under the same VRAM, an improvement of about 50%.


Prime Intellect previously provided training infrastructure such as reinforcement learning, evaluation, and sandboxes. With the addition of inference services, it can connect model training and actual deployment into the same platform, and data generated from actual operation can also continue to be used for subsequent training.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish