header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

WASTE Inference Engine Open Sourced: Running Kimi K3 on a 64GB MacBook

According to Dotion Beating monitoring, the edge database company SQLite AI has open-sourced the WASTE inference engine. It allows the preservation of all layers and the expert Kimi K3 to run on a 64GB MacBook Pro. The converted model is approximately 1TB in size, with a speed of 0.49 to 0.54 Tokens per second.

WASTE keeps around 27GB of the model backbone in memory. Over 80,000 experts are placed on the built-in SSD. For each generated Token, only the currently invoked expert is read from the disk.

This version does not involve distillation, pruning, or expert removal. However, the expert weights have been re-quantized to 3 bits, while the model backbone adopts 4-bit and 8-bit.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish