← Back to issue16 / 22 · Week of Jul 27, 2026

Deltafin makes local MoE serving an I/O tradeoff

The Deltafin repo exposes an OpenAI-compatible server for Kimi K3 and offers a full local download or on-demand expert streaming. Its README lists roughly 1.7 TB for the full mode and about 215 GB for streaming. Why it matters: A local endpoint does not remove the operating constraint; it moves it into disk, network, and token latency. The streaming path lowers the storage commitment but can take 3+ minutes to fetch uncached experts.

Try this: Inspect the hardware, disk, network, and latency table before installing. Treat the maintainer's speed figures as a starting point for a reproducible test, not a capacity plan.

GitHub 608 stars · Aug 2verify ↗
Source
gavamedia/deltafin
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime