06Serve Kimi K3, a 2.8 trillion parameter model. What is the minimum viable configuration?▼hard★ EssentialNewTogether AIFireworks AIBaseten4 repliesunlockedTwo point eight trillion parameters fits in one node, because the experts ship in a four-bit format. The footprint arithmetic that reproduces the project's own published hardware minimum, the hybrid cache that needs specific engine support, and why batch-one throughput is a tenth of the bound.Open full answer →
27A model interleaves linear-attention and full-attention layers. What changes about serving it?▼hardNewTogether AIFireworks AIBaseten4 replies◆ premiumMost layers stop having a growing cache, which changes capacity planning by a large factor and demands a cache manager most engines did not have. What the memory model becomes, what the engine must implement, and the two features that were disabled by default while it settled.Open full answer →