DeepSeek V4 Flash on a Single AMD MI300X
- majke - 34276 sekunder sedanI don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
- GTP - 23599 sekunder sedanStrange that in the prior art they didn't list DwarfStar, as it is able to run the same model (probably quantized differently though) in less memory. Maybe the author isn't aware of it?
- fergusfinn - 20916 sekunder sedannice! i think the higher HBM on Mi300x is really useful for this kind of thing
we did some work on this for 2xMi300x (kindly referenced in the readme) https://blog.doubleword.ai/deepseek-v4-flash-mi300x. https://hotaisle.xyz/quick-start hotaisle is great for getting Mi300x to experiment with
- Tepix - 24441 sekunder sedanUnfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB.
Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.
- WhitneyLand - 21551 sekunder sedanAnother headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.
Dumbed down quantization?
No. Full intended inference weights preserved, so far so good.
Slow performance?
No again. Looks like you could get over 150 tokens/second.
Give up context window size?
Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.
- xorfish - 26637 sekunder sedanThis is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.
- PrimeAli - 8440 sekunder sedanGreat
- sylware - 23675 sekunder sedanIs their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?
BTW, how many many params open weight frontier models have? A few teras, 100s of teras?
- pop3zxcv - 9570 sekunder sedan[dead]
- jkwang - 33106 sekunder sedan[flagged]
- hn0tdqaek4 - 26594 sekunder sedan[dead]
Nördnytt! 🤓