World
Github Repository Details How to Run Deepseek V4 Flash on a Single AMD Mi300x
The 304B-parameter checkpoint runs without extra quantization or offload. The MI300X's 192 GB HBM3 and 5.3 TB/s bandwidth enable single-GPU deployment, with roughly half the list price of an H100 SXM5. Key fixes address FP8 format incompatibilities, MoE routing bugs, and kernel shapes. The stack uses a digest-pinned vLLM RO Cm nightly and AITER 0.1.19. Un-cached prefill hits 7.9–8.5K tok/s depending on scheduler budget; production profile with a 2,048-token budget delivers 6,988–7,019 tok/s. Warm recall of 380K cached tokens takes 0.64–2.65 s after a 120–125 s cold prefill. The repository is Apache-2.0 licensed, with the model MIT-licensed.
Source: Hacker News






