| |
DeepSeek V4 Flash on a Single AMD MI300X
A developer has successfully configured DeepSeek V4 Flash to run on a single AMD MI300X GPU in production, achieving 168.6 tokens/second for single-stream decoding and 542 tokens/second across 8 concurrent streams without additional quantization. The setup required custom fixes for FP8 format compatibility, MoE routing, and kernel tuning specific to MI300X hardware, as the official vLLM recipe only targets NVIDIA and newer AMD GPUs. The MI300X's 192 GB of memory enables the full 304-billion-parameter model to fit in VRAM while supporting a 256K context window and burst loads up to 64 concurrent streams.
Read Full Article →
← More Tech news