Loading video...
Video Failed to Load
I tested MTPLX v2 with QWEN 3.6 27B and compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and... show more
15,846 views • 1 month ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
