Wësche's banner
Wësche's profile picture

Wësche

@WescheNex1q1,588 subscribers

Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston

Shorts

Qwen3.6-27B dense quality ladder. Same tasks. Same grader. Same day. How much does compression cost? FP8 - 83.3 · 29GB NVFP4 - 83.0 · ~15GB BF16 - 81.9 · 56GB GPTQ-Pro - 79.4 · 13GB IQ2_XXS - 79.0 · 9.4GB Who wins where: 🥇 Raw quality → FP8 (83.3) 🥇 Practical default → NVFP4 (83.0 at ~15GB) 🥈 Full precision → BF16 is NOT better (81.9) ❌ "Just use any 4-bit" → GPTQ-Pro loses 3.6 pts vs NVFP4 🛟 Emergency tiny → IQ2 still holds at 79.0 Overall winner: If you care about the absolute number: FP8 If you care about running it on a real box: NVFP4 is the overall pick. Same quality class as full precision. Half the size of FP8. ~4× smaller than BF16.

Qwen3.6-27B dense quality ladder. Same tasks. Same grader. Same day. How much does compression cost? FP8 - 83.3 · 29GB NVFP4 - 83.0 · ~15GB BF16 - 81.9 · 56GB GPTQ-Pro - 79.4 · 13GB IQ2_XXS - 79.0 · 9.4GB Who wins where: 🥇 Raw quality → FP8 (83.3) 🥇 Practical default → NVFP4 (83.0 at ~15GB) 🥈 Full precision → BF16 is NOT better (81.9) ❌ "Just use any 4-bit" → GPTQ-Pro loses 3.6 pts vs NVFP4 🛟 Emergency tiny → IQ2 still holds at 79.0 Overall winner: If you care about the absolute number: FP8 If you care about running it on a real box: NVFP4 is the overall pick. Same quality class as full precision. Half the size of FP8. ~4× smaller than BF16.

24,599 просмотров

Videos

Больше нет контента для загрузки