正在加载视频...
视频加载失败
ollama run llama3.1:405b Tested in with AMD MI300X 🤯
98,184 次观看 • 2 年前 •via X (Twitter)
10 条评论

@TensorWaveCloud @AMD Does it run with CPU offloading, as the AMD MI300X only brings 192 GB of memory? Asking, because your model pages states that the 405B model of Llama 3.1 has a size of 231 GB?

@TensorWaveCloud @AMD 231GB of memory is needed to run it in 4bit quantization. 🥹 That is only a magnitude more than we currently have in the best consumer GPUs. 🤣

@TensorWaveCloud @AMD Wai- that is actually insane I am testing Llama 3.1 8B, it's so fast! @Meta really cooked with this one!

@TensorWaveCloud @AMD FYI! Unless you’ve patched L3.1 on your own, the generations for all L3.1 405B, 70B & 8B are not accurate and off:

@TensorWaveCloud @AMD 🙏 Thank you! We are upstreaming the patch! 🙏🙏🙏 Please support open-source!! When you get early access again, please help make the fixes! ❤️❤️❤️

@TensorWaveCloud @AMD I'm going to need more VRAM

@TensorWaveCloud @AMD Amazing

@TensorWaveCloud @AMD Can you run crysis?

@TensorWaveCloud @AMD Brb gonna get a giga cluster to run this

@TensorWaveCloud @AMD 👍

