
Xuan-Son Nguyen
@ngxson • 6,687 subscribers
Engineer @huggingface
Shorts
Videos

Real-time webcam demo with Hugging Face SmolVLM and ggml llama.cpp server. All running locally on a Macbook M3
Xuan-Son Nguyen973,865 görüntüleme • 1 yıl önce

llama.cpp is likely the first LLM runtime in the world to allow "interrupt" reasoning without stopping the whole response. We also added a small "skip" button on the Web UI, the model gives the final response as soon as you click the button. The response is no longer bound to reasoning budget!
Xuan-Son Nguyen31,511 görüntüleme • 1 ay önce
Daha fazla içerik yok.