
Xuan-Son Nguyen
@ngxson • 6,843 subscribers
Engineer @huggingface
Shorts
Videos

Real-time webcam demo with Hugging Face SmolVLM and ggml llama.cpp server. All running locally on a Macbook M3
Xuan-Son Nguyen974,221 просмотров • 1 год назад

llama.cpp is likely the first LLM runtime in the world to allow "interrupt" reasoning without stopping the whole response. We also added a small "skip" button on the Web UI, the model gives the final response as soon as you click the button. The response is no longer bound to reasoning budget!
Xuan-Son Nguyen31,511 просмотров • 3 месяцев назад
Больше нет контента для загрузки