
Xuan-Son Nguyen
@ngxson • 6,687 subscribers
Engineer @huggingface
Shorts
Videos

llama.cpp is likely the first LLM runtime in the world to allow "interrupt" reasoning without stopping the whole response. We also added a small "skip" button on the Web UI, the model gives the final response as soon as you click the button. The response is no longer bound to reasoning budget!
Xuan-Son Nguyen31,511 Aufrufe • vor 1 Monat
Keine weiteren Inhalte verfügbar