Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub:

2,252,490 Aufrufe • vor 6 Tagen •via X (Twitter)

56 Kommentare

Profilbild von White Circle
White Circlevor 6 Tagen

Every model we train at White Circle now uses Halo. The same codebase runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL. Technical blog:

Profilbild von White Circle
White Circlevor 6 Tagen

Other training frameworks often require a separate implementation for each model family. In Halo, adding a new model family takes just 100 lines of connecting a wrapper instead of rewriting the model.

Profilbild von White Circle
White Circlevor 6 Tagen

One YAML file contains the model, training method, GPU setup and checkpoint settings. One command starts training.

Profilbild von White Circle
White Circlevor 6 Tagen

Halo supports data, tensor, context, expert and expert-tensor parallelism. It also includes optimized attention, grouped-GEMM and fused loss kernels.

Profilbild von White Circle
White Circlevor 6 Tagen

Halo is not limited to supervised fine-tuning. It supports async RL and training with external environments. We are excited to partner with the @sgl_project team to make it the primary engine for rollouts.

Profilbild von White Circle
White Circlevor 6 Tagen

On @OpenAI gpt-oss-20b, Halo delivered 2.3–2.8x the throughput of stock TRL with less peak memory. Both runs used the same FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations.

Profilbild von White Circle
White Circlevor 6 Tagen

We partnered with @liquidai to add native Halo support for LFM2.5-8B-A1B. On one B300, Halo was up to 20% faster than next best open-source framework.

Profilbild von White Circle
White Circlevor 6 Tagen

We also used Halo to fine-tune @Zai_org GLM-4.7-Flash on 177M tokens of agentic traces. The resulting model improved SWE-rebench-V2 by @nebiusai from 33% to 42%. Halo reached up to 1.63× TRL throughput on the same model precision and data. Model and write-up:

Profilbild von White Circle
White Circlevor 6 Tagen

Thanks for reading to the end! Star us on GitHub: Read more:

Profilbild von Daniel Smidstrup
Daniel Smidstrupvor 6 Tagen

Open source is the way, and it's only going to get more and more relevant in the coming time! :)

Profilbild von White Circle
White Circlevor 6 Tagen

True

Profilbild von leo
leovor 6 Tagen

insane way to kick off the week -- congrats @whitecircle team

Profilbild von White Circle
White Circlevor 6 Tagen

Lets go!

Profilbild von neural nets.
neural nets.vor 6 Tagen

crazy reading through the docs

Profilbild von Vishal Singh 🥑
Vishal Singh 🥑vor 6 Tagen

love that this is open source, congrats team ❤️

Profilbild von Catalin
Catalinvor 6 Tagen

Congrats on the awesome work, team! Good to see training infra being open source 👏 Starred the repo to contribute to the growth.

Profilbild von White Circle
White Circlevor 6 Tagen

thanks!

Profilbild von Paul Mit
Paul Mitvor 6 Tagen

great idea you've got my star, guys

Profilbild von giulia
giuliavor 6 Tagen

took me one read to understand exactly who this is for. that almost never happens with infra launches. congrats team!!! starred 👀

Profilbild von Mac
Macvor 6 Tagen

Peak improvement, I will try it from first hand

Profilbild von sui
suivor 6 Tagen

so we get Grok 4.7 along with this today? nice.

Profilbild von White Circle
White Circlevor 6 Tagen

Crazy day!

Profilbild von Maxime Labonne
Maxime Labonnevor 6 Tagen

Congrats guys!

Profilbild von Virgile RIETSCH
Virgile RIETSCHvor 6 Tagen

This is huge !!

Profilbild von Levan Kvirkvelia
Levan Kvirkveliavor 6 Tagen

haloshi framework

Profilbild von 𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆
𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆vor 6 Tagen

wait what this is incredible

Profilbild von White Circle
White Circlevor 6 Tagen

ty Brian! show the repo some love ⭐

Profilbild von ivy
ivyvor 6 Tagen

so excited!!!

Profilbild von White Circle
White Circlevor 6 Tagen

🥰

Profilbild von Ted | unfair.so
Ted | unfair.sovor 6 Tagen

Congrats team!

Profilbild von Soraia
Soraiavor 6 Tagen

Great work! You got my star ⭐️

Profilbild von Nidhi Singh
Nidhi Singhvor 6 Tagen

love that this is open source, congrats!

Profilbild von Rushil Chopra
Rushil Chopravor 6 Tagen

I have always wanted to train a custom math model, let’s see how this goes!

Profilbild von White Circle
White Circlevor 6 Tagen

Keep us posted, any feedback is welcome

Profilbild von AshutoshShrivastava
AshutoshShrivastavavor 6 Tagen

love seeing more training infrastructure get open-sourced. congrats on the release!

Profilbild von Ion
Ionvor 6 Tagen

Congrats! Impressive what you delivered there 👀

Profilbild von Nikita Andersson
Nikita Anderssonvor 6 Tagen

@sarapenrique This is fabulous

Profilbild von Hauke 🌱🦀
Hauke 🌱🦀vor 6 Tagen

Uh training my own models for my products would be pretty neat. I need better hardware 👀

Profilbild von White Circle
White Circlevor 6 Tagen

💚

Profilbild von Prasenjit
Prasenjitvor 6 Tagen

oh this is sick actually, well cooked

Profilbild von Kalash
Kalashvor 6 Tagen

banger stuff 💥

Profilbild von Dan Kulkov
Dan Kulkovvor 6 Tagen

insane start of the week congrats team!

Profilbild von Aditi
Aditivor 6 Tagen

wait… 2.8x faster than stock trl ??

Profilbild von Vasko
Vaskovor 6 Tagen

congrats to the team!

Profilbild von Lucas Valbuena
Lucas Valbuenavor 6 Tagen

congrats!

Profilbild von White Circle
White Circlevor 6 Tagen

Thank you

Profilbild von Sameer
Sameervor 6 Tagen

congrats! def keeping an eye on this one.

Profilbild von Suraj Sharma
Suraj Sharmavor 6 Tagen

The interesting part is keeping models in native HuggingFace format while getting that throughput boost.

Profilbild von White Circle
White Circlevor 6 Tagen

Worked hard on this

Profilbild von Suraj Sharma
Suraj Sharmavor 6 Tagen

LFG 🫡🙌

Profilbild von Jackson
Jacksonvor 6 Tagen

Really cool. Really unique and impressive numbers :•)

Profilbild von Marco Franzon
Marco Franzonvor 6 Tagen

The only limit now is the hardware. Everyone can now train or finetune a model without rewriting it for the training phase. Awesome!

Profilbild von White Circle
White Circlevor 6 Tagen

thanks Marco!

Profilbild von Marco Franzon
Marco Franzonvor 6 Tagen

Great work !

Profilbild von Deni
Denivor 6 Tagen

Congrats on the launch team! Starred.

Profilbild von Ole Lehmann
Ole Lehmannvor 6 Tagen

happy launch day guys"

Ähnliche Videos