Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub:

2,252,490 görüntüleme • 6 gün önce •via X (Twitter)

56 Yorum

White Circle profil fotoğrafı
White Circle6 gün önce

Every model we train at White Circle now uses Halo. The same codebase runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL. Technical blog:

White Circle profil fotoğrafı
White Circle6 gün önce

Other training frameworks often require a separate implementation for each model family. In Halo, adding a new model family takes just 100 lines of connecting a wrapper instead of rewriting the model.

White Circle profil fotoğrafı
White Circle6 gün önce

One YAML file contains the model, training method, GPU setup and checkpoint settings. One command starts training.

White Circle profil fotoğrafı
White Circle6 gün önce

Halo supports data, tensor, context, expert and expert-tensor parallelism. It also includes optimized attention, grouped-GEMM and fused loss kernels.

White Circle profil fotoğrafı
White Circle6 gün önce

Halo is not limited to supervised fine-tuning. It supports async RL and training with external environments. We are excited to partner with the @sgl_project team to make it the primary engine for rollouts.

White Circle profil fotoğrafı
White Circle6 gün önce

On @OpenAI gpt-oss-20b, Halo delivered 2.3–2.8x the throughput of stock TRL with less peak memory. Both runs used the same FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations.

White Circle profil fotoğrafı
White Circle6 gün önce

We partnered with @liquidai to add native Halo support for LFM2.5-8B-A1B. On one B300, Halo was up to 20% faster than next best open-source framework.

White Circle profil fotoğrafı
White Circle6 gün önce

We also used Halo to fine-tune @Zai_org GLM-4.7-Flash on 177M tokens of agentic traces. The resulting model improved SWE-rebench-V2 by @nebiusai from 33% to 42%. Halo reached up to 1.63× TRL throughput on the same model precision and data. Model and write-up:

White Circle profil fotoğrafı
White Circle6 gün önce

Thanks for reading to the end! Star us on GitHub: Read more:

Daniel Smidstrup profil fotoğrafı
Daniel Smidstrup6 gün önce

Open source is the way, and it's only going to get more and more relevant in the coming time! :)

White Circle profil fotoğrafı
White Circle6 gün önce

True

leo profil fotoğrafı
leo6 gün önce

insane way to kick off the week -- congrats @whitecircle team

White Circle profil fotoğrafı
White Circle6 gün önce

Lets go!

neural nets. profil fotoğrafı
neural nets.6 gün önce

crazy reading through the docs

Vishal Singh 🥑 profil fotoğrafı
Vishal Singh 🥑6 gün önce

love that this is open source, congrats team ❤️

Catalin profil fotoğrafı
Catalin6 gün önce

Congrats on the awesome work, team! Good to see training infra being open source 👏 Starred the repo to contribute to the growth.

White Circle profil fotoğrafı
White Circle6 gün önce

thanks!

Paul Mit profil fotoğrafı
Paul Mit6 gün önce

great idea you've got my star, guys

giulia profil fotoğrafı
giulia6 gün önce

took me one read to understand exactly who this is for. that almost never happens with infra launches. congrats team!!! starred 👀

Mac profil fotoğrafı
Mac6 gün önce

Peak improvement, I will try it from first hand

sui profil fotoğrafı
sui6 gün önce

so we get Grok 4.7 along with this today? nice.

White Circle profil fotoğrafı
White Circle6 gün önce

Crazy day!

Maxime Labonne profil fotoğrafı
Maxime Labonne6 gün önce

Congrats guys!

Virgile RIETSCH profil fotoğrafı
Virgile RIETSCH6 gün önce

This is huge !!

Levan Kvirkvelia profil fotoğrafı
Levan Kvirkvelia6 gün önce

haloshi framework

𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆 profil fotoğrafı
𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆6 gün önce

wait what this is incredible

White Circle profil fotoğrafı
White Circle6 gün önce

ty Brian! show the repo some love ⭐

ivy profil fotoğrafı
ivy6 gün önce

so excited!!!

White Circle profil fotoğrafı
White Circle6 gün önce

🥰

Ted | unfair.so profil fotoğrafı
Ted | unfair.so6 gün önce

Congrats team!

Soraia profil fotoğrafı
Soraia6 gün önce

Great work! You got my star ⭐️

Nidhi Singh profil fotoğrafı
Nidhi Singh6 gün önce

love that this is open source, congrats!

Rushil Chopra profil fotoğrafı
Rushil Chopra6 gün önce

I have always wanted to train a custom math model, let’s see how this goes!

White Circle profil fotoğrafı
White Circle6 gün önce

Keep us posted, any feedback is welcome

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava6 gün önce

love seeing more training infrastructure get open-sourced. congrats on the release!

Ion profil fotoğrafı
Ion6 gün önce

Congrats! Impressive what you delivered there 👀

Nikita Andersson profil fotoğrafı
Nikita Andersson6 gün önce

@sarapenrique This is fabulous

Hauke 🌱🦀 profil fotoğrafı
Hauke 🌱🦀6 gün önce

Uh training my own models for my products would be pretty neat. I need better hardware 👀

White Circle profil fotoğrafı
White Circle6 gün önce

💚

Prasenjit profil fotoğrafı
Prasenjit6 gün önce

oh this is sick actually, well cooked

Kalash profil fotoğrafı
Kalash6 gün önce

banger stuff 💥

Dan Kulkov profil fotoğrafı
Dan Kulkov6 gün önce

insane start of the week congrats team!

Aditi profil fotoğrafı
Aditi6 gün önce

wait… 2.8x faster than stock trl ??

Vasko profil fotoğrafı
Vasko6 gün önce

congrats to the team!

Lucas Valbuena profil fotoğrafı
Lucas Valbuena6 gün önce

congrats!

White Circle profil fotoğrafı
White Circle6 gün önce

Thank you

Sameer profil fotoğrafı
Sameer6 gün önce

congrats! def keeping an eye on this one.

Suraj Sharma profil fotoğrafı
Suraj Sharma6 gün önce

The interesting part is keeping models in native HuggingFace format while getting that throughput boost.

White Circle profil fotoğrafı
White Circle6 gün önce

Worked hard on this

Suraj Sharma profil fotoğrafı
Suraj Sharma6 gün önce

LFG 🫡🙌

Jackson profil fotoğrafı
Jackson6 gün önce

Really cool. Really unique and impressive numbers :•)

Marco Franzon profil fotoğrafı
Marco Franzon6 gün önce

The only limit now is the hardware. Everyone can now train or finetune a model without rewriting it for the training phase. Awesome!

White Circle profil fotoğrafı
White Circle6 gün önce

thanks Marco!

Marco Franzon profil fotoğrafı
Marco Franzon6 gün önce

Great work !

Deni profil fotoğrafı
Deni6 gün önce

Congrats on the launch team! Starred.

Ole Lehmann profil fotoğrafı
Ole Lehmann6 gün önce

happy launch day guys"

Benzer Videolar