正在加载视频...

视频加载失败

Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub:

2,252,490 次观看 • 6 天前 •via X (Twitter)

56 条评论

White Circle 的头像
White Circle6 天前

Every model we train at White Circle now uses Halo. The same codebase runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL. Technical blog:

White Circle 的头像
White Circle6 天前

Other training frameworks often require a separate implementation for each model family. In Halo, adding a new model family takes just 100 lines of connecting a wrapper instead of rewriting the model.

White Circle 的头像
White Circle6 天前

One YAML file contains the model, training method, GPU setup and checkpoint settings. One command starts training.

White Circle 的头像
White Circle6 天前

Halo supports data, tensor, context, expert and expert-tensor parallelism. It also includes optimized attention, grouped-GEMM and fused loss kernels.

White Circle 的头像
White Circle6 天前

Halo is not limited to supervised fine-tuning. It supports async RL and training with external environments. We are excited to partner with the @sgl_project team to make it the primary engine for rollouts.

White Circle 的头像
White Circle6 天前

On @OpenAI gpt-oss-20b, Halo delivered 2.3–2.8x the throughput of stock TRL with less peak memory. Both runs used the same FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations.

White Circle 的头像
White Circle6 天前

We partnered with @liquidai to add native Halo support for LFM2.5-8B-A1B. On one B300, Halo was up to 20% faster than next best open-source framework.

White Circle 的头像
White Circle6 天前

We also used Halo to fine-tune @Zai_org GLM-4.7-Flash on 177M tokens of agentic traces. The resulting model improved SWE-rebench-V2 by @nebiusai from 33% to 42%. Halo reached up to 1.63× TRL throughput on the same model precision and data. Model and write-up:

White Circle 的头像
White Circle6 天前

Thanks for reading to the end! Star us on GitHub: Read more:

Daniel Smidstrup 的头像
Daniel Smidstrup6 天前

Open source is the way, and it's only going to get more and more relevant in the coming time! :)

White Circle 的头像
White Circle6 天前

True

leo 的头像
leo6 天前

insane way to kick off the week -- congrats @whitecircle team

White Circle 的头像
White Circle6 天前

Lets go!

neural nets. 的头像
neural nets.6 天前

crazy reading through the docs

Vishal Singh 🥑 的头像
Vishal Singh 🥑6 天前

love that this is open source, congrats team ❤️

Catalin 的头像
Catalin6 天前

Congrats on the awesome work, team! Good to see training infra being open source 👏 Starred the repo to contribute to the growth.

White Circle 的头像
White Circle6 天前

thanks!

Paul Mit 的头像
Paul Mit6 天前

great idea you've got my star, guys

giulia 的头像
giulia6 天前

took me one read to understand exactly who this is for. that almost never happens with infra launches. congrats team!!! starred 👀

Mac 的头像
Mac6 天前

Peak improvement, I will try it from first hand

sui 的头像
sui6 天前

so we get Grok 4.7 along with this today? nice.

White Circle 的头像
White Circle6 天前

Crazy day!

Maxime Labonne 的头像
Maxime Labonne6 天前

Congrats guys!

Virgile RIETSCH 的头像
Virgile RIETSCH6 天前

This is huge !!

Levan Kvirkvelia 的头像
Levan Kvirkvelia6 天前

haloshi framework

𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆 的头像
𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆6 天前

wait what this is incredible

White Circle 的头像
White Circle6 天前

ty Brian! show the repo some love ⭐

ivy 的头像
ivy6 天前

so excited!!!

White Circle 的头像
White Circle6 天前

🥰

Ted | unfair.so 的头像
Ted | unfair.so6 天前

Congrats team!

Soraia 的头像
Soraia6 天前

Great work! You got my star ⭐️

Nidhi Singh 的头像
Nidhi Singh6 天前

love that this is open source, congrats!

Rushil Chopra 的头像
Rushil Chopra6 天前

I have always wanted to train a custom math model, let’s see how this goes!

White Circle 的头像
White Circle6 天前

Keep us posted, any feedback is welcome

AshutoshShrivastava 的头像
AshutoshShrivastava6 天前

love seeing more training infrastructure get open-sourced. congrats on the release!

Ion 的头像
Ion6 天前

Congrats! Impressive what you delivered there 👀

Nikita Andersson 的头像
Nikita Andersson6 天前

@sarapenrique This is fabulous

Hauke 🌱🦀 的头像
Hauke 🌱🦀6 天前

Uh training my own models for my products would be pretty neat. I need better hardware 👀

White Circle 的头像
White Circle6 天前

💚

Prasenjit 的头像
Prasenjit6 天前

oh this is sick actually, well cooked

Kalash 的头像
Kalash6 天前

banger stuff 💥

Dan Kulkov 的头像
Dan Kulkov6 天前

insane start of the week congrats team!

Aditi 的头像
Aditi6 天前

wait… 2.8x faster than stock trl ??

Vasko 的头像
Vasko6 天前

congrats to the team!

Lucas Valbuena 的头像
Lucas Valbuena6 天前

congrats!

White Circle 的头像
White Circle6 天前

Thank you

Sameer 的头像
Sameer6 天前

congrats! def keeping an eye on this one.

Suraj Sharma 的头像
Suraj Sharma6 天前

The interesting part is keeping models in native HuggingFace format while getting that throughput boost.

White Circle 的头像
White Circle6 天前

Worked hard on this

Suraj Sharma 的头像
Suraj Sharma6 天前

LFG 🫡🙌

Jackson 的头像
Jackson6 天前

Really cool. Really unique and impressive numbers :•)

Marco Franzon 的头像
Marco Franzon6 天前

The only limit now is the hardware. Everyone can now train or finetune a model without rewriting it for the training phase. Awesome!

White Circle 的头像
White Circle6 天前

thanks Marco!

Marco Franzon 的头像
Marco Franzon6 天前

Great work !

Deni 的头像
Deni6 天前

Congrats on the launch team! Starred.

Ole Lehmann 的头像
Ole Lehmann6 天前

happy launch day guys"

相关视频