Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub:

2,252,490 просмотров • 6 дней назад •via X (Twitter)

Комментарии: 56

Фото профиля White Circle
White Circle6 дней назад

Every model we train at White Circle now uses Halo. The same codebase runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL. Technical blog:

Фото профиля White Circle
White Circle6 дней назад

Other training frameworks often require a separate implementation for each model family. In Halo, adding a new model family takes just 100 lines of connecting a wrapper instead of rewriting the model.

Фото профиля White Circle
White Circle6 дней назад

One YAML file contains the model, training method, GPU setup and checkpoint settings. One command starts training.

Фото профиля White Circle
White Circle6 дней назад

Halo supports data, tensor, context, expert and expert-tensor parallelism. It also includes optimized attention, grouped-GEMM and fused loss kernels.

Фото профиля White Circle
White Circle6 дней назад

Halo is not limited to supervised fine-tuning. It supports async RL and training with external environments. We are excited to partner with the @sgl_project team to make it the primary engine for rollouts.

Фото профиля White Circle
White Circle6 дней назад

On @OpenAI gpt-oss-20b, Halo delivered 2.3–2.8x the throughput of stock TRL with less peak memory. Both runs used the same FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations.

Фото профиля White Circle
White Circle6 дней назад

We partnered with @liquidai to add native Halo support for LFM2.5-8B-A1B. On one B300, Halo was up to 20% faster than next best open-source framework.

Фото профиля White Circle
White Circle6 дней назад

We also used Halo to fine-tune @Zai_org GLM-4.7-Flash on 177M tokens of agentic traces. The resulting model improved SWE-rebench-V2 by @nebiusai from 33% to 42%. Halo reached up to 1.63× TRL throughput on the same model precision and data. Model and write-up:

Фото профиля White Circle
White Circle6 дней назад

Thanks for reading to the end! Star us on GitHub: Read more:

Фото профиля Daniel Smidstrup
Daniel Smidstrup6 дней назад

Open source is the way, and it's only going to get more and more relevant in the coming time! :)

Фото профиля White Circle
White Circle6 дней назад

True

Фото профиля leo
leo6 дней назад

insane way to kick off the week -- congrats @whitecircle team

Фото профиля White Circle
White Circle6 дней назад

Lets go!

Фото профиля neural nets.
neural nets.6 дней назад

crazy reading through the docs

Фото профиля Vishal Singh 🥑
Vishal Singh 🥑6 дней назад

love that this is open source, congrats team ❤️

Фото профиля Catalin
Catalin6 дней назад

Congrats on the awesome work, team! Good to see training infra being open source 👏 Starred the repo to contribute to the growth.

Фото профиля White Circle
White Circle6 дней назад

thanks!

Фото профиля Paul Mit
Paul Mit6 дней назад

great idea you've got my star, guys

Фото профиля giulia
giulia6 дней назад

took me one read to understand exactly who this is for. that almost never happens with infra launches. congrats team!!! starred 👀

Фото профиля Mac
Mac6 дней назад

Peak improvement, I will try it from first hand

Фото профиля sui
sui6 дней назад

so we get Grok 4.7 along with this today? nice.

Фото профиля White Circle
White Circle6 дней назад

Crazy day!

Фото профиля Maxime Labonne
Maxime Labonne6 дней назад

Congrats guys!

Фото профиля Virgile RIETSCH
Virgile RIETSCH6 дней назад

This is huge !!

Фото профиля Levan Kvirkvelia
Levan Kvirkvelia6 дней назад

haloshi framework

Фото профиля 𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆
𝗕𝗿𝗶𝗮𝗻 𝗥𝗲𝘆6 дней назад

wait what this is incredible

Фото профиля White Circle
White Circle6 дней назад

ty Brian! show the repo some love ⭐

Фото профиля ivy
ivy6 дней назад

so excited!!!

Фото профиля White Circle
White Circle6 дней назад

🥰

Фото профиля Ted | unfair.so
Ted | unfair.so6 дней назад

Congrats team!

Фото профиля Soraia
Soraia6 дней назад

Great work! You got my star ⭐️

Фото профиля Nidhi Singh
Nidhi Singh6 дней назад

love that this is open source, congrats!

Фото профиля Rushil Chopra
Rushil Chopra6 дней назад

I have always wanted to train a custom math model, let’s see how this goes!

Фото профиля White Circle
White Circle6 дней назад

Keep us posted, any feedback is welcome

Фото профиля AshutoshShrivastava
AshutoshShrivastava6 дней назад

love seeing more training infrastructure get open-sourced. congrats on the release!

Фото профиля Ion
Ion6 дней назад

Congrats! Impressive what you delivered there 👀

Фото профиля Nikita Andersson
Nikita Andersson6 дней назад

@sarapenrique This is fabulous

Фото профиля Hauke 🌱🦀
Hauke 🌱🦀6 дней назад

Uh training my own models for my products would be pretty neat. I need better hardware 👀

Фото профиля White Circle
White Circle6 дней назад

💚

Фото профиля Prasenjit
Prasenjit6 дней назад

oh this is sick actually, well cooked

Фото профиля Kalash
Kalash6 дней назад

banger stuff 💥

Фото профиля Dan Kulkov
Dan Kulkov6 дней назад

insane start of the week congrats team!

Фото профиля Aditi
Aditi6 дней назад

wait… 2.8x faster than stock trl ??

Фото профиля Vasko
Vasko6 дней назад

congrats to the team!

Фото профиля Lucas Valbuena
Lucas Valbuena6 дней назад

congrats!

Фото профиля White Circle
White Circle6 дней назад

Thank you

Фото профиля Sameer
Sameer6 дней назад

congrats! def keeping an eye on this one.

Фото профиля Suraj Sharma
Suraj Sharma6 дней назад

The interesting part is keeping models in native HuggingFace format while getting that throughput boost.

Фото профиля White Circle
White Circle6 дней назад

Worked hard on this

Фото профиля Suraj Sharma
Suraj Sharma6 дней назад

LFG 🫡🙌

Фото профиля Jackson
Jackson6 дней назад

Really cool. Really unique and impressive numbers :•)

Фото профиля Marco Franzon
Marco Franzon6 дней назад

The only limit now is the hardware. Everyone can now train or finetune a model without rewriting it for the training phase. Awesome!

Фото профиля White Circle
White Circle6 дней назад

thanks Marco!

Фото профиля Marco Franzon
Marco Franzon6 дней назад

Great work !

Фото профиля Deni
Deni6 дней назад

Congrats on the launch team! Starred.

Фото профиля Ole Lehmann
Ole Lehmann6 дней назад

happy launch day guys"

Похожие видео