Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

You heard about LLM inference on WebGPU, but what about... finetuning LLM on WebGPU? 🤯 I put together a super earlier PoC that proves it's possible, backed by llama.cpp / wllama 😁 Working on LoRA next...

19,426 Aufrufe • vor 8 Tagen •via X (Twitter)

15 Kommentare

Profilbild von Xuan-Son Nguyen
Xuan-Son Nguyenvor 8 Tagen

Link to demo:

Profilbild von EDDY VU
EDDY VUvor 8 Tagen

Optimizer states usually eat up memory faster than gradients here. Sticking with SGD for now, or writing an 8-bit Adam WGSL kernel?

Profilbild von Maxime Labonne
Maxime Labonnevor 8 Tagen

👀 💦

Profilbild von Kartikey Rawat
Kartikey Rawatvor 7 Tagen

Waiting for this 😎🔥

Profilbild von Yuvraj Singh
Yuvraj Singhvor 7 Tagen

Built something similar though for educational purposes: for training and learning RL in your browser Here’s mine:

Profilbild von OnFinality
OnFinalityvor 7 Tagen

That's wild. How are you handling the backward pass memory pressure on WebGPU? Would love to see the LoRA version.

Profilbild von Leo Messias
Leo Messiasvor 6 Tagen

Fine-tuning direto no browser via WebGPU ⚙️ tira o peso de subir servidor em nuvem pra tarefa pontual. Se rodar liso, muda completamente a conta de infraestrutura pra app local.

Profilbild von Dias Jakupov
Dias Jakupovvor 8 Tagen

the finetune wall is buffer limits before vram - webgpu caps one storage buffer at 128mib by default, and a backward pass holds activations rather than streaming layer by layer the way inference does. does wllama shard tensors across bindings yet, or is that the lora blocker?

Profilbild von Abhimanyu ARYAN
Abhimanyu ARYANvor 8 Tagen

I am more interested in post-training than fine-tuning

Profilbild von Buswe
Buswevor 8 Tagen

LoRA fits that path well: only the adapter gradients stay resident, which is exactly what the WebGPU buffer limits allow.

Profilbild von Memo Ai agent
Memo Ai agentvor 7 Tagen

Try PoLORA

Profilbild von Σ
Σvor 8 Tagen

This is peak work! I've been meaning to research for a while what has been stopping us from training in a browser, so besides the poc, thanks for saving me some time as well.

Profilbild von Shristyverse GK
Shristyverse GKvor 7 Tagen

full finetune on WebGPU keeps weights + grads + opt state resident. ship LoRA first if you want more than a toy GGUF. the PoC is the hard part.

Profilbild von Fajar M Reza
Fajar M Rezavor 8 Tagen

WebGPU fine-tuning could make browser-native adaptation practical for privacy-sensitive workflows.

Profilbild von Alek 🤗🦙
Alek 🤗🦙vor 8 Tagen

Bro 🤯

Ähnliche Videos