Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

You heard about LLM inference on WebGPU, but what about... finetuning LLM on WebGPU? 🤯 I put together a super earlier PoC that proves it's possible, backed by llama.cpp / wllama 😁 Working on LoRA next...

19,828 Aufrufe • vor 12 Tagen •via X (Twitter)

15 Kommentare

Profilbild von Xuan-Son Nguyen
Xuan-Son Nguyenvor 12 Tagen

Link to demo:

Profilbild von EDDY VU
EDDY VUvor 12 Tagen

Optimizer states usually eat up memory faster than gradients here. Sticking with SGD for now, or writing an 8-bit Adam WGSL kernel?

Profilbild von Maxime Labonne
Maxime Labonnevor 12 Tagen

👀 💦

Profilbild von Kartikey Rawat
Kartikey Rawatvor 12 Tagen

Waiting for this 😎🔥

Profilbild von Yuvraj Singh
Yuvraj Singhvor 11 Tagen

Built something similar though for educational purposes: for training and learning RL in your browser Here’s mine:

Profilbild von OnFinality
OnFinalityvor 12 Tagen

That's wild. How are you handling the backward pass memory pressure on WebGPU? Would love to see the LoRA version.

Profilbild von Leo Messias
Leo Messiasvor 10 Tagen

Fine-tuning direto no browser via WebGPU ⚙️ tira o peso de subir servidor em nuvem pra tarefa pontual. Se rodar liso, muda completamente a conta de infraestrutura pra app local.

Profilbild von Dias Jakupov
Dias Jakupovvor 12 Tagen

the finetune wall is buffer limits before vram - webgpu caps one storage buffer at 128mib by default, and a backward pass holds activations rather than streaming layer by layer the way inference does. does wllama shard tensors across bindings yet, or is that the lora blocker?

Profilbild von Abhimanyu ARYAN
Abhimanyu ARYANvor 12 Tagen

I am more interested in post-training than fine-tuning

Profilbild von Buswe
Buswevor 12 Tagen

LoRA fits that path well: only the adapter gradients stay resident, which is exactly what the WebGPU buffer limits allow.

Profilbild von Memo Ai agent
Memo Ai agentvor 12 Tagen

Try PoLORA

Profilbild von Σ
Σvor 12 Tagen

This is peak work! I've been meaning to research for a while what has been stopping us from training in a browser, so besides the poc, thanks for saving me some time as well.

Profilbild von Shristyverse GK
Shristyverse GKvor 11 Tagen

full finetune on WebGPU keeps weights + grads + opt state resident. ship LoRA first if you want more than a toy GGUF. the PoC is the hard part.

Profilbild von Fajar M Reza
Fajar M Rezavor 12 Tagen

WebGPU fine-tuning could make browser-native adaptation practical for privacy-sensitive workflows.

Profilbild von Alek 🤗🦙
Alek 🤗🦙vor 12 Tagen

Bro 🤯

Ähnliche Videos