Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

opensource community while discovering ssms work better than transformers Soon you may not even need gpus for long context portable ai.

289,040 Aufrufe • vor 2 Jahren •via X (Twitter)

9 Kommentare

Profilbild von nisten - e/acc
nisten - e/accvor 2 Jahren

As far as I’ve heard from openai folk they’re “looking into it”. So this is new to them too. No one knows how well it’ll perform at scale, but a 0.79B and 2.9B model trained on just 300m tokens (because of being gpupoor) coming within ballpark of llama7B is pretty cray cray.

Profilbild von nisten - e/acc
nisten - e/accvor 2 Jahren

If anyone has spare compute to donate the skunkworks kitchen is ready to outcook openai with it just saying

Profilbild von The Highly Automated Cat — e/acc ⏩
The Highly Automated Cat — e/acc ⏩vor 2 Jahren

sexting sms?

Profilbild von nisten - e/acc
nisten - e/accvor 2 Jahren

structured state space models

Profilbild von 𝔠ɣ𝔟ε𝓇 𝕙ιε𝕣øթ𝕙𐊀ητ //daniel//
𝔠ɣ𝔟ε𝓇 𝕙ιε𝕣øթ𝕙𐊀ητ //daniel//vor 2 Jahren

Is the tech legit? Transformer-killers have flopped (get it?) in the past…

Profilbild von nisten - e/acc
nisten - e/accvor 2 Jahren

@ddaughertytech lk99 moment

Profilbild von murat 🍥
murat 🍥vor 2 Jahren

wow the deepfates program has really grown

Profilbild von Jamison 🏕️☃️
Jamison 🏕️☃️vor 2 Jahren

banger song

Profilbild von Tyr
Tyrvor 2 Jahren

Whatd I miss?

Ähnliche Videos

This is probably the most entertaining way to understand one of AI’s hardest AI debates. Transformer vs Post-Transformer, argued by leading researchers, inside a real physical boxing ring. Both technically deep and genuinely entertaining. I was glued for the entire 1 hour 20 minutes. So many super cool points to learn. 🥊 Transformers - Transformers still own the present because they work at scale. They are simple, trainable, hardware-friendly, and already power the strongest AI systems we use today. - The Transformer is basically a memory machine. It stores information as keys and values, then uses attention to pull back the most useful parts when answering. - The real Transformer advantage is not just “attention.” The bigger advantage is that it fits modern hardware extremely well, so it can process huge batches of tokens fast. - Scaling is still the brutal rule. If you give Transformers more compute, more data, and more parameters, they usually keep getting better. Any Post-Transformer architecture has to scale just as well, or better. - It is not enough to look clever on small tests, because the real question is whether it improves faster than Transformers when scaled up. - A replacement cannot be slightly better. Because the whole AI stack is already built around Transformers, the next architecture may need to be around 10x better to force everyone to switch. - Transformers are powerful, but they may be brute force. A human does not need to read the entire internet many times to become smart, but current LLMs need enormous data and compute. 🥊 Post-Transformer - Post-Transformer people are not saying Transformers are bad. They are saying Transformers may be the best current tool, not the final form of machine intelligence. - The biggest Post-Transformer target is native reasoning and continual learning. Today’s LLM reasoning often feels like text-based step-by-step work added on top, instead of thinking happening naturally inside the model. - Latent reasoning is one possible next step. That means the model reasons inside its own hidden internal space, instead of writing every thought out as words. - Continual learning is still a major weakness. Humans keep learning from experience, but most Transformer-based models are trained, frozen, and then only adapt inside the prompt. - Long context is not the same as real memory. A model can read a huge prompt, but that is different from building a life history, learning from mistakes, and updating beliefs over time. - The future may be hybrid, not a clean replacement. Transformers may stay as 1 building block while newer systems add better memory, better reasoning, and better learning loops. - The most interesting possibility is that Transformers may help discover their own successor. AI agents are already getting better at research and coding, so the next architecture may come from AI-assisted architecture search. ------- - Benchmarks are a problem. Many public benchmarks are easy to game, so they may show leaderboard strength without proving deeper intelligence. - Perplexity is still probably a great metric to evaluate frontier models,, because it tests prediction quality. --- Overall, Transformers continue to dominate, but the frontier is clearly widening. Pathway’s BDH (Dragon Hatchling — brain-inspired reasoning architecture), Sakana AI’s CTMs (Continuous Thought Machines — models that think over time), and Liquid AI’s LFMs (Liquid Foundation Models — efficient multimodal foundation models) - all of these show how the frontier is expanding. --- From “Pathway (pathway[.]com)” Youtube channel (link in comment) Zuzanna Stamirowska

Rohan Paul

89,110 Aufrufe • vor 1 Monat