Loading video...
Video Failed to Load
🧱 Transformers has hit the scaling wall 🧱 💰 GPT 4.5 cost billions, with no clear path to AGI for 10x$ more 📘 Facebook, Yann LeCun, is now saying we need new architectures 🔎 Deepmind CEO, Demis Hassabis, is saying we need 10 years We have another path to... show more
90,772 views • 1 year ago •via X (Twitter)
12 Comments

At the heart of it, todays top models are - Capable: Of incredible PhD level tasks & beyond - (Un)Reliable: Maybe 1-out-of-30 time What everyone want, is not a smarter model But a more reliable model, doing basic college level task Longer write up:

To do the more boring things in life like - organize emails and receipts - fill up forms - order groceries - be a friend The things that actually matter .... All tasks which a 72B model is more than capable of, if only it was more reliable

And thats where our work in Qwerky comes in. Because the one thing that is holding these AI models, and agents back.... Is simply the lack of reliable understanding, in memories. Memories which is at the heart of recurrent model like RWKV...

Instead of scaling bigger more expensive models, which is unable to bring an ROI to investors What if we iterate faster, at <100B active parameters. To make these already capable models. More reliable instead. More personalizable. At a size with ROI

A model which can be "Memory tuned" without Catastrophic Forgetting Overcoming the barrier, which makes finetuning out of reach for the vast majority of teams With quick efficient personalization of AI models, to unlock reliable commercial AI agents. Without compounding errors

Memories is the secret to AGI Once memories for personalized AI is mastered. Where it can be reliably tuned, with controlled datasets by AI Engineers easily... The next step is to get the AI model to prepare their continuous training dataset without compounding loss

Its a binary question, is recurrent memory the path to AGI? If so, this path to AGI is inevitable As all the critical ingredients is already here, and is not bound by hardware, only by software You can read more in details in our long form writing...

Wall Street ain't ready for this... Coinbase launched Base ~1 year ago This Layer 2 blockchain has raked in ~$1.2M every week on average Now Wall Street is FOMOing into crypto. Front run them by reading Milk Road. 5 minutes. Every day. For free.

Interesting read, reliability of daily fine-tuning with new memories is something I'm looking forward too. I hope you succeed !

Thank you 🙏 i believe all of us wish for more AI to do the boring chores in life... reliably to our prefences

I've been following RWKV for years, and I believe its approach is more elegant than the brute-force method transformers use to make LLMs work. The only potential weakness might be the importance of prompt sequencing, but this can be mitigated by users. Great Project!

As per the latest RWKV v7 paper - a properly trained 1.5B and 3B model is able to handle ~32k context length, without any major issues Prompt sequencing would be an issue after that If scaled correctly, for larger models, this is mostly mitigated within std context length

