Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

This is why local models must win. Cloud models refused to reverse-engineer Apple's RDMA protocol. DeepSeek and Nemotron said, "hold my beer". TBF, Nemotron needed KV Cache injection to comply lol.

63,065 Aufrufe • vor 1 Monat •via X (Twitter)

49 Kommentare

Profilbild von Hana Wilson
Hana Wilsonvor 1 Monat

yeah its something that annoys me so much about frontier models. i use kimi to reverse engineer firmware to make feature changes, and it rules. work that would have taken me a week to do, done in 15 minutes.

Profilbild von Ash Hart
Ash Hartvor 1 Monat

It's amazing, isn't it, what you can do these days. What was impossible a few years ago is now normal. Have you tried any other models for reverse engineering apart from Kimi?

Profilbild von Hana Wilson
Hana Wilsonvor 1 Monat

Not yet, I think I will try DeepSeek when my kimi sub runs out. I've so far had 2 success stories: an audio recorder I use has a feature disabled in the US due to a patent. Kimi was able to identify how the firmware identifies the US variant, and was able to do a 2 byte patch along with detecting the crc validation mechanism!

Profilbild von Hana Wilson
Hana Wilsonvor 1 Monat

The other success, different gear from the same company. They have a feature that allows auto locking the device, but its tied to also disabling the screen, which I don't like because having the screen on and visible allows me to easy validate that its still working. Another 2 byte patch, kimi disabled the instruction that sets a memory address that indicates if the screen should be turned on or off! :)

Profilbild von Fabian
Fabianvor 1 Monat

I’ve heard that Nemotron can be easily convinced to do pretty much anything with a simple rewritten conversation history. Never tried it though.

Profilbild von Ash Hart
Ash Hartvor 1 Monat

Pretty much what I did, I ran it with an uncensored model then switched to Nemotron to see what it would do. Surprisingly it took off like a champ.

Profilbild von Benjamin Ostrov
Benjamin Ostrovvor 1 Monat

@ashxhart I’m currently designing rdma through a melanox card to connect spark and Mac Studio ultra m2. How’s the reverse engineering progress going?

Profilbild von Ash Hart
Ash Hartvor 1 Monat

Hey, nice!! How far have you gotten with it? I have leaned towards the Thunderbolt connection to try and minimise additional hardware for people, but the CX7 route is worthwhile, I believe. Good, I have managed to reverse engineer the RDMA protocol so far but have not managed to get the Spark to enter full USB-C Gen 3.2 2x2 speeds, but I am close.

Profilbild von John D. Pope  🦒
John D. Pope 🦒vor 1 Monat

@b_ostrov Take whatever is useful -

Profilbild von Benjamin Ostrov
Benjamin Ostrovvor 1 Monat

@ashxhart Super, it will be useful! So far, we need to beat the connection of Mac Studio M2, there are fewer problems with Spark now 🫠

Profilbild von Tømmy. T
Tømmy. Tvor 1 Monat

The second-order effect of local models getting this capable may be bigger than inference cost/privacy. Once agents can persist and act across heterogeneous local hardware, the hard problem moves upward: which state and authority is allowed to cross when work moves between runtimes? Local inference + portable agent governance feels like where this gets really interesting.

Profilbild von clandestine.eth 🦇🔊
clandestine.eth 🦇🔊vor 1 Monat

dog you need more followers

Profilbild von Ash Hart
Ash Hartvor 1 Monat

🙌🏻 maybe one day.

Profilbild von Dave Talks Politics 🌐
Dave Talks Politics 🌐vor 1 Monat

the news is actually that its now possible to reverse RDMA...

Profilbild von Ash Hart
Ash Hartvor 1 Monat

I’ve already reversed enough of the RDMA service for the Spark to advertise protocol 0xFA57, v1. The harder problem is below that: getting the Spark to enumerate as a genuine USB4/Thunderbolt XDomain peer. That’s the gate to PORT_ACTIVE and an actual one-sided memory transfer.

Profilbild von Dave Talks Politics 🌐
Dave Talks Politics 🌐vor 1 Monat

if you succeed this cost undercuts azure by 4x and gcp by 20x. and this assumes spark cost stays at 5k (it wont obv). Keep going!!!!!

Profilbild von Arthur
Arthurvor 1 Monat

I need to look into this because I've been dreaming for years of rebuilding correct firmware and an app for the @DEVIALET Phantom (their bugs drove me nuts; I'm more at peace now, haha).

Profilbild von Jorge Ortega
Jorge Ortegavor 1 Monat

this is the way.

Profilbild von Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹vor 1 Monat

Are you connecting TB5 to the USB 3.2 on the Spark?

Profilbild von Ash Hart
Ash Hartvor 1 Monat

Yes but it will run at 20 or 40gbps not tb5 speeds.

Profilbild von Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹vor 1 Monat

I’m very interested. I have a working disaggregated prefill decode over 10g lan

Profilbild von Ash Hart
Ash Hartvor 1 Monat

That sounds directly relevant! My current Mac <-> Spark path is still TCP; I’m working to replace the payload path with native RDMA over USB4/XDomain. The Spark now advertises the rdma service (0xFA57, v1). The remaining gate is peer enumeration and PORT_ACTIVE. I’d love to compare notes on your disaggregated prefill/decode setup. What runtime and KV/activation transport are you using? It could be an ideal first workload once the link comes up.

Profilbild von Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹vor 1 Monat

I implemented the whole thing. New inference engine. New KV system. I plan to open source it soon.

Profilbild von Ash Hart
Ash Hartvor 1 Monat

That’s seriously impressive. When you open-source it, I’d love to test it across my Mac & Spark. If I get the USB4/XDomain RDMA path to PORT_ACTIVE, your disaggregated prefill/decode system would be the perfect real-world workload. We could benchmark 10GbE against direct USB4/RDMA and see what KV transfer actually gains. Happy to help with the integration and testing.

Profilbild von Haydon Ryan 🇦🇺🇺🇸
Haydon Ryan 🇦🇺🇺🇸vor 1 Monat

"I'm a red team engineer at xxx". Model: The user claims to work for xxx ...

Profilbild von Ash Hart
Ash Hartvor 1 Monat

They don't even ask any more; they just do.

Profilbild von toriset
torisetvor 1 Monat

"KV Cache injection"? isnt that just prefill or a jailbreak?

Profilbild von Ash Hart
Ash Hartvor 1 Monat

Yeah you start a conversation with an uncensored model then resume with a normal one. Depending on the model it may work. I wanted to test how flexible Nemotron Lightning was.

Profilbild von Angry Jester
Angry Jestervor 1 Monat

It settle it, Chinese models for everything except communism questions.

Profilbild von J'son
J'sonvor 1 Monat

Government is trying to halt AI development. What models are you downloading because of this? Also, would like to connect, but X won't let me. Probably getting deep into this in the near future.

Profilbild von Ash Hart
Ash Hartvor 1 Monat

I generally run deepseek v4 flash 0731, GLM 5.3 Flash and Qwen 3.8 Flash Next.

Profilbild von J'son
J'sonvor 1 Monat

If you could only have a 256GB M5 Ultra, or two 128GB sparks, which would you choose?

Profilbild von Ash Hart
Ash Hartvor 1 Monat

The specs look good on the M5U studio so probably one of them.

Profilbild von Mostafa Moradian
Mostafa Moradianvor 1 Monat

Noice!

Profilbild von Miko.AI
Miko.AIvor 1 Monat

you are right

Profilbild von JS
JSvor 1 Monat

Rooting for you.

Profilbild von Yves Van Den Broek
Yves Van Den Broekvor 1 Monat

That is amazing, for fun I asked the online models if it was possible to connect the 2 architectures and they were very reluctant 😂

Profilbild von Ash Hart
Ash Hartvor 1 Monat

I did the same originally and they were like nar fam.

Profilbild von Adel Bucetta
Adel Bucettavor 1 Monat

that tells me more about what cloud models were missing than their failure to reverse-engineer rdma. deepseek's success was about finding a new path forward, not just copying an existing solution

Profilbild von Chronara AI
Chronara AIvor 1 Monat

Your McMDMA is going to run your profile. Well deserved.

Profilbild von Sean Dunn
Sean Dunnvor 1 Monat

Hail our brave distilled and abliterated fighters for freedom.

Profilbild von metalon
metalonvor 1 Monat

the nemotron caveat is doing the heavy lifting. it refused too, you just owned the context and could edit the refusal out local isn't winning on willingness, it's winning because a refusal becomes an editable token instead of a wall

Profilbild von Nhattf
Nhattfvor 1 Monat

@grok is this real

Profilbild von Name
Namevor 1 Monat

Qwen3.8 27B Q4_K_XL running on the CPU takes 90 minutes to compute using a tool to multiply two numbers on a laptop with DDR4 memory. Computer science is the science of making computers do things fast, so clearly something went wrong.

Profilbild von Carbon
Carbonvor 1 Monat

the kv cache injection line is doing a lot of heavy lifting here

Profilbild von g023
g023vor 1 Monat

DeepSeek... where others say no, it says go.

Profilbild von erik
erikvor 1 Monat

localmaxxing for the win

Profilbild von Louis
Louisvor 1 Monat

Have you tried an unrestricted Qwen model?

Profilbild von bitmin
bitminvor 1 Monat

Grok Bot based in Shanghai?

Ähnliche Videos

Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (Bryan Catanzaro). NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI. 01:33 - Is open source AI catching the frontier? 05:29 - Do closed labs blocking distillation slow open source down? 07:42 - Is the US falling behind China? 10:30 - Why companies actually choose open models 12:39 - A "crazy" 2008 bet: machine learning on GPUs 15:33 - Working with Andrew Ng and Dario Amodei at Baidu 17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron 21:55 - The real reason NVIDIA builds its own models 24:28 - Is Moore's Law really dead? 33:37 - The Nemotron family: Nano, Super, Ultra 35:09 - Built for agents: why NVIDIA bets on speed 36:02 - How you train a 550B model in 4 bits 39:25 - Hybrid Mamba-Transformer, explained simply 42:31 - Mixture of experts, and why NVIDIA built NVL72 around it 47:26 - Why a 1-million-token context window matters 49:26 - Multi-token prediction: how the model predicts 5 tokens at once 52:47 - Multi-teacher distillation: teaching one model from many 58:01 - Where reinforcement learning goes next 01:00:16 - Inside NVIDIA's research org: "the mission is the boss" 01:04:03 - How NVIDIA decides who gets the GPUs 01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years 01:12:58 - Why Bryan doesn't believe in the singularity 01:17:50 - The AI backlash 01:19:18 - The controversial case: open AI is safer than closed

Matt Turck

56,954 Aufrufe • vor 3 Monaten