
0xSero
@0xSero • 57,901 subscribers
I want to be free
Shorts
Videos

Here's my conversation with Philip Kiely. A deep dive into inference engineering, his book his available online for free, and it's worth reading. Inference is becoming more important to people and organisations by the day, learning about the topic will prepare you well.
0xSero24,290 views • 17 days ago

Local AI can be so good, but you’d need about 12k USD to get it. Then it’s not so great Here’s a Q4 of Qwen3.5-262B-REAP Weights 131GB KVCache 50GB 256,000 context 350 tokens/s prefill warmup 4,000 tokens/s prefill cache 36 tokens/s generation Vision enabled REAP is good
0xSero87,959 views • 3 months ago

MiniMax-M3 running locally 100+ tok/s Good questions not often that models ask good questions.
0xSero32,537 views • 1 month ago

Finally, full precision MiniMax-M2.7 running at home. 100 tokens/s decode 5050 tokens/s prefill
0xSero72,465 views • 3 months ago

Guess who got GLM-5.2 at home? You can't deny I'm elite at inference.
0xSero27,083 views • 1 month ago

Let me save you hours of testing frontends. If you're ever working on a front-end, instead of writing tests, and adding puppeteer slop to your repo 1. Get an llm to write you with whatever needs to be tested 2. Copy that, go to browser 3. Open localhost with your selected app 4. Use Claude Chrome Extension or Parchi 5. Send it the prompt 6. QA engineering, there you go. Use models results and pass it back to your coding agent to fix whatever is flagged.
0xSero109,943 views • 5 months ago

I yapped about LLM Compression for 40 minutes, how much misinformation did i spread this time (,:
0xSero35,181 views • 1 month ago

I LOVE Deepseek-v4-flash, incredibly reliable and capable, logical. It's lacking in frontend but I have MiMo for that. I would recommend any company spending 100k+ a year on AI to purchase 8-10~ 6000s and have a few of the works to have them blind test these models for work.
0xSero51,035 views • 2 months ago