Is a single accuracy number all we can get... from model evals?🤔 🚨Does NOT tell where the model fails 🚨Does NOT tell how to improve it Introducing EvalTree🌳 🔍identifying LM weaknesses in natural language 🚀weaknesses serve as actionable guidance (paper&demo 🔗in🧵) [1/n]show more

Zhiyuan Zeng
65,424 views • 1 year ago
🚨 Introducing Nomis Pass: All the Scores in a... Single Subscription. Easy to Get. Even Easier to Get Airdrops! You asked for less chaos. We built it. Start Here -> 1/ The price — $5. Potential rewards = $15K. Learn more 👇🧵show more

Nomis | Onchain Reputation Protocol
48,383 views • 1 year ago
🚀New paper out - We present Video-MSG (Multimodal Sketch... Guidance), a novel planning-based training-free guidance method for T2V models, improving control of spatial layout and object trajectories. 🔧 Key idea: • Generate a Video Sketch — a spatio-temporal plan with background, foreground, and motion in the pixel space. • Encode this structure directly into the latent space of the diffusion model during generation, which does not require fine-tuning or additional memory during inference. 🧵show more

Jialu Li
35,258 views • 1 year ago
🚨 BREAKING: This is how our Sun will die.... Not explode. Not collapse into a black hole. It will shed itself. As the core runs out of fuel, the structure holding the star together begins to fail. So it does something unexpected: It releases its outer layers… to stabilize. What we see as a “nebula” isn’t just debris. It’s a star rebalancing itself. In my view: Stars don’t just die. They resolve instability. And what’s left behind… is the most stable state it can reach. So the real question is: If stars follow this rule… Does everything?show more

TheNewPhysics
24,742 views • 5 months ago
I MADE MY AI AGENT 10X FASTER WITHOUT CHANGING... THE MODEL not a smarter model, not a bigger context window, not another clever prompt the same kind of AI that designs vaccines for viruses we have not even met yet was spending two minutes opening the wrong files just to hand me a brief from three months ago the problem was never capability, it was the scaffolding that piled up around my agent by accident, folder by folder an agent does not think in your categories, it searches from scratch every single time, and your tidy human folders are a maze to it the fix was almost stupidly small, one index file at the root of each big folder and a few numbers in front of the folder names slowest task dropped from 2 minutes to 26 seconds, fastest ones hit 10, zero model changes capability is cheap when the scaffolding around it is broken the article breaks down the whole system in 15 minutes ↓show more

shmidt
36,479 views • 3 months ago
NEW: Megyn Kelly delivers cryptic message, says there are... going to be a lot more Epstein developments in the coming year and we may hear about them from “him directly.” “We’re not done with Jeffrey Epstein. I can tell you that for a fact. Can't tell you how I know, but I can tell you for a fact.” “We're gonna hear a lot more about Jeffrey Epstein in the coming year... and you may be even hearing from him directly. More on that, as I'm allowed to tell it.”show more

Collin Rugg
4,444,513 views • 2 years ago
🚨 THE BITCOIN BOTTOM IS CLOSE. BUT NOT HERE... YET. Look at the cycle ROI chart. Every cycle takes time to fully bottom out. The give back phase does not end in a single candle. It grinds. We are inside that phase now, but the cycles show there is still road left before the real low prints. That is not bad news. It is time. Time to prepare. Time to position.show more

Crypto Rover
77,839 views • 3 months ago
This woman taking her baby daughter for a walk... in a stroller sees a new Chevy suv that she likes. She proceeds to leave the stroller on the sidewalk and trespass onto someone else’s property just to go and see what model it is. In a day and age of technology, all you really have to do is snap a photo and AI will tell you what it is and where to get one. Do you think she’s being nosy or does the homeowner have a case of trespassing?show more

SonnyBoy🇺🇸
386,495 views • 8 months ago
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
79,140 views • 1 year ago
🚨🚨 CNN JUST DID IT AGAIN 🚨🚨 A Black... woman is found HANGING FROM A TREE in Mississippi… And what does CNN lead with? “The imagery of a Black person in the South hanging from a tree evokes painful memories…” PAINFUL MEMORIES?! They left out the ONE detail that destroys the entire narrative: 👉 HER BLACK BOYFRIEND WAS ARRESTED IN CONNECTION WITH HER DEATH. Not a word. Not a chyron. Not a single mention. Just pure, calculated race-baiting while a woman is dead and the truth sits right there. This isn’t journalism. This is propaganda. This is how they keep the fire burning. How many more times are they going to do this before people finally wake up? Share this. Let them feel the heat.show more

CONSTITUTION X 🇺🇸
553,965 views • 1 month ago
🚨 GPT-5.6 features flag can be enabled in the... codex > So we can see ui but can't use the model directly > Only trusted users are given access for now Personal life update: my adhd get worse with vacation as it fuck up my routine that's why I was not active lately, suggest something to solve this 🫂show more

Chetaslua
354,989 views • 2 months ago
HydraFusion Explained. Part I: How does the Copilot engine... know what to optimize for? Your prompt is evaluated across 4 dimensions: ➡ Does it require deep reasoning? (aka. reasoning depth) ➡ Is it a sophisticated problem? (aka. code generation complexity) ➡ Is it untangling a complicated mess? (aka. debugging difficulty) ➡ Is it dominated by tool-use? (aka. tool orchestration needs) Based on this evaluation, a HyDRA score is assigned to determine the capability profile your task needs the most and to establish a quality bar. Part II: How does it choose a model? Note: It doesn' t pick one model to handle the entire job e2e, (that's Auto mode). Instead, it selects 1 of 3 execution workflows and assigns the best model at different stages based on the HyDRA score: 1️⃣ Single ⚙️ How it works: A single model completes the task from start to finish. ⚖️ Rationale: The task comfortably meets the quality bar with one model. Multi-model orchestration would add latency and cost with no meaningful quality gain. 2️⃣ Cascade ⚙️ How it works: A lightweight, cost-efficient model generates the solution. This draft is evaluated against a quality gate and if it falls short of the quality bar, the entire task escalates to a stronger, frontier model. ⚖️ Rationale: Only bring in the big guns when there is concrete evidence that a lightweight model won't meet the quality threshold. 3️⃣ Critique ⚙️ How it works: A lightweight model drafts the initial code and tool interactions. An independent, read-only frontier model reviews that draft and provides feedback. The original lightweight model then performs any targeted revision(s) before the final response is sent to the user. ⚖️ Rationale: Writing code (output tokens) is expensive while reviewing code (input tokens) is cheap. Instead of incurring the cost of a powerhouse writing hundreds of lines from scratch, a cost-efficient model writes the first draft, and the frontier model just reviews it and points out fixes. HydraFusion is available in experimental preview on the GitHub Copilot CLI: /experimental on, /model and select Hydrafusion (Research Preview)show more

Julia Muiruri
12,687 views • 21 days ago
🚨The Difference in Range Between Messi and Ronaldo is... Massive, Where No Distance or Angle seems to Matter for Ronaldo, but Messi Scores only from near the Box The Difference in the Number & Angle Of Green Spots on the Map Says it All. But Ronaldo is Portrayed as a Poacher ?🤯show more

Preeti
19,264 views • 10 days ago
Qwen3.6 35B A3B can't fill out a paper form... on its own. But give it NVIDIA's LocateAnything-3B — the #1 trending model on HuggingFace — as its eyes, and the two small models get it done together. (The test: place each element at the right pixel position on a blank form image, not type into a field.) Setup: > Qwen is the brain (main model), LocateAnything is the eyes (helper model acting as a tool). > I gave Qwen a new tool: ask "where's the email field?" and LocateAnything returns the exact x, y, width, height. > The blue boxes on the screen are its detections. Look how tight they are — it nails every field. Result: > Qwen3.6 35B A3B + LocateAnything-3B: form completed, all info correct. > Name, DOB, ID, gender, marital status, nationality, email, phone, address, postal code: all landed in the right field areas. > Character-box alignment still a touch loose, but every value is where it belongs. > 9m10s, 224.5k input, 24.3k output, 21 turns. Why it matters: > Qwen alone can't finish this test. Bolt on a 3B model that does exactly one thing > locate > and suddenly it can. > A combination of small models can do the work of a single large one.show more

stevibe
150,714 views • 3 months ago
You've probably scrolled past a dozen posts about Jev... this week without anyone telling you what it actually is. It's the first model from TypeSafe, a lab started by one of the researchers behind ChatGPT. It's a decision engine: you give it options, it picks one and tells you how sure it is. It cannot write a single word, and that is the interesting part. Every other AI you use writes. That is the whole interface. So when software needs a plain yes or no, we make a model write a paragraph and then dig the answer back out of it. Fine in a chat window where a human reads it. Bad inside software, where code has to act on it. The bet is that the valuable half was never the writing. It was the deciding. That problem showed up in WordPress years ago, and it is the reason WPVibe works the way it does. The AI does the work. Anything permanent stops and waits, because a delete that skips the trash is not something software should decide on its own.show more

John Turner
21,029 views • 12 days ago
🚨 JUST IN: Netanyahu admits Israel DOES NOT need... help reaching their “goals” in Iran This is EXACTLY WHY we anti-war advocates have been so vocal this past week. This is NOT OUR WAR — it’s Israel’s. And if they claim they can do it themselves, be my guest. “Israel has the capability to achieve all its goals alone when it comes to Iran’s nuclear facilities and that it is up to President Trump if he wants to join in or not.” – Bibi Netanyahushow more

Nick Sortor
788,275 views • 1 year ago
Robotics keeps hitting the same wall. Single task RL... works, but... it does not scale to hundreds of tasks or new embodiments. This new paper looks like a real step toward fixing that. The team introduces MMBench, a benchmark with 200 tasks across many domains and robots, and Newt, a language conditioned world model trained online across all 200 tasks at once. The simple idea behind Newt: The model learns from demos to get the right priors It trains across many tasks through online interaction It uses language to ground the goal It adapts fast when a new task shows up What stood out to me: ✅ One model trained on 200 tasks at the same time ✅ Language conditioned control for both states and RGB ✅ Better data efficiency than strong baselines ✅ Strong open loop control ✅ Fast adaptation to new tasks and embodiments ✅ Full release of 200 checkpoints, 4000 demos, code, and benchmark This is a good push toward general control instead of one model per task. If you want the full paper: Project page: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
70,090 views • 10 months ago
What the hell does Quant actually do? A model... is trained in BF16, released in Q8, and we use Q4 because we're told 'its nearly lossless'. Fine, but I was curious about the 'nearly' part, what is actually being lost between BF16 all the way down to Q2? As always, its complicated. Needed a way to visualize the loss, so I had a Qwen 3.8 27B BF16 create a paragraph then had every quant below it do the same and marked the mutations. When a model writes a word, it assigns 'how sure' percentage to each. For example, 'The weather tomorrow is *windy*, a model is 90% sure all the words are correct except the windy part, since it can also be sunny, cold, hot, cloudy, etc. Any words a BF16 model was 90% sure of, every quant down to Q2 rarely changes it. On the other hand, any word that it was only 40~50% sure of, such as windy in our example, are the types of data that gets impacted by the quant process. Okay so what, some words changed, how does this impact anything? I set out to find out, and its all below if you are just as weird as I am about these things.show more

Killy
19,592 views • 21 days ago