Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

computer use agents are slow because they think too much they make round trips to an LLM. every single time. even for a command you've said a hundred times before. that's why it always feels slow. my mac now remembers. the first time i say something new, like "search...

14,212 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

your agent reviewing its own work is not a check. it is a second opinion from the same source. this is the most common gap in agent systems and it hides in plain sight, because the step exists. there is a review. it just cannot do the thing you think it does. here is the mechanism. the model produced an output from a context. you then ask the same model, holding the same context, whether that output is correct. it answers fluently, because that is what it does. and the answer is drawn from the same distribution that produced the thing being judged. same weights, same window, same blind spots. if the reason the output is wrong is something the model does not know, the review does not know it either. if the reason is something the context does not contain, the review has the same context. the failure mode and the detector share a cause. > why it feels like it works because most of the time the output is fine, and the review says fine. agreement is not evidence of detection. a reviewer that says pass on everything agrees with reality most of the time too. what you actually want to measure is what happens on the cases that are wrong. that is the only place a check earns its name, and it is exactly the place where a self-review is weakest. there is research on this. Huang and colleagues at DeepMind showed at ICLR 2024 that intrinsic self-correction, revising without external grounding, does not reliably help and often makes things worse. > what to actually do move the check outside the model. a test that runs, a schema that validates, a file that exists or does not, an exit code from something you did not write. these are not smarter than the model. they are just not correlated with it, and that is the entire value. when the judgement genuinely needs a model, at minimum use a different family. same family means shared blind spots, and frontier judges measurably inflate scores for outputs that look like their own. and split the work by kind. anything objectively checkable goes to code. only the genuinely semantic calls go to a judge, and those get a rubric written as one line. a review inside the loop tells you the model is confident. a check outside it tells you whether the work is done. save this - then read the eval setup below

Hanako

14,325 Aufrufe • vor 24 Tagen

Japan just changed what an AI model even is. New Sakana Fugu doesn't try to out-think GPT-5, Claude, or Gemini. It conducts all three at once - and beats every one of them. A trader in Tokyo unleashed it on the fastest market alive - 5min Bitcoin binary and turned $6,200 into $304,865. His wallet: The frontier just stopped being which model is smartest. It's who's conducting them - and the market hasn't priced that in yet. Sakana Fugu isn't a bigger model - it's a full multi-agent orchestration system. The coordinator behind it carries about 10,000 parameters, evolved rather than hand-coded, and it runs the most capable models on earth like a single instrument. Pointed at Bitcoin, here's what it does every five minutes. It assembles a team from a pool of frontier models and assigns each one a role: > Thinker - reads the candle, the order book, the news, builds the plan > Worker - turns the plan into one call: up or down, and how much > Verifier - votes ACCEPT or REVISE before a cent moves If the Verifier says REVISE, nothing trades. Fugu reads its own miss, reroutes, even calls itself for a corrective round, and runs it again. No look-ahead, ever - the next candle only appears after it commits. This is what should worry every lab still chasing a bigger model: the edge was never scale. It's orchestration - and Fugu does it better than anything alive. Bookmark this - when the whole timeline is chasing orchestration in six months, you'll already have the breakdown. You're not going to wire up an orchestra of frontier models yourself. Mirror the wallet Fugu runs instead:

cvxv666

72,821 Aufrufe • vor 2 Monaten

Video of my first drive impressions of the Tesla Model Y L Premium Long Wheelbase. It has adaptive damping that noticeably improves ride comfort on bad roads versus the shorter Model Y Premium (LA is a great place to test this lol). While the new Model Y Performance also has adaptive damping, that is tuned more for sportiness. The Y L dampening is tuned for comfort, and even on the 20" Uberhelix wheels, it soaks up the bumps really well. The Model Y L also benefits from updates to Tesla's adaptive system controls, which improve the accuracy and latency of their actuator control. The Model Y L doesn't feel all that noticeably larger than the Model Y while driving it, but what is immediately noticeable is the improved rear view out the back window from the review mirror. Due to the shape of the glass and the rear, you get a taller, more expansive view. It's very quiet in every row, and the Y L even features improvements in this area versus the current Model Y. There is a new "Rear Comfort" setting that is exclusive to the Model Y L. Basically what that does is it prioritizes ride comfort for second and third rows over optimal steering and handling and first row comfort. My wife said it was her favorite second-row experience of any Tesla she's ever been in (which is all of them). The Model Y L just feels really refined overall. The driving dynamics, sound dampening, suspension, FSD, etc, it all just feels really dialed in, and I think when you test drive it, you will agree. As for FSD, I got to try FSD V14.3.7, and it's as you'd expect: amazing. I'm going to post a couple separate dedicated FSD videos later today.

Sawyer Merritt

122,919 Aufrufe • vor 26 Tagen

I USE THIS DEVICE FROM 26 YEARS AGO DAILY. In fact I wrote this on it. There is a quiet magic in the AlphaSmart 3000 that no modern laptop or glowing rectangle can touch. I carry one into the forest the way some people carry a trusted notebook. Three AA batteries keep it alive for months. Literally months. The screen shows only four lines of text at a time. There is no Wi-Fi. No notifications. No browser tab whispering that I should check something else. Just a full-size keyboard, a soft click under the fingers, and the words as they arrive. The story behind it is pure garage-to-classroom ingenuity. In the early 1990s two Apple engineers, Joe Barrus and Ketan Kothari, kept hearing the same complaint from teachers: kids were spending more time fighting the computer than writing. So they left Apple, brought in Ketan’s brother Manish, and built a device that did one thing with absolute focus. The first one was released in 1993 and the style is very Macintosh Platform language. I would not recommend that unit as it is very limited. The AlphaSmart 3000 arrived in 2000 wearing that iconic translucent bondi-blue shell that matched the first iMac and little too much like the eMate. It was never meant to be a computer. It was meant to be a writing instrument. Thousands were sold to schools around the world. Twenty-six years later it still is an amazing writing device. I bought a stack of six on eBay for about sixteen dollars each. They live in my bag the way a good pen used to. When I sit under the trees or at a picnic table with no signal, I open a file and the words come cleaner. I write some of my X Aricles and some of the ReadMultiplex articles on these machines. Later I connect the AlphaSmart to my computer or iPhone, dump the text, and the modern world takes over only when I am ready for it. The process is so freaking simple. Plug it in and it looks like a keyboard to the computer, than press (SEND) and it literally types out the text like you are typing it. It will likely be compatible decades into the future. I also use the AlphaSmart Neo2 as it is more powerful, ha, you can get more fonts on the screen. It is the rare piece of technology that feels more honest with every year that passes. No planned obsolescence. No forced updates. Just three batteries, a keyboard, and the freedom to think without interruption. Absolutely worth it. If you find one on eBay let me know. The best part is to see some kid’s homework still stuck in the machine from 2000.

Brian Roemmele

33,382 Aufrufe • vor 1 Monat