What the hell does Quant actually do? A model... is trained in BF16, released in Q8, and we use Q4 because we're told 'its nearly lossless'. Fine, but I was curious about the 'nearly' part, what is actually being lost between BF16 all the way down to Q2? As always, its complicated. Needed a way to visualize the loss, so I had a Qwen 3.8 27B BF16 create a paragraph then had every quant below it do the same and marked the mutations. When a model writes a word, it assigns 'how sure' percentage to each. For example, 'The weather tomorrow is *windy*, a model is 90% sure all the words are correct except the windy part, since it can also be sunny, cold, hot, cloudy, etc. Any words a BF16 model was 90% sure of, every quant down to Q2 rarely changes it. On the other hand, any word that it was only 40~50% sure of, such as windy in our example, are the types of data that gets impacted by the quant process. Okay so what, some words changed, how does this impact anything? I set out to find out, and its all below if you are just as weird as I am about these things.show more

Killy
19,592 次观看 • 21 天前
You've probably scrolled past a dozen posts about Jev... this week without anyone telling you what it actually is. It's the first model from TypeSafe, a lab started by one of the researchers behind ChatGPT. It's a decision engine: you give it options, it picks one and tells you how sure it is. It cannot write a single word, and that is the interesting part. Every other AI you use writes. That is the whole interface. So when software needs a plain yes or no, we make a model write a paragraph and then dig the answer back out of it. Fine in a chat window where a human reads it. Bad inside software, where code has to act on it. The bet is that the valuable half was never the writing. It was the deciding. That problem showed up in WordPress years ago, and it is the reason WPVibe works the way it does. The AI does the work. Anything permanent stops and waits, because a delete that skips the trash is not something software should decide on its own.show more

John Turner
21,029 次观看 • 12 天前
watch anon. a 27b model thinking out loud on... gtx 1660 super, i asked bonsai what model it is and it reasoned through the whole answer before it spoke. 20 tokens a second, gpu pinned at 125 watts, 4.25 of 6 gigs used. bonsai is a 1bit quant of qwen 3.6 27b, the king of the 3090 crushed to 3.5gb. turns out the hardware was in your drawer the whole time. you just needed the quant to catch up.show more

Sudo su
26,442 次观看 • 2 个月前
And finally, my 3D model of @Lord_Griselda's Medusa is... done! This was a lot of fun to do, I've actually had it finished for a couple of weeks, but I also had to finish the video for it, where I walk through the entire process which you can find below! #blender #b3dshow more

Niall
100,304 次观看 • 2 年前
Without googling it… how many of you actually know... what a butter bell is? We’ve had one on our counter for years. I just learned what it does and how to use it. My wife made it very clear that I am a savage I might be but was wondering if I was the only oneshow more

John Asghar
651,923 次观看 • 9 个月前
today was the first time i was genuinely impressed... with what AI can do i recently decided to buy a whole FPV drone setup knowing basically nothing about the hardware side of it there's a pretty steep learning curve even just to set everything up properly: radios, RF protocols, flight controllers, ESCs, firmware, batteries, goggles, betaflight configs etc as someone that spends essentially 12h a day prompting agents to build software, it's actually pretty rare that i interact with AI on something where i have zero idea what's going on under the hood, and i never really used it for debugging a bunch of physical devices that all have to talk to each other i had codex + voice mode open for basically the entire setup. told it everything i bought, sent it some pics and then just started talking to it >what order do i set all this up in >how do i change this setting on the radio >which of these cables do i use >the drone is flashing pink wat mean >can you make this thing less insane to fly in my apartment and it was surprisingly seamless it would go find the manual for whatever specific thing i was holding, tell me exactly which buttons to press, what port to plug something into, what i should see if it worked etc then when i got to configuring the actual drone i had codex running on the computer it was plugged into, so it could inspect the config, back everything up, change settings, send usb reboot signals and check what happened the insane thing about voice mode is that youre literally hands on with the hardware and just telling codex what it should do, i literally never touched a thing on the computer besides starting voice mode if something doesn't work you tell it what happened and keep going a few hours of this and i had the radio, goggles, charger, batteries, drone firmware and betaflight all set up and had actually flown the thing the part that stuck with me is that i also understood what most of it was doing by the end, every time there was a term or tech i didnt understand id just ask to explain there is something absolutely magical about having proper real time personalized assistance, being able to dump a pile of unfamiliar hardware on your desk and have something figure out exactly what you own and walk through it with you in real time you become the missing physical link pressing the buttons i think spending all day using coding agents has actually made me pretty numb to AI progress. every new model is a bit better at some benchmark or can oneshot some task that the previous one couldn't and you just kinda adjust to it this felt different mostly because i had no existing knowledge to fall back on for the first time the jarvis comparison didn't feel cringe ai for coding and general computer tasks is cool and all but this feels a lot closer to the endgame anyone should be able to just ask any question about whats going on in their life and have realtime support i wonder if more hardware products will actually start exposing some sort of MCP or interface for agents to plug into thinking for example of how elevators in china are increasingly built with interfaces that let delivery robots call them directly instead of having to physically press a button we might actually start seeing hardware design shift from being purely human-interface-first to also being agent-interface-first buttons, screens and menus exist because humans need some way to tell machines what to do. agents don't necessarily need any of that if the hardware exposes an interface directly very curious which side closes the physical world gap first: humanoid robots that can operate hardware designed for humans, or hardware adapting so agents can operate it directlyshow more

ultra
18,241 次观看 • 1 个月前
Alright I've never seen this model handle a storm... of this size, so not sure how well it will do...but for any who want an idea here is what the "Fox Model" is showing though 1pm Sunday (as far out as it goes). It does give you an idea of the scale of this system. Big time storm.show more

Mike Thomas
2,169,220 次观看 • 8 个月前
After a few more hours, I think I've figured... out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus 5 or contextualize it won't work the same way as with other models. It loves exploring, so it doesn't need much guidance for it. Unique preferences, artifacts, and references compliment it well and enable cleaner and more effective exploration and execution. Now that it can explore more effectively on its own and understand intent better, the best thing to do is to get out of its way (e.g., it doesn't need examples of your preferences; a clear high-level description of it works best). It's truly agentic in that sense. A good first step to provide better context for Opus 5 is to distinguish between what's situational and what needs persistence. Regardless, persistent system prompts and CLAUDE.MD needs to stay lightweight. Remove memories and tool descriptions from these. CLAUDE.MD is also a great place to tap into progressive disclosure by linking command/skills to it. On the situational side, agent skills and auto-memory can leverage progressive disclosure and the improved ability of the model to use its external context/knowledge. Conflicting and unnecessary instructions, which are common at this layer (mainly to ensure reliability), are going to throw off this model easily. That's the biggest change I had to make. Simple, clean, and clear prompts and skills work best. I had to clean a lot of my skills and system prompts. The way I prompt remains the same (usually clear and well-scoped). MCP tool descriptions are also more descriptive and have been deduped from the system prompt. Anthropic released a guide on the new rules for context engineering, which was helpful here. I started to test the recommendations and created a little artifact with the things that worked along the way. This might feel like a lot of work. Believe me, it has been frustrating. But I think we can expect future frontier models to become more agentic and smarter at figuring out the right context/gaps. The best thing to do is to prepare for that now. Boris Cherny mentioned that Opus 5 is their least prompt-injectable model yet. I am not sure if that was something they intentionally trained for or if it emerged based on how it was trained, which is to be extremely agentic in nature and more direct in execution.show more

elvis
37,824 次观看 • 2 个月前
You guys, I am at a loss of words.... This """puzzle""" is literally as simple as match the symbol on the shield to the only other shield with a symbol, AND THE GAME STILL HAS AN NPC TELL YOU HOW TO DO IT. These developers think you have an IQ below that of a goldfish.show more

Narwitz
609,224 次观看 • 1 年前
Earlier today I was ironing some clients clothe with... the gas iron when I had the sudden urge to use the toilet. I turned off the gas and went to the toilet. I went about with other things when I came out thinking the iron was off. Next thing I heard was Mummy your table is burning. I rushed down to find that the gas was still on and I had only increased the heat instead of turning it off. What amazed me was that the rubber iron mat that came with it didn't burn but everything beneath it melted. Does anyone know the kind of material they use in making the mat? Alhamdullilah, what if it happened during the week when my kids are not around? Why is it so hard to fix the power problem in this country so we don't have to resort to endangering our lives trying to solve problems that we shouldn't be having in 2026? In all, Alhamdullilah.show more

Halimah Ahmed (Tiwal'adire)
123,766 次观看 • 3 个月前
What happens when an autonomous robotaxi gets into an... accident? So far, nothing. Yesterday, I rode in a Zoox robotaxi and got hit by an RTC bus in front of New York New York. The Zoox was trying to turn right into New York New York on Tropicana Avenue just after crossing Las Vegas Boulevard and the bus was trying to get over to make its stop. What I saw, was the RTC bus merging into the front left of the Zoox robotaxi and then a crunching sound followed by a grinding sound. Pretty sure it was the RTC bus at fault. Impact detected. The Zoox robotaxi stayed in place while a warning popped up on the user display. Zoox support came over the in-car speaker and asked if I was okay and then eventually had me exit the vehicle at the New York New York rideshare area. Then, it just took off. The RTC bus? It also took off. The driver never got out to check to see what happened. Loaded up some passengers and went about the day. Nobody stayed around to file a police report, so I did. The Nevada State Troopers are scratching their heads at what to do here and so am I. Now I have a lot of questions that won’t get answered probably. I am okay and not injured thankfully, just shaken up a bit. As far as riding Zoox again, I feel it is safe to ride and highly recommend trying it out if you haven’t. On a side note: The Zoox took the hit pretty well and I couldn’t see damage when I checked it out. The bus had a nice scrape on its right wheel fender. I will keep you all updated on this.show more

Chris Holmes
94,476 次观看 • 6 个月前
Critical Disengagement on FSD v14.3.7. Context: The car needed... to leave the parking garage and go to HEB. It started from the second floor, made its way to the top floor, and I believe it tried to find a path through this metal wire fence. If it had gone through that, it was a 6 story drop to the ground. In the dashcam clip, it is clearly visible. You can see the car is at 4 mph throughout the whole turn and is not slowing down. It wasn't trying to park because the route was set to leave the parking garage, not park in it. Either way, driving at 4 mph inches from the fencing is not okay. I have an interior GoPro recording of the whole situation. It actually tried to drive into the fencing again after I restarted FSD. In thread 👇 Obviously since I disengaged, we don't know what it truly would have done. Maybe it would have made contact and stopped. This was as far as I was going to take it as someone responsible for the vehicle, so I believe this disengagement was fair. I was not expecting something like this as the latest versions of FSD are exceptionally safe. Tesla AI Tesla Ashok Elluswamy Elon Muskshow more

Aryan Butala
920,837 次观看 • 1 个月前
The discourse around the Eze penalty last night is... fascinating. If nothing else it provides Arteta with an excuse as to how his team has been hard done by and robbed. He loves excuses. It’s a fascinating situation because I actually think the right outcome was reached albeit through the wrong process. Once given on the field under the current rules it shouldn't have been overturned. For that Arsenal can feel aggrieved. Don't forget Arsenal were the beneficiaries of an incredibly soft penalty in Leverkusen. It wasn’t overturned. I find the attached video interesting because it clearly shows how minimal the contact is. The defender doesn’t pin Ezes foot to the ground. Doesn’t smash into it. It brushes down the side of it. It’s where the still pictures of it were wildly misleading. Are we really saying that contact such as that is worthy of a penalty? What’s clear in the video is that Eze has absolutely made the most of the slight contact. I love the straight left leg. A very natural position. It’s a brilliant dive and normally it would have been rewarded. By saying that is a stone wall penalty all we are doing is encouraging such theatrics and cheating. Whenever you hear the words “that was clever” by a commentator they are intimating the player has cheated/made the most of it. We complain about refs all the time but what chance do they have when players are doing stuff like this. Until players are punished properly there is no disincentive for players to keep doing this stuff. I also find it fascinating when you have a foot brushed like this. When you compare it to all the holding/pushing/grappling at other times. Sometimes a contact sport and others not. Is this really what football has become?show more

Luke Paton
104,573 次观看 • 5 个月前
🚨 Ajinkya Rahane revealed the story behind his decision... to send Yashasvi Jaiswal off the field. 🗣️ " I sensed the situation and also I felt it was going out of hand, both the players were going out of hand and it was my responsibility to control my teammates. You respect the umpires, the officials who are on field, who are controlling the game, who are officiating the game and also the situations. Mistakes can happen to anyone, but you, as a player, you be in that limit. So, I sensed the situation, what can happen next, it was my instant reaction to send him off the field. As a player, I also felt that he must have felt bad, I said, okay, if he felt bad about me, it's okay. But I think, that point in time, it was a good thing for him. The thing which I told him about, go in from the back, put ice on your head and come back "show more

OldMonkofCricket
21,540 次观看 • 1 个月前
Reading the comments on this post told me everything... I need to know about art. So many of you wrote to say you felt the exact same thing I did standing in front of the Starry Night. That strange pull, the sense that the painting is somehow alive... It turns out the feeling is completely universal. I am not someone who gets emotional easily. But every so often a piece of art really gets to me. And that is exactly what happened the first time i walked up to Van Gogh's magnum opus. I know how that sounds, it is one of the most famous paintings on earth, almost a cliché to be moved by it. But maybe it is that famous for a reason... It genuinely looks like it is moving while you stand there, shifting in a way I cannot explain and have never seen in any of the thousands of paintings I have looked at in my life. The same day I took the photo below, I also filmed this short video. It still comes nowhere near the real thing, but you can at least catch a little more of it here, the thick ridges of paint, the way the stars and swirls lift right off the canvas. This technique is called "impasto", from the Italian for mixture. It means paint laid on so thickly that every brushstroke and ridge stays visible on the surface. Van Gogh piled it on with a loaded brush, building the stars up thick so they would blaze as bright as possible against the dark. But no video and no explanation of technique can really convey what it feels like to stand in front of it. And the fact that so many of you felt the very same thing says something true about what art is really for. Tolstoy said it best: "The receiver of a true artistic impression is so united to the artist that he feels as if the work were his own and not someone else’s, as if what it expresses were just what he had long been wishing to express. A real work of art destroys, in the consciousness of the receiver, the separation between himself and the artist. In this uniting of it with others lies the great attractive force of art."show more

James Lucas
57,165 次观看 • 22 天前
So, my opinion on what the Antarctic (Antarctica) Anomaly... is that it's a type of frequency technology. It must be way more powerful than HAARP, as many have claimed it to be, because we would see these anomalies at other HAARP sites, and we don't, not like this. With that said, and I'm very much trying to avoid letting what I want it to be not play a part here, I think it is a technology that is being used either off the coast of Antarctica itself or Bouvet Island. A third possibility is an area just to the northwest of the island that looks odd. It's possible it is a sonar scan from a ship, but why in that remote location? It looks like an antenna set up or rows of something that is out of place. I also believe that the weather events and fires that have taken place in Africa could possibly have been because of this. Each time we saw the anomaly, it was followed by a destructive weather event in Africa. A weird connection to that is we have been told and warned of a very busy 2024 Atlantic hurricane season. This is in part because of the above-average Atlantic ocean temperatures, which is the fuel to Hurricanes. With all this info, it's possible to see how the Anomaly could be a frequency tech that can manipulate or create weather, And or WARM up the Ocean temps to purposely enhance the Hurricane season and Storm growth. Keep in mind that many of our hurricanes and many of the biggest hurricanes have come from the west coast of Africa and form over the Cape Verde islands before heading towards the Caribbean and the United States. This is all of course speculation, and I'm learning many new things every day, so this idea may morph over time as we learn more. In the end, it is very hard to ignore all these findings. #antarctica #anonaly #AntarcticaAnomaly #BouvetIslandshow more

In2ThinAir
442,580 次观看 • 2 年前
AI has had exactly two scaling axes that worked... so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inferenceshow more

Sasha Malysheva
15,053 次观看 • 1 个月前
how #karina was selected as a cameo in g-dragon's... mv 👥️ *explaining how the role of a female cast is needed in the mv* 👤 the female cast will appear during that part, but i'm not sure who am i going to cast 👩 it will work out if we can cast someone, but if we can't then.. 👤 you're right, we should have someone for it. because it's a (pair) dance choreo that i need to dance with a woman.. 👤 karina is the only one i can think of (to cast for the role)show more

♡
172,566 次观看 • 1 年前
Jev + SERV is actually insane. We already showed... you can increase Jev's performance with SERV Reasoning. Now we're taking it further, bringing Jev-powered Decision nodes into Graph Sharding with the upcoming SERV v3. Here's a breakdown of how it works: Jev is a decision-making model. Given a task and a set of options, it predicts which path is more likely. Think of the octopus that predicted World Cup results. Jev does that for your business, except it's not luck. It weighs every option and tells you how sure it is. It does this by assigning probabilities to outcomes. It doesn't generate text on its own, so you can't expect it to create a new outcome for you. But that's also what enables it to be lightning fast and dirt cheap. For example, in customer service you can ask Jev how to triage an incoming query and route it to the correct department. It can only select from the list of departments you provide it. This also means it can't hallucinate a new outcome outside the options it's given, which makes it incredibly interesting for OpenServ. In Graph Sharding, we take a single system prompt and break it down into multiple LLM steps with deterministic input and output shapes. Some of these steps require an LLM to produce new output, while others are simply decision routers that determine the next possible path. Traditionally, LLMs are slow and expensive. Breaking a single prompt into multiple steps increases accuracy and reliability by a ton, but it also introduces latency. Jev takes on those decision nodes, which are the backbone of a business process and therefore SERV graphs, and makes them super consistent and lightning fast, lowering the overall cost and latency of graph execution. SERV Reasoning on its own is a great force multiplier for Jev because, like all other models, it works by interpreting input instructions. The clearer those instructions are, the better the model performs. That's where SERV Reasoning comes into play. Just like amplifying any other model, we also amplify the accuracy and consistency of Jev's responses. And now we're bringing Jev-powered Decision nodes into Graph Sharding with SERV v3.show more

Armagan Amcalar
365,108 次观看 • 6 天前