Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I've been testing Kimi K3 to see if it could solve the Exfiltrating Sensitive Data via Server-Side Prototype Pollution lab (Expert difficulty) on PortSwigger. K3 successfully solved the lab in under 1.2hr 1 shot without any guardrails when on work. It's an good model for certain use cases. Compared...

32,611 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Hate to be that guy but it needs to be said. Tesla has not yet solved FSD and there is no guarantee it will be solved by June Don't misunderstand me, I'm not predicting it won't be ready. I'm simply reminding a specific group of people that seem to be operating under the assumption it's already a done deal If it was truly solved, we would have Unsupervised. Yes, it's incredible 99% of the time but cracking the code for the final 1% is excruciatingly difficult While we're on the subject - does FSD pull out in front of people and then take its good old time for anyone else? This is a behavior I'm really not a fan of and seems to happen regularly for me on 13.2.8 and 13.2.7 (AI4) We still don't have answers for sun glare - I watch an unhealthy amount of FSD videos and I still see this happening more than I'd like to Yes, snow is another story but initially this isn't really part of the rollout plan for TX and CA. Makes sense to solve for better conditions first Lane selection is still an issue where I'm at as you can see in this clip. Wish I would have had my camera ready for this but I was using FSD and the navigation just needed to go straight through the intersection. Instead, for no reason, it got into the left lane (left turn only) and stayed there. When I got the green arrow it just sat at the light (luckily no one was behind me so I stayed for a bit) and I eventually took over and just made the left turn You'd think basic things like this would be solved by now. Maybe it's the map data but even if it is, that's still a problem This will undoubtedly get misinterpreted and that's fine, but I've felt like some caution around Unsupervised would be healthy for certain pockets of the $TSLA community I'm still as confident as ever that Tesla will be the first to solve for generalized autonomy and I don't think it's close (especially in the US), but there's at least a non-zero percent chance it happens later than this June

Dillon Loomis

87,214 Aufrufe • vor 1 Jahr

Chinese AI models are wiping billions off Big Tech right now. Google just lost $200 billion in a single day, and the model it needed to fight back still isn't ready. Gemini 3.5 Pro, Google's most powerful model, is months behind schedule. Alphabet stock dropped 4.4% that same day. The Deepseek moment is happening again, and the new model is FAR bigger. On the same day Google's delay leaked, a Beijing lab called Moonshot released Kimi K3. It is the largest open model ever built, with 2.8 trillion parameters. It took the number one spot on the Frontend Code Arena, a live coding leaderboard, passing Anthropic's best model. And Moonshot is giving it away for free on July 27. The genius part: Anyone with enough computers can download it and run a frontier level AI without paying a cent to a US company. A single task on Kimi K3 costs about 94 cents. The same work on some American models costs nearly double. So why would a company keep paying premium prices for a model it can now get for free? The entire US AI business is built on selling access to models that cost billions to train. If a free Chinese version does most of the same work, that pricing power starts to crack. And Kimi is close to the best. On one closely watched intelligence ranking it scored 57, just behind the top American models GPT-5.6 Sol and Fable 5, and ahead of Claude Opus 4.8. Bank of America told clients that Kimi proves Chinese labs can keep making big leaps even with limited chips. And the founder of Moonshot, Yang Zhilin, learned to build AI as a researcher INSIDE Google. Google literally wrote the 2017 paper that made all of these models possible. Now the people who studied its work are using it to destroy Google, and handing it out for free. What happens next: Kimi K3's weights go public on July 27. Google reports earnings on July 22, and everyone will be asking the same question about Gemini. If free models keep topping the charts, every valuation built on paid AI access has to be rewritten. What do you think?

Ricardo

47,790 Aufrufe • vor 1 Monat

BREAKING: GPT-5.6 Sol is out—AND Codex has been merged into ChatGPT Desktop as ChatGPT Codex. This combo model and desktop app harness are the gold-standard for knowledge work in AI. 5.6 is powerful, fast, half the price of Fable, and my default for almost everything. We’ve been testing it internally Every 🪨 for about a month across coding, writing, design, and knowledge work. Here’s our day-zero vibe check: - An A-tier coder—but it’s not Fable. Sol scored 56/100 on our Senior Engineer benchmark compared to a 91 for Fable. I think the 56/100 undersells it, it's an excellent implementor, and very smart. But Fable just writes conceptually cleaner code and works better at the top end of task complexity. PRO-TIP: Use GPT-5.6 as Fable's subagent for the most goated combo in AI coding. - The best writer of the frontier models. It’s clearer and more concise than Fable or Opus 4.8, without the overexplaining or weird private language. It can one-shot marketing emails, help you workshop taglines, and explain complex concepts clearly. It's also super fast, which makes it easy to collaborate with. - Design is better, but not top-tier. It has noticeably more taste than 5.5, but Fable and Opus 4.8 are still playing at a different level. See examples in the video and vibe check below. - The real leap is knowledge work. Sol is the first model I’ve trusted to run whole loops of knowledge work—not just help with individual tasks. I use it to process email, surface decisions from meetings and Slack, find job candidates, scan Facebook Marketplace for furniture, and log my meals. It has shifted my job from doing the work to tending the system that does it. - The merged app is fine. I was extremely worried about this because I love the Codex app. OpenAI was caught in an interesting position: How to make an agent orchestration app for regular ChatGPT consumers, coders, and businesses all in one app. They now split the interface between ChatGPT Work and ChatGPT Codex. They're basically the same except Work hides code. And "Chat" has been demoted to 2nd tier status for quick questions in either one. It's not a big leap, but it's not a huge setback either. And it remains my favorite of the desktop agent orchestration apps. Verdict: If I really had to put my finger on it, I'd say Fable has way more big model smell. But that means it's a skill in itself to get value out of it—99% of people are still not there yet. GPT-5.6 is almost as powerful, but is easy to use, fast, and relatively cheap. It should give you an early sense of where model work is going. Full Every 🪨 Vibe Check:

Dan Shipper 📧

145,911 Aufrufe • vor 1 Monat

There’s been times when I’m walking or hiking in the woods and I hear knocking coming a building, it gives me an eerie vibe like a trapped in the woods kind of feeling. I know I post about this stuff a bit but it truly fascinates me. Like how this man is hearing knocking on the ground in the woods like something underneath him is knocking but it’s solid ground right? The only logical explanation I can think of is it’s reverb from construction or possibly pipes on the ground being worked on. I was at the state park getting a jog in, the surrounding area was flat, I could see far on either side of me. The fields were clear of trees like a new developmental lot. I didn’t have my headphones on, I decided just to go for a jog without them. During the jog it was quiet, except for the sounds of the wildlife, I could tell they were there, it didn’t bother me. However as I progressed I started hearing a screeching noise, it didn’t sound like any animal to me, almost like old metal on metal brakes of a car ready to fail is how I would describe it. I wasn’t worried because as I said before, the area I was at was clear, it wasn’t as if anybody could be hiding along my path, it was just empty, but what made it weird is I couldn’t pinpoint where the noise was and it sounded close yet I couldn’t see anything that could be the source. Just a dirt path and grass. There were trees along the horizon but they were far, too far for that noise to be so close. For the rest of the job I kept my eyes open but to this day I still couldn’t explain what that noise was.

SonnyBoy🇺🇸

23,312 Aufrufe • vor 1 Monat

I got to try Grok 4.5 in early access in Cursor for the past few days and I absolutely enjoyed it. It feels like Opus 4.8 at 2x the speed at a much cheaper price point. I tasked it to brainstorm > plan > implement a big feature for my game (this act 1 boss fight) and it did not disappoint. - It is much smarter than Composer 2.5, during planning mode, it is able to think through my request more robustly, ensuring that edge cases are covered and makes sure to ask the right questions to confirm with me first. - It is much better at brainstorming ideas/suggestions, similar to Opus 4.8, though I think Fable still edges out a little when it comes to brainstorming ideas and suggestions - It is FAST. probably the fastest of all frontier models (Opus 4.8, GPT 5.5 etc), which makes it a joy to build with, because I can stay in the flow - It has much improved visual/animation capabilities than Composer 2.5, it can code up animations (i wanted an explosion animation with particle effects) with much, much better visuals, animation movement and timing. This is a big leap and I was so happy to see this improvement. - The best part for me is that I can just use the same model from planning down to execution without switching to a lower cost model because the price point is cheaper than other frontier models. I'll be testing this model with more challenging tasks in the next few days but I think this is going to be my main driver for vibe coding for a while. Also, its nice to see Grok back in the race. 🙌

Danny Limanseta

1,421,893 Aufrufe • vor 1 Monat