Загрузка видео...

Не удалось загрузить видео

На главную

I've been testing Kimi K3 to see if it could solve the Exfiltrating Sensitive Data via Server-Side Prototype Pollution lab (Expert difficulty) on PortSwigger. K3 successfully solved the lab in under 1.2hr 1 shot without any guardrails when on work. It's an good model for certain use cases. Compared...

32,704 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Hate to be that guy but it needs to be said. Tesla has not yet solved FSD and there is no guarantee it will be solved by June Don't misunderstand me, I'm not predicting it won't be ready. I'm simply reminding a specific group of people that seem to be operating under the assumption it's already a done deal If it was truly solved, we would have Unsupervised. Yes, it's incredible 99% of the time but cracking the code for the final 1% is excruciatingly difficult While we're on the subject - does FSD pull out in front of people and then take its good old time for anyone else? This is a behavior I'm really not a fan of and seems to happen regularly for me on 13.2.8 and 13.2.7 (AI4) We still don't have answers for sun glare - I watch an unhealthy amount of FSD videos and I still see this happening more than I'd like to Yes, snow is another story but initially this isn't really part of the rollout plan for TX and CA. Makes sense to solve for better conditions first Lane selection is still an issue where I'm at as you can see in this clip. Wish I would have had my camera ready for this but I was using FSD and the navigation just needed to go straight through the intersection. Instead, for no reason, it got into the left lane (left turn only) and stayed there. When I got the green arrow it just sat at the light (luckily no one was behind me so I stayed for a bit) and I eventually took over and just made the left turn You'd think basic things like this would be solved by now. Maybe it's the map data but even if it is, that's still a problem This will undoubtedly get misinterpreted and that's fine, but I've felt like some caution around Unsupervised would be healthy for certain pockets of the $TSLA community I'm still as confident as ever that Tesla will be the first to solve for generalized autonomy and I don't think it's close (especially in the US), but there's at least a non-zero percent chance it happens later than this June

Dillon Loomis

87,214 просмотров • 1 год назад

Chinese AI models are wiping billions off Big Tech right now. Google just lost $200 billion in a single day, and the model it needed to fight back still isn't ready. Gemini 3.5 Pro, Google's most powerful model, is months behind schedule. Alphabet stock dropped 4.4% that same day. The Deepseek moment is happening again, and the new model is FAR bigger. On the same day Google's delay leaked, a Beijing lab called Moonshot released Kimi K3. It is the largest open model ever built, with 2.8 trillion parameters. It took the number one spot on the Frontend Code Arena, a live coding leaderboard, passing Anthropic's best model. And Moonshot is giving it away for free on July 27. The genius part: Anyone with enough computers can download it and run a frontier level AI without paying a cent to a US company. A single task on Kimi K3 costs about 94 cents. The same work on some American models costs nearly double. So why would a company keep paying premium prices for a model it can now get for free? The entire US AI business is built on selling access to models that cost billions to train. If a free Chinese version does most of the same work, that pricing power starts to crack. And Kimi is close to the best. On one closely watched intelligence ranking it scored 57, just behind the top American models GPT-5.6 Sol and Fable 5, and ahead of Claude Opus 4.8. Bank of America told clients that Kimi proves Chinese labs can keep making big leaps even with limited chips. And the founder of Moonshot, Yang Zhilin, learned to build AI as a researcher INSIDE Google. Google literally wrote the 2017 paper that made all of these models possible. Now the people who studied its work are using it to destroy Google, and handing it out for free. What happens next: Kimi K3's weights go public on July 27. Google reports earnings on July 22, and everyone will be asking the same question about Gemini. If free models keep topping the charts, every valuation built on paid AI access has to be rewritten. What do you think?

Ricardo

47,790 просмотров • 1 месяц назад

BREAKING: Anthropic just dropped Fable 5.1—and CLAUDE IS SO BACK. We’ve spent the last week testing it at Every 🪨 across coding, writing, and knowledge work. Our verdict: It's finally Fable for everyone. It’s the strongest coding model we’ve used, but now it's fast, token-efficient, and CRUCIALLY actually speaks like a normal person. Here’s our vibe check: - A monster at coding. Kieran Klaassen rebuilt a working version of Proof, our document editor, from one prompt. It added useful details he hadn’t requested, and it handles enormous coding jobs that run for days at a time. It built a computer use Mac app for me called Hands in one-shot that other models failed at. - A Claude our writers want to use again. It has clearer prose, fewer AI tells, and it takes an edit without arguing. It's a significant upgrade over Opus 5. And won Katie Parrott's heart back. - About half the tokens as Opus 5, and much faster. In our Slack-agent tests, it delivered comparable results to Opus 5 using about half as many tokens, in about 60% of the time. - Knowledge work you can delegate. It can produce great knowledge work—like slide decks—end to end without making slop. And flew threw hammer's tests with flying colors. - It now supports zero-data-retention agreements. Now businesses can actually use it! A big barrier to Fable adoption is gone. Net Result: It's obviously an Opus 5 killer. If that was your daily driver you should switch today. If you're using GPT-5.6 in ChatGPT for Work, it's spinning the wheels on for big delegated tasks. I still use ChatGPT for Work more day to day, but I use way more tokens in Fable 5.1. I send it off at the beginning of the day to do big programming projects, like end to end MVP builds, and check in every once in a while. State of Play: The big knock on Anthropic was they built a supergenius in a datacenter that was almost unusable. It was too slow, argued back, and talked in technical gibberish. They've managed to solve those problems and more with Fable 5.1!

Dan Shipper

202,709 просмотров • 6 дней назад

BREAKING: GPT-5.6 Sol is out—AND Codex has been merged into ChatGPT Desktop as ChatGPT Codex. This combo model and desktop app harness are the gold-standard for knowledge work in AI. 5.6 is powerful, fast, half the price of Fable, and my default for almost everything. We’ve been testing it internally Every 🪨 for about a month across coding, writing, design, and knowledge work. Here’s our day-zero vibe check: - An A-tier coder—but it’s not Fable. Sol scored 56/100 on our Senior Engineer benchmark compared to a 91 for Fable. I think the 56/100 undersells it, it's an excellent implementor, and very smart. But Fable just writes conceptually cleaner code and works better at the top end of task complexity. PRO-TIP: Use GPT-5.6 as Fable's subagent for the most goated combo in AI coding. - The best writer of the frontier models. It’s clearer and more concise than Fable or Opus 4.8, without the overexplaining or weird private language. It can one-shot marketing emails, help you workshop taglines, and explain complex concepts clearly. It's also super fast, which makes it easy to collaborate with. - Design is better, but not top-tier. It has noticeably more taste than 5.5, but Fable and Opus 4.8 are still playing at a different level. See examples in the video and vibe check below. - The real leap is knowledge work. Sol is the first model I’ve trusted to run whole loops of knowledge work—not just help with individual tasks. I use it to process email, surface decisions from meetings and Slack, find job candidates, scan Facebook Marketplace for furniture, and log my meals. It has shifted my job from doing the work to tending the system that does it. - The merged app is fine. I was extremely worried about this because I love the Codex app. OpenAI was caught in an interesting position: How to make an agent orchestration app for regular ChatGPT consumers, coders, and businesses all in one app. They now split the interface between ChatGPT Work and ChatGPT Codex. They're basically the same except Work hides code. And "Chat" has been demoted to 2nd tier status for quick questions in either one. It's not a big leap, but it's not a huge setback either. And it remains my favorite of the desktop agent orchestration apps. Verdict: If I really had to put my finger on it, I'd say Fable has way more big model smell. But that means it's a skill in itself to get value out of it—99% of people are still not there yet. GPT-5.6 is almost as powerful, but is easy to use, fast, and relatively cheap. It should give you an early sense of where model work is going. Full Every 🪨 Vibe Check:

Dan Shipper 📧

146,015 просмотров • 2 месяцев назад

There’s been times when I’m walking or hiking in the woods and I hear knocking coming a building, it gives me an eerie vibe like a trapped in the woods kind of feeling. I know I post about this stuff a bit but it truly fascinates me. Like how this man is hearing knocking on the ground in the woods like something underneath him is knocking but it’s solid ground right? The only logical explanation I can think of is it’s reverb from construction or possibly pipes on the ground being worked on. I was at the state park getting a jog in, the surrounding area was flat, I could see far on either side of me. The fields were clear of trees like a new developmental lot. I didn’t have my headphones on, I decided just to go for a jog without them. During the jog it was quiet, except for the sounds of the wildlife, I could tell they were there, it didn’t bother me. However as I progressed I started hearing a screeching noise, it didn’t sound like any animal to me, almost like old metal on metal brakes of a car ready to fail is how I would describe it. I wasn’t worried because as I said before, the area I was at was clear, it wasn’t as if anybody could be hiding along my path, it was just empty, but what made it weird is I couldn’t pinpoint where the noise was and it sounded close yet I couldn’t see anything that could be the source. Just a dirt path and grass. There were trees along the horizon but they were far, too far for that noise to be so close. For the rest of the job I kept my eyes open but to this day I still couldn’t explain what that noise was.

SonnyBoy🇺🇸

23,312 просмотров • 2 месяцев назад