Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Modern Claudefare Built fully with Opus 5 on High Mode over a few days using principles from Matt Shumer’s Gauntlet Loop - 84,100 lines of code Includes remakes of 4 beloved maps: 0:00 - Rust 0:40 - Highrise 0:57 - Nuketown 1:26 - Terminal Solo & multiplayer (w/...

252,608 просмотров • 1 день назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

one thing that has saved my projects more time than I can count is evals boy was I excited when florian, quite literally an expert in benchmarks, agreed to hop into a ~2h interview to do a walkthrough of what the eval landscape looks like in 2026 (and also answer my personal business questions on the subject) given that now running frontier model through benchmarks is a vector for hacking other systems in order to avoid doing work (looking at you sol), I think it's more important than ever to educate folks on the evals situation. had a lot of fun throughout this session and I hope that you learn a thing or two! enjoy! 🌹 table of content: 0:00:00: are AI Benchmark broken? 0:05:45: Florian Brand background 0:09:00: what motivates florian to work on evaluation? 0:13:33: what is the mirrorcode benchmark about? 0:18:20: cheating in agent benchmark is insaneeeee 0:24:08: LLM benchmarks in era of agents 0:26:30: what’s up with the pelican man 0:28:27: evals are about capabilities 0:31:46: components of running evals 0:35:30: the volume of things to audit is huge!!! 0:40:20: expert answers are wrong hahahaha 0:46:00: api providers aren’t the same 0:48:00: benchmark narrow capabilities (synthetically) 0:50:56: link between eval and environment 0:53:45: small validated benchmark or massive bench? 0:56:11: what is your flow to review a benchmark? 0:58:30: tracking work capabilities with evaluation 1:00:20: slide deck in industry is all vibecoded 1:03:30: harness impact in the evaluation 1:07:39: hardware/sandboxes impact evaluation too! 1:11:00: “is it going to get worse?” 1:12:40: all components influence the final score 1:13:50: training models on different harnesses? 1:17:20: is the model just the weights or it’s all of it? 1:19:30: how to craft benchmark that prevent to cheating and undereliciting models in 2026 1:23:19: ways agents cheat and steal 1:26:00: correct elicitation of capabilities is important 1:36:00: building evaluation on prime intellect 1:45:10: how do you design interactivity benchmarks? 1:48:40: do you think evals are well set to reflect real world performance? 1:52:50: what will the benchmarking landscape will look like in 1 year

Yacine Mahdid

12,185 просмотров • 6 дней назад

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 просмотров • 1 год назад

🚨Our episode with Anne-Laure Le Cunff is now live! Anne-Laure Le Cunff is the founder of Ness Labs and author of Tiny Experiments: How to Live Freely in a Goal-Obsessed World. She's a neuroscience PhD and writes a newsletter that has 100,000+ subscribers. She's used those insights to develop a new model of success — one built around conducting “tiny experiments” that help her build a life on her own terms. She joins me to discuss how we get trapped in cognitive scripts, the hidden dangers of productivity culture, how we can experiment our way to a better life and MUCH more! Timestamps 0:00:00 Intro 0:00:48 Guest introduction 0:05:15 How do you know you are bored out? 0:17:18 People who love us the most might turn out to be our biggest blockers 0:22:20 Don't confuse activity with effectiveness 0:27:39 We will do virtually anything to gain what is really an illusion of control 0:32:16 The map is not the territory, the menu is not the meal. And yet, words are magic spells 0:39:16 The Winner’s Script and the Loser’s Script 0:47:34 "You gotta run at the top speed if you just want to stay in place.” 0:50:27 Let go of the linear and replace it with the loop- a more cyclical approach for growth 0:55:20 Can you sit alone in a room for 15 minutes? 0:58:24 Procrastination is just a signal from your brain that something is not quite working right now 1:20:38 We know nothing 1:24:26 AI is a rocket ship for the mind 1:28:20 In 100 years, nobody will remember you

Infinite Loops 🎙

22,154 просмотров • 1 год назад

Today: a 1.5 hour interview with the co-founders of Coherence Neuro: Ben Woodington and Elise Jenkins They are, as far as i can tell, the only (neurotechnology x oncology) startup that exists today. 'Neurotechnology? For cancer?' you may ask. Yes! As it turns out, tumors interact with the nervous system a fair bit, and you can use the very same neuromodulation toolbox that exists for neuropsychiatric conditions, for monitoring and treating cancer. Coherence has built an invasive device to place at the site of a tumor to do exactly this. Their first indication is a form of brain cancer called glioblastoma; one of the most fatal subtypes of cancer to exist today. The standard of care (with one exception that we discuss) has not changed in 25 years. If Coherence works out, and there is a very real chance they will, that may change. Most interesting of all is that Coherence believes that the bioelectric properties of cancer are not just worth poking at for brain cancers, but for all cancers. And maybe even for diseases outside of it! This conversation covers how Coherence’s first neurotech device (SOMA) works, the molecular reasons behind why neuromodulation affects cancer at all, what the biomarker readouts look like, the obvious Michael Levin comparison, and a lot more. Also: shout to Nicole for setting up the connection here in the first place! Crazy to think that a meeting in mid-2025 ended up leading to this Youtube/Spotify/Apple Podcasts links in replies 0:00:00 - Introduction 0:01:42 - How is SOMA different from Novocure’s Optune? 0:08:57 - Why does neuromodulation affect cancer at all? 0:13:28 - How was cancer-nervous system crosstalk first discovered? 0:15:42 - Anti-epileptics and beta blockers as accidental cancer drugs 0:17:38 - What is molecularly happening when you block cancer-neuron crosstalk? 0:19:50 - What is SOMA actually reading out as a biomarker? 0:20:44 - What does it mean that cancer is “very electric”? 0:22:02 - Can you derive universal biomarkers across patients? 0:23:09 - How is the device placed? 0:24:45 - How does the blocking stimulation regime work? 0:26:43 - Is it fair to say this is closed loop? 0:29:05 - Why not just spam the tumor with constant stimulation? 0:32:31 - Why MRI safety is non-negotiable for oncology devices 0:33:35 - Walk us through the patient journey from diagnosis to implantation 0:36:13 - The Michael Levin question: can you reprogram cancer back to normal? 0:42:29 - Efficacy, hospice settings, and the utility of the neuromodulation literature 0:45:52 - Why start with glioblastoma instead of an easier cancer? 0:48:57 - Regulatory strategy and the reimbursement threat 0:55:37 - How well does mouse-to-human translation work for neuromodulation? 0:55:57 - What do in silico models of neuromodulation look like? 0:58:09 - Why didn’t this exist 10 years ago? 1:01:48 - The founding story 1:06:38 - Why build your own device instead of using off-the-shelf arrays? 1:08:35 - Speaking with glioblastoma patients 1:12:04 - What was it like to raise money for this? 1:13:56 - Beyond cancer: TBI, lung disease, and the pan-disease argument 1:17:40 - Hiring at Coherence + what is the hardest type of talent to find 1:23:17 - What would you do with $100M equity-free? 1:27:15 - Are you a neurotech company or a cancer company?

owl

37,299 просмотров • 5 месяцев назад

I just built a Claude Code plugin that runs your entire Meta ads workflow 🤯 5 skills, one plugin: it spies on competitors, maps your whole category, writes your ad copy, grades it before launch, and audits your live ad account like a $10K agency. All inside Claude Code. And now, with Claude Fable 5, it's even more insane. Perfect for DTC brands, media buyers, and agencies tired of paying for 5 different ad tools that don't talk to each other. If you're spending hours scrolling the Ad Library by hand, exporting CSVs from Ads Manager every Monday, staring at a blank doc trying to write ad variation #14, launching creative and praying it converts... This plugin eliminates the entire loop: → /spy pulls every active ad a competitor is running, ranked by run-time (longevity = proven winners) → /competitors-extractor maps 3-5 brands head-to-head and finds the angles nobody's running → /bulk-creative spins 20 on-brand copy variations off the winning angle → /ad-score grades every ad 0-100 across 6 dimensions before you spend a cent → /ad-matter audits your live Meta account through Meta's official MCP and hands you a prioritized fix list No more Ad Library scrolling. No more $300/mo spy tools. No more launching blind. What you get: → Competitor intel ranked by what's actually been running longest → The open angles in your category nobody is using yet → 20 ad variations in your brand voice, on command → A 0-100 health score on your live ad account with a fix-this-week list Built 100% in Claude Code. I wrote a free step-by-step Playbook that walks you through building this entire plugin yourself: every skill, every script, every scoring rubric. Want access for free? >Like this post > Comment "SCALE" And I'll send it over (must be following so I can DM)

Mike Futia

33,458 просмотров • 1 месяц назад

THIS GUY VIBE CODED A FULL CAPYBARA FOOD DELIVERY GAME IN 2 WEEKS WITH CLAUDE CODE you play as a capybara delivering food on a bike. orders stack on the back, you have a phone with apps in-game and the whole delivery system is realistic 2 weeks, zero game dev experience, and ENTIRELY AI generated the full stack: > claude code for all the code > three.js for the 3D engine > suno for original music > elevenlabs for sound effects and voice > GPT images-2 and grok for textures and illustrations > tripo3d for generating all the 3D assets the cinematics are all in-game too. he asked claude to build a cinematic editor with timeline controls, camera animation, and transitions. then he just placed the cameras himself his workflow was more planning than coding (obviously): > come up with the core mechanic > plan every feature using claude /plan mode > generate assets with AI tools > spend most of his time on the final polish, prop placement, and making the design feel right he said the human part is what most vibe coded games are missing. AI can generate everything but having taste for what looks good and what feels right is still on you the game is playable right now in the browser this is what vibe coding is actually capable of in the game dev space right now a year ago this would have taken a small team of developers, a sound designer, and an artist working together for months now one person with no experience can ship a polished playable game with story, music, and mechanics in 14 days the tools keep getting better and the barrier to making real games keeps getting lower

Om Patel

241,259 просмотров • 3 месяцев назад

WARP SPEED: EPISODE 4 - starring Daksh Gupta, CEO of Greptile Daksh Gupta The core difference between human productivity and every other species? Tools. On the margin, slightly better tools create long-term efficiencies that unlock time for more important things. Daksh, Soohoon Choi , and Vaishant Kameswaran built Greptile into one of the fastest-growing AI code review platforms in the world. $25M Series A led by Benchmark . YC-backed. 16 people. 500M+ lines of code reviewed this month alone for companies like Brex, Substack, and PostHog. The insight: AI coding tools are exploding. Cursor, Claude Code, Devin. Everyone's writing more code than ever. But the systems for validating that code before it ships? Breaking down. Greptile is the independent, centralized validation layer - AI that reviews pull requests, catches bugs, enforces standards. Daksh didn't plan to start a company. He was in senior year when he realized he'd lost his way. Remembered why he chose CS in the first place: to build things. He then convinced his college roommate - "the smartest person I know" - to move to San Francisco with barely enough angel money to survive. Now they're one of the most craft-obsessed teams in SF - using Linear, Claude Code, Cursor, and Warp to ship at warp speed. In this conversation: (0:00) - From losing his way in college to moving to San Francisco (1:24) - Applying to Y Combinator and building the foundations of Greptile (2:37) - Hiring more senior than typical: arbitraging ageism in Silicon Valley (3:40) - Building a company worthy of people who could go anywhere (4:14) - The new bottleneck: producing code is worthless without validating it (4:52) - The Greptile stack: Linear, Claude Code, Cursor, Warp (5:55) - "People talked about Warp with a passion you wouldn't expect for back-office software" (6:17) - The magic of heavy automation: "I'm never in it" (6:42) - Built for engineers, not HR people: "When you're used to confusing software, simple things are just delightful" (7:46) - Scaling from 15 to 50: "Whatever will go wrong, Warp will have functionality for it" (8:55) - Discovering Warp on Twitter: "Why are people so excited about payroll?" (9:35) - "It inspires you to build more intuitive things"

Ayush S

27,984 просмотров • 8 месяцев назад