Revisiting this game a few months later: (i) It's... pretty impressive that general-purpose image generators can estimate meaningful depth from a single photo. (ii) Monocular depth remains pretty awful for 3D reconstruction! Here Gemini > GPT > Qwen. (At least I hope so!…🫣)show more

Keenan Crane
34,217 Aufrufe • vor 4 Monaten
I had an early access to GPT-6 Astra and... I can say, everyone in the world now has a 3D designer at their fingertips. I gave it an image of a house and asked to create it in 3D with all the details including toys, appliances and furniture. This a full 3D model reconstruction in Blender with geometry that you can manually tweak and run 60fps as a locally rendered "game" on device.show more

Tom Krcha
1,433,857 Aufrufe • vor 9 Tagen
A preview of what's next, visualized with Rerun and... PlayCanvas supersplat ✨ (Also, feel free to send me a DM 📩; I’ll be in San Francisco from July 21–29, and I'm looking to meet like-minded folks!) I'm convinced that Gaussian Splats will be an integral part of any data engine as an underlying representation. So I've started putting together a repo that: 1. Given a single image, perform image outpainting 🖼️🖌️ 2. Estimate a monocular depth map on the outpainted image 📏 3. Train a Gaussian Splat initialized from the monocular depth 🎓✨ 4. Warp to new views, perform inpainting on the missing masks -> Train new splat 🔄🎨 This is going to be integrated into exo-egoforge, but I wanted to start with the simple single-image version before moving to a multi-video implementation There's some weirdness in the final rerun visualization, but the trained splat looks great 🎉! This is all based on the very cool VistaDream paper ( .github.io/) More on this next week!show more

Pablo Vela
26,036 Aufrufe • vor 1 Jahr
idk what where this person is, and i know... it's been almost a year, but i still cant help but watch this video and get extremely upset in general NOT AT THEM TO BE CLEAR, they can do whatever (and im pretty sure already did), but it's just sad to see that fade away from someone ykshow more

Delta !!
23,657 Aufrufe • vor 4 Monaten
Two months ago, at an RSS workshop panel, I... said that general real2sim reconstruction is still far from solved, especially for articulated objects, which will be a major bottleneck for robotics. But now that has totally changed with Astra; using several photos as input, we can easily reconstruct vivid articulated scenes. The result is so impressive! Excited to see how it will shift the paradigm of real2sim2real and agentic learning for robotics and humanoids.show more

Siyuan Huang
18,339 Aufrufe • vor 3 Tagen
Taehyung sound bites from BTS comeback live is being... used extensively by kmedia news channels. 🐯it’s been quite a while right? We waited so long for this moment, and now I really feel like it is 'finally' here! 🐯While it is Korean, the word ‘Arirang’ itself is very lovely and meaningful, and it has a lot of depth. So beyond that, there are also many things we can express through it.show more

Taehyung Naver
14,150 Aufrufe • vor 5 Monaten
Some updates on the multiview vistadream pipeline with Rerun!... Rerun came in extremely useful here, as being able to visualize depths at each stage of the pipeline allowed me to debug some nasty bugs. Since the last time, I was only working with a single image input. I've added in VGGT as my multiview pose + depth estimator. It works REALLY well for getting camera poses, but the depths are not that great. To try and fix that, I estimated depth maps from MoGeV2 for each of the views, and scale+shift aligned them so that they would match up to the confident sections of VGGT's depth predictions. You can see in the video just how much sharper the visualized 2d depth maps are! The biggest issue continues to be the multiview consistency 🫠 That's up next, along with actually training the Gaussian splat. Lots of work went into actually understanding inputs+outputs for VGGT. I had some funky bugs where the confidence values would all collapse to true I'm also really excited for this pipeline to use Difix3D+ Nvidia instead of Flux Inpainting, it seems like a better suited for a multiview pipeline.show more

Pablo Vela
29,904 Aufrufe • vor 1 Jahr
I bet you've never noticed this ultra-cool but rather... complicated "cone of prediction" at the center of Windows context menus? It allows the system to predict that you're aiming for a flyout. Well, that's because Windows (at least in the old days) didn't do any such fanciness. It just uses a timeout of 400ms, and it's always worked pretty well. It's even adjustable (SPI_SETMENUSHOWDELAY) these days, but used to just be 80% of your double-click speed setting. So the faster your reflexes were - per that setting, at least - the shorter your menu flyout allowance. No cone needed. ( I stole this from a Raymond Chen hockey card that I keep tucked up in my 50 Mission Cap ).show more

Dave W Plummer
45,336 Aufrufe • vor 6 Monaten
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,992 Aufrufe • vor 5 Monaten
Selling candles on Etsy, Amazon, Shopify, or eBay? Then... you know the struggle — product videos are expensive and a huge time sink. I tried Pollo AI's "Photo to Video Ads" feature right inside the Mobile APP with just one candle photo. A few minutes later, that single image turned into a cinematic, cozy video ad that looked ready to go live. No studio. No crew. No editing skills. All from my phone. At this point, Pollo Agent honestly feels like my creative team. AI-generated product ads are getting scary good.show more

Aurelia Vance
33,828 Aufrufe • vor 3 Monaten
#JIHYO’s opening message at the Kiwoom Heroes vs. Lotte... Giants game: “Hello, I'm Jihyo from TWICE. It's an honor for us to be here today to throw the ceremonial first pitch and first hit on this meaningful day, brought together by the Kiwoom Heroes and Bumin Hospital I hope all the players stay healthy and give us a great game” JIHYO FIRST PITCH #JIHYOxKiwoomHeroesshow more

JIHYO GLOBAL UNION
30,973 Aufrufe • vor 2 Monaten
+113k No words can describe the month that I... had to be honest and all I can thank is God and the relations that I have made. To say that 3 months ago I went from -3k -> 5.6k -> 23k to now is just so amazing for me. Im happy with the turnout of every single day as well as I can't thank 67FnF anymore and all the boys that are in there 30-1 pretty damn happy and my 1 off day is the day after I ended a 4-year relationship so i'm happy to say I kept my head down and kept grinding. To 2026, i'm happy to be in the circle I am and good luck to everyone in this Q1 Cycle!show more

Cottage
59,425 Aufrufe • vor 8 Monaten
Astra (GPT-6) is here!!! I've had early access and... tested it like crazy with things like games, code, writing, browser control, presentations and general knowledge work. This is the best model I've ever used. Period. (Incredible demos below in this thread ⬇️) Here's my take on Astra: > It's insanely capable. This feels like a massive improvement, not just an incremental change. This is especially true with zero-shot prompts. > It's all about knowledge work. Slide creation, analysis, writing, and browser control. And oh my...it's so good at browser control. GPT-5.6 was already fantastic at doing things in the browser, Astra is another level and significantly faster. > We're closer than ever (arrived?) at prompt-to-playable game. And I don't just mean only playable, these are actually fun games. I bet if someone with a great eye for games used Astra, they could create a viral game within 1-2 weeks. > Astra is better at writing but not perfect. It removes much of the "AI Smell" we're all familiar with but some stink still survived. > It has a tendency to use the same design colors and look/feel as GPT-5.6 (forrest green anyone?) but it is more steerable in design than previous models. > It's highly steerable in general. A little nudge goes a long way. When I first started using Astra, almost every task I gave it would go for ~30 minutes. I wanted it to keep working. Adding more specifics to a prompt helped greatly with it's ability to work for a long time. > Astra's 3D understanding is unmatched. 3D asset creation was consistent and easy and its spacial awareness while building complex 3D worlds blew me away. I'm still getting familiar with Astra but this will now be my go-to model for any difficult work I have. Check out the demos below: 👇show more

Matthew Berman
1,892,422 Aufrufe • vor 9 Tagen
And while so many of you are treating the... fanchant like it‘s a way to give “credit” (which isn’t the purpose of a fanchant AT ALL) I hope the majority remember what ENHA themselves shared in the Blood Saga behinds. They talked about how intense the preparation for this tour was—from having to rearrange songs, relearn formations and rework choreography. These 6 poured their blood, sweat, and tears into making sure ENGENEs would have an amazing experience despite everything that has happened these past few months. Put fanwars, engagement, and number games aside. When you walk into that venue and see those 6 giving everything they have on stage, I hope it reminds you that the least we can do is respect and support the current lineup in front of us.show more

Trinᥫ᭡
13,821 Aufrufe • vor 2 Monaten
introducing a new, very fun, LLM benchmark- the Game-of-Life... Bench! the rules are simple: given an 8x8 grid following Conway's game of life rules, the goal is to create an initial pattern with at most 32 cells that can last the longest number of turns before dying/repeating. some results to highlight (with caveats detailed below): - gpt 5.1 lasts the longest with a 106 step run - claude models are really bad at this! they refuse to reason about this task and score < 25 points - deepseek r1 is the best open model with 102 steps. why? because i wanted to create a benchmark that has (i think) no practicality, but is still fun to look at, cheap, and still measures something interesting. i also am a big fan of the game of life. its absurdly simple rules leading to intractability is extremely cool to me. also, i saw a lot of work with LLMs trying to "predict" the next state in Conway's game of life, I think game-of-life bench is more fun because it's pretty open ended and only asks the LLM for the initial state. I also think this could be an RL env? but idk why you would ever train on this task haha i don't think this is a "serious" benchmark because it doesnt measure anything practical, but i still think it's a hard benchmark exactly because you can't predict what happens with your initial state many turns into the future; this is why i was initially expecting all LLMs to be bad at it, but turns out, some are clearly better than the others (the ordering may surprise you!) reminder: this is still a work-in-progress; (1) i am gpu-poor so could only do 10 runs for each model, even though total running cost is relatively low. maybe with some more credits i can run more seeds for each model. (2) i handpicked models which i think are at the frontier right now, plus some others that were on my mind. so, if you'd like to see a model on here, let me know. (3) i currently only do an 8x8 grid because i thought that by itself would be pretty hard for current LLMs, but of course we can increase grid sizes! (4) the coolest thing is, i dont think we can calculate the max possible number of states (yay undecidability!) you can go without repeating, so this is essentially a no-ceiling task, which is pretty cool! again, i did this mostly out of a desire to make LLMs do something fun. if this keeps me entertained for a few more days, i'd likely release a blog post on it. if it keeps me entertained for a week (and someone sponsors me), i'll put more work into it :P lastly, this is fully open sourced, so feel free to run this on your own!show more

Akshit
13,775 Aufrufe • vor 6 Monaten
I don’t think people realize how big a jump... Grok 4.6 is. Look at this comparison with Fable on the same prompt. They’re pretty close, but Grok took half the time at ONE TENTH the cost. And I prefer a lot of choices Grok made: - fills the hero image to the top instead of leaving an ugly white strip - flat sharp tiles instead of the usual rounded card slop - fewer pointless labels and elements, more restraint - lighter text, less dense and overwhelming Fable certainly has more depth if you’re trying to do crazy edge-of-distribution stuff. It’s a much bigger model. But for the majority of work, you can get Fable quality for a 90% discount. Just insane how much of a step up it is from Grok 4.5.show more

Anshu
53,085 Aufrufe • vor 29 Tagen
Today I will send my last appeal, I'm just... done with YouTube. I'm completely burned out. After fighting for months to save my 13-year-old Panzer Picture channel from that ridiculous false "Child Sexual Abuse" flag, I pretty much lost all hope it's ever coming back. Bigger channels stay silent, when so many channel are being removed for CP and Sexual abuse, I can tell you, your time will come. The stress has been brutal on my health (Crohn’s hasn’t helped). If you're still here and believe in preserving WW2 history, a like or RT would mean the world right now. Thank you for not forgetting ❤️ #RestorePanzerPicture #WW2History #YouTubeCensorshipshow more

Panzerpicture
117,758 Aufrufe • vor 2 Monaten
JADE gets emotional reflecting on her North American tour... in a new TikTok: “I’m having so much fun on this tour. I just feel so grateful that I’m getting to do this at this point in my career. The fact that I get to tour [North America] after being in the industry for 15 years, and only now just getting to do my own headline tour, is incredible. It’s been a long time coming. What’s really beautiful about these shows is that when I look into the audience, I recognise so many of the fans here from back in the day, who’ve literally waited for years – like me! – for this moment. It just makes me feel so loved and supported to know people have believed in me enough to stick around for years waiting for this to happen. Anyone that’s bought a ticket, dressed up, made their own costumes… It’s just such a lovely, beautiful thing. I hope they can see when I’m on stage just how much that means to me, and how much I love performing and putting on the best show that I possibly can every single night. I will never ever take it for granted. I’m just so chuffed that I get to do this for a living… be a silly pop girlie, write and create music that brings people – and myself – a lot of joy… Thank you for believing in me. I literally get on the bunk on my bus most nights from the tour and just lie there like, ‘Oh my god! As if this is my life!’ It means a lot. I hope I get to do this for the rest of my life… Full of gratitude and lots of all the lovely emotions. Thank you so much.”show more

JADE tea room ☕️
47,114 Aufrufe • vor 6 Monaten