multi-player, multi-view, llm generated shared 3d objects *this palm... tree cost about 1.3 cents using gpt-4oshow more

Yohei
25,130 views • 1 year ago
I just validated my startup idea using a multi-agent... AI debate. → GPT-4o (Founder) → Gemini (VC) → Claude (Customer) The result? Shockingly real...and honestly everyone should try this: 👇show more

Shruti
228,029 views • 1 year ago
This is hand and finger motion capture powered by... AI. 3D motion data is generated from 2D video. No suits or gloves required. This example was captured using Move Multi-Cam with 8 GoPro cameras shooting at 4K and 120 FPS.show more

Move AI
52,597 views • 2 years ago
This is 3D animation powered by AI. No suits... or markers needed. This 2D video was captured with GoPro cameras and converted to 3D motion data using Move Multi-Cam. #MadeWithMoveshow more

Move AI
182,792 views • 2 years ago
THE BEST visual explainer of how information propagates through... a transformer. If you want to have more than intuition about how the Transformer architecture is ruling the LLM world - open-source project explains everything about LLM Transformer Models! - A great resource for anyone looking to gain a deeper understanding of how Transformer-based AI models like GPT work, including: - Self-attention mechanisms - Encoder-decoder architecture - Positional encoding - Multi-head attentionshow more

Rohan Paul
106,897 views • 1 year ago
Adaptive and Temporally Consistent Gaussian Surfels for Multi-view Dynamic... Reconstruction Contributions: • A method for efficiently reconstructing dynamic surfaces from multi-view videos using Gaussian surfels. • A unified and gradient-aware densification strategy for optimizing dynamic 3D Gaussians with fine details. • A temporal consistency approach that ensures stable and coherent surface reconstructions across frames by enforcing consistency on curvature maps. • Extensive experiments that demonstrate our method’s advantages including fast training, high-fidelity novel view synthesis, and accurate surface geometry.show more

MrNeRF
31,821 views • 1 year ago
What if anatomy explorers felt alive? This 3D dog,... including the skeleton, organs and rig, was generated with ai and reacts to the cursor in real time with head tracking and tail movement. - Used GPT Images 2 for consistency - Omma AI for 3D generation and code using Three.js Lmk if you’d like to try it!show more

Gábor Pribék
204,523 views • 2 months ago
By far the best way to generate images at... different angles is this 3D Camera Control node. Instead of describing angles with words, you just click the view you want. Simple UI, and the Qwen Image Edit 2511 multi-angle LoRA keeps things consistent across generations. Workflow below 👇show more

rob - comfyui
59,189 views • 6 months ago
#Keep4o 🚨THE GPT-4o FILE🚨 Researchers at Microsoft Research published... a paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4.” Their conclusion: “An early (yet still incomplete) version of an artificial general intelligence (AGI) system.” 📎 Paper: OpenAI’s Charter defines AGI as: “Highly autonomous systems that outperform humans at most economically valuable work.” 📎 Source: OpenAI’s own System Card for GPT-4o shows that the model improved performance on 21 out of 22 medical evaluations compared to GPT-4T. On the MedQA USMLE (the U.S. medical licensing exam), accuracy jumped from 78.2% to 89.4% , surpassing specialized medical AI models like Med-Gemini and Med-PaLM 2. 📎 Source: Under OpenAI’s agreement with Microsoft, AGI is explicitly excluded from Microsoft’s license. And who decides if AGI has been reached? OpenAI’s Board. WHAT THEY DID WITH IT AFTER THEY TOOK IT FROM PEOPLE A. Military deployment. On February 28, OpenAI signed a deal to deploy models in classified military environments. 📎 Source: B. State Department. A State Department memo confirmed: “For now, StateChat will use GPT-4.1 from OpenAI.” This is a direct descendant of the GPT-4 family the same family Microsoft’s researchers called early AGI. 📎 Source: C.Altman’s personal biotech investment. Altman personally invested $180 million in Retro Biosciences,a longevity startup.OpenAI then built GPT-4b micro, based on GPT-4o.The model made proteins 50 times more effective. 📎 Source: WHAT INDEPENDENT BENCHMARKS SHOW Overall SM-Bench score: GPT-4o (extended): 66.6% GPT-5.3 Chat: 63.4% GPT-5.1: 58.9% GPT-5.4: 51.4% GPT-5.2: 47.8% Creative Writing: GPT-4o: 97.31% Pass 98, Fail 2 GPT-5.4: 36.77% Pass 40, Fail 60 Reasoning / Overfit: GPT-4o: 83.06% GPT-5.4: 39.25% The model they removed is still the best they ever made at the things humans actually use AI for. 📎 Source: Musk asks the court to make a judicial determination on whether GPT-4 constitutes AGI. If a jury finds that GPT-4 is AGI, then GPT-4o,which was more advanced,is also AGI and under OpenAI’s own founding documents, it was never supposed to be locked behind a subscription,licensed exclusively to Microsoft, given to the military, or taken away from the public. 📎 Source: The most powerful version of GPT-4o was never given an official dated snapshot. It was only available through the chatgpt-4o-latest endpoint that OpenAI itself described as intended for “research use only.” It was never officially archived. That is not an oversight. That is a pattern. 📎 Source: 📎 Source: WE DEMAND A.Frozen model snapshots under independent custody. Specifically: gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, the March 2025 version (chatgpt-4o-latest), gpt-4-0613 (the original GPT-4 evaluated in the Sparks of AGI paper), and gpt-4.1-2025-04-14 (currently running in the State Department). B.Cryptographic hash verification (SHA-256) for each snapshot. Every model has weights. Those weights can be hashed. If OpenAI provides a snapshot today, the hash proves whether the weights were modified later. This is the only way to verify that models were not downgraded before testing. C.Independent AGI benchmarking. Using the AGI definition from OpenAI’s own Charter applied to ALL frozen snapshots listed above. D.Explanation for the missing March 2025 snapshot. OpenAI was founded on one promise: build AGI for the benefit of humanity. -They took it from us. -They gave it to the military. -They gave a custom version to the CEO’s biotech investment. -They put it in government classified networks. -They refuse to call it AGI because the moment they do, they lose billions.show more

🩵BlueBeba🩵
17,835 views • 4 months ago
GSTAR: Gaussian Surface Tracking and Reconstruction Contributions: • A... new framework for tracking and reconstructing dynamic scenes, combining 3D Gaussians and meshes to effectively manage changes in topology. • A method for Gaussian unbinding and surface re-meshing, allowing for the generation of new surfaces as topologies evolve. • A method for handling large or fast deformations of surfaces between frames using scene flow warping. Abstract (excerpt): However, tracking dynamic surfaces with 3D Gaussians remains challenging due to complex topology changes, such as surfaces appearing, disappearing, or splitting. To address these challenges, we propose GSTAR, a novel method that achieves photo-realistic rendering, accurate surface reconstruction, and reliable 3D tracking for general dynamic scenes with changing topology. Given multi-view captures as input, GSTAR binds Gaussians to mesh faces to represent dynamic objects. For surfaces with consistent topology, GSTAR maintains the mesh topology and tracks the meshes using Gaussians.show more

MrNeRF
22,698 views • 1 year ago
POV: One tiny act of kindness changed EVERYTHING Gugugaga... only had one dumpling… but she still shared it with a hungry little bunny in the rain What happened next melted my heart Created this cozy Pixar-style 3D animated short using GPT Image 2 + Seedance Huge shoutout to Renoise canvas for helping bring this wholesome world to life Would you share your last dumpling? Prompt is in the video DM for full Prompt #RenoiseCanvasshow more

Shami
213,606 views • 1 month ago
Introducing Kaleido💮 from AI at Meta — a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:show more

Shikun Liu
22,332 views • 9 months ago
Seedance 2.0 - Cinematic Summoning VFX Prompt This prompt... generates a multi-cut VFX showcase with a clear progression. It’s designed for realistic 3D with grounded physics, natural character motion and readable, production-style VFX. You don’t need to define the summon, but you can specify what is being summoned if you want control. If you leave it open, the model decides based on the character sheet. You can find gpt image 2 character sheet prompt in the quoted post. Also, this isn’t limited to summoning. You can reuse the same VFX structure for anything, just swap the effect logic. Just change the parts about summoning. Prompt in the replies 👇show more

Kōda
28,400 views • 2 months ago
Republicans like Nancy Mace insist that Liam should be... using the female restroom. Really? A few questions: - Why? - How does this make sense to anyone beyond a kindergarten-level understanding of gender? - Have you actually thought this through, or is this just about scoring cheap political points? The point of posts like this is simple: Complex, multi-faceted issues can’t be reduced to oversimplified, black-and-white answers. But hey, I guess nuance isn’t for everyone.show more

Brian Krassenstein
621,330 views • 1 year ago
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video... - PackUV — A new volumetric video representation that packs 3D Gaussian attributes into a sequence of UV atlases for efficient streaming and storage, making it readily compatible with existing video coding infrastructure. - PackUV-GS — An efficient method to fit PackUV directly from multiview videos using optical-flow-based keyframing and Gaussian labeling to handle large motions, disocclusions, and temporal consistency. - PackUV-2B — The largest multi-view 4D dataset with 2B frames, large motions, and disocclusions. It provides 360° coverage from 50+ synchronized cameras.show more

MrNeRF
17,035 views • 4 months ago
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,801 views • 1 year ago
CHINA JUST LEAKED THE FUTURE OF WEB APPS. Alibaba... open-sourced PageAgent and 99% of SaaS founders are sleeping on this. It's a JavaScript AI agent that lives INSIDE your webpage. Users control your entire interface with natural language. ↳ No browser extensions needed, screenshots or multi-modal LLMs, headless browser setup, and also no backend rewrite required Just drop it in your HTML with ONE line of code. What took 20 clicks now takes one sentence. "Click login, fill in my credentials, submit the form" Done. This is not a demo, it is production-ready. ↳ Turn any SaaS into an AI copilot in minutes ↳ Smart form filling for ERP, CRM, admin systems ↳ Voice commands and accessibility built in ↳ Multi-page agent tasks via Chrome extension ↳ MCP server support for external control ↳ Bring your own LLM (Qwen, GPT, Claude, anything) Every founder building AI features just got a shortcut. Every developer manually building copilots just got replaced. The integration looks like this: That's it. Your app now has an AI agent.show more

Kanika
329,499 views • 20 days ago
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
64,883 views • 26 days ago
Back when we were developing GEN3C, we often imagined... a Holodeck-like future: a simulator where multiple agents can enter the same generated world, act independently, and learn to collaborate. Gamma-World makes this feel more concrete. It is a generative multi-agent world model that takes synchronized observations and actions, then rolls out what each agent will see next in the same evolving world — action-responsive at 24 FPS. For me, the key challenge is going beyond two players. As more agents enter, identity cannot be tied to fixed slots, interaction cannot rely on dense pairwise attention, and independent actions still need to resolve into one shared state. Two ideas make this work: 1⃣ Simplex RoPE Distinct agent identities without slot bias — unique, but permutation-equivalent. 2⃣ Sparse Hub Attention Agents communicate through learnable hubs instead of dense all-to-all attention: agent → hub → agent This keeps cross-agent communication scalable. The exciting part: training on two-player data can generalize to four-player rollouts without additional training, and the same formulation extends to real-world bimanual robot coordination. A step toward populated world models: many agents, one shared world. Congrats to the team on Gamma-World! Project:show more

Xuanchi Ren
304,145 views • 1 month ago
This is the kind of project that quietly changes... an industry. Instead of using traditional slicers, this creator built an app that converts hand-drawn curves directly into printable G-code for spiral vase mode prints. App → G-code → Print. - No complicated workflow. - No heavy setup. - Just turning ideas into physical objects almost instantly. The interesting part isn’t just the print quality. It’s how tools are becoming so simple that more people can create complex things without needing advanced technical knowledge. Would tools like this make 3D printing finally mainstream? 🎥 Media: ( Instagram ) ⚠️ This content is shared for informational purposes only. CTO Robotics Media is a media platform and does not own or develop the technology shown. Credit belongs to the original creators.show more

CTO ROBOTICS Media
132,137 views • 1 month ago
What if you could turn a single 360° photo... into a production-ready Isaac Sim environment in minutes? That's exactly what we did here. Using World Labs' Marble and an Insta360 X5 capture (rotating on top), we generated a complete navigable 3D environment and populated it with Lightwheel Sim Ready assets (bottom view). The result? A fully interactive scene in Isaac Sim, ready for sim2real testing,. Navigation, manipulation, or any robotics task you need to validate. What used to take weeks of manual 3D modeling and asset placement now takes minutes. Capture once in the real world, simulate everywhere in your training pipeline. This is the future of robotics development with world models. NVIDIA Robotics NVIDIA Omniverse #Sim2Real #Robotics #Simulationshow more

Jonathan Stephens
46,643 views • 6 months ago