Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Super excited for #ECCV'26 in Malmö! Our group will present a series of cutting-edge works, including two orals and a spotlight, around generative AI that cover 3D scene generation, photo-realistic avatars, agentic modeling, and much more! I'll also give several workshop talks about our recent research, including a few...

14,628 görüntüleme • 6 gün önce •via X (Twitter)

2 Yorum

Animesh Karnewar profil fotoğrafı
Animesh Karnewar6 gün önce

@LukasHollein Congrats @ErkocZiya and @LukasHollein! Still going strong 💪 🙂. Solid line-up. Congrats to all 🎉!

Noiz Ai profil fotoğrafı
Noiz Ai3 gün önce

The 3D scene generation and lighting transition combination feels especially relevant to Helix. It is one thing to make a plausible view, another to keep illumination and geometry consistent as the observer moves.

Benzer Videolar

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,810 görüntüleme • 1 yıl önce

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,903 görüntüleme • 9 ay önce

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360°, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360° head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data 🏠 🌍 🎥 Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,334 görüntüleme • 8 ay önce

🚨 SIGGRAPH Asia 2025 Paper Alert 🚨 ➡️Paper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes 🌟Few pointers from the paper 🎯Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. 🎯 To this end, authors of this paper proposed “WorldExplorer”, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. 🎯They initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. 🎯Then, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. 🎯Concretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. 🎯Their novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. 🎯Finally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. 🎯Compared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. 🎯They believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. 🏢Organization: TU München 🧙Paper Authors: Manuel-Andreas Schneider, Lukas Höllein , Matthias Niessner 📝 Read the Full Paper here: 🗂️ Project Page: 🧑‍💻 Code: 🎥 Be sure to watch the attached Technical Summary Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️QT and teach your network something new Follow me 👣, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 görüntüleme • 11 ay önce

We’re thrilled to share that our MERFISH+ preprint is now live on bioRxiv!👉 In this work, the Bintu and Zhu labs (UCSD) developed MERFISH+, a next-generation spatial genomics platform that combines genome-wide RNA and epigenetic imaging over a large field of view. By introducing acrydite-modified probes covalently anchored to hydrogels, MERFISH+ achieves remarkable imaging stability and enables >1,800-gene, multi-modal, and multi-month experiments. With this platform, they, together with the Chi lab at UCSD, profiled a whole developing human heart at 12 post-conception week with merely two slides, resulting in a total of 53 slides, 3.1 million single cells and more than 30 cell types. Building upon our previous 3D reconstruction and modeling framework, Spateo ( we reconstruct the 3D human heart that nicely captures the anatomical structure of the heart, including the intricate vasculature network. Sophisticated analyses provide a holistic view of an entire organ and enable systematic characterization of 3D cellular neighborhoods and transcriptional gradients of substructures such as the descending arteries. Furthermore, using a generative integration framework for spatial multimodal data (Spateo-VI), we harmonized these MERFISH+ transcriptomic and chromatin data to reconstruct a 3D spatially-resolved multi-omics atlas of the developing human heart, shared at and MERFISH+ thus sets a new standard for large-format, multi-omic spatial profiling, enabling holistic, 3D characterization of organs at subcellular resolution. Huge congratulations to first authors Colin Kern, qingquan Zhang, Yifan Lu , and Jacqueline Eschbach, and to all collaborators from the Bintu, Zhu, Chi, and Qiu labs for this amazing team effort. Thanks for your diligence, creativity, and hard work on this project. We’re grateful for support from Arc Institute and our generous donors. Our lab is expanding—if you’re excited about building the next generation of single-cell and spatial genomics techniques and predictive single cell and spatial foundation models, we’re hiring! If you are interested, please reach out to me via direct message or email at [email protected]. We are excited for any potential collaborations along this line of research in Stanford, UCSF and Berkeley and other labs as well.

evo-devo

42,326 görüntüleme • 10 ay önce

3D scanning and rendering is moving so fast - got my splats up and running and I'm mind blown getting ~100fps for this complex 3D scene ⬇️ 🤯 1. WAY faster than NeRF: For comparison, NeRFs would takes around 10 seconds per frame (!) Instead I'm zipping around with FPV controls without breaking a sweat - though I do crash a few times towards the end of the video lol 2. Old Meets New: Gaussian Splatting is cool in that it fuses classical graphics and deep learning techniques. Like NeRFs, this is still a radiance field - just without the slower (ne)ural rendering part. 3. Explicit Representation: Instead you represent a 3D scene as a collection of ellipsoidal "splats" called gaussians. Each gaussian has a position, size, and color. Rendering in real-time is done by projecting into the image plane and alpha blending. 4. Photorealistic Effects: Gaussian splatting use spherical harmonics to represent the view-dependent effects and lighting - allowing surfaces to change color when viewed from different angles, enabling greater photorealism. It doesn't use a neural network, but the training loop is similar to deep learning. 5. Enables Direct Editing: But it's not just speed - with Gaussian Splatting you also get 3D editing support! So you can select, move, and delete stuff, even relight stuff. This type of editing has been more tedious to do with NeRFs and their implicit black box representations. 📲 More tests cooking! Much more to unpack here including simpler explanations. If you enjoyed this post, you might enjoy my feed: Bilawal Sidhu

Bilawal Sidhu

337,090 görüntüleme • 3 yıl önce

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,821,114 görüntüleme • 1 yıl önce

Fast Company just published a great piece on World Labs , Fei-Fei Li , Marble, and the idea that spatial intelligence / world models may be one of the next big shifts in AI. I was happy to be quoted in the article, but I also wanted to share more context about my own experience with World Labs and Marble, and why this direction is especially interesting to me. My starting point: volumetric capture — For the past few years I’ve been exploring and using volumetric capture and reconstruction (photogrammetry, NeRFs, 3D Gaussian Splats) mostly capturing locations around Montreal. Alleys, museums, urban interiors. I love every step of it: the capture itself, the pipeline, and what can be done with the output. Turning real spaces into real-time explorable systems. I do this personally, sharing explorations here, and professionally as chief technologist, and co-founder of Dpt. Physical reality + generative manipulation — In my work I’m especially drawn to mixing physical reality with generative and digital manipulation: using physical interfaces (light, clay, ink, ... ) to drive generative AI pipelines, building mixed reality prototypes that reshape your surroundings, or starting from real captured spaces and transforming them using tools like Marble. Like many people, I saw the World Labs announcement on Twitter in September 2024, and Marble when it surfaced in early December. But by then, I already had a sense something was coming. The first conversation — As someone deep into volumetric capture and radiance fields, I obviously knew about Ben Mildenhall and his pioneering work on NeRF. To my surprise, Ben reached out to me in late June 2024. He’d been following some of my experiments and wanted to chat about my process and workflows and how I was using this “stuff” creatively. At that point he didn’t share what he was building, but we had a genuinely great conversation about radiance fields, AI, and my work. He was curious about the creative perspective, not just the technical one. When the World Labs announcement dropped a few months later, it all made sense. I understood what Ben had been working on, and why the creative angle mattered to them. Then in August 2025, he invited me to try the Marble beta, and I’ve been experimenting with it since. Experimenting with Marble — The first thing I used Marble for was materializing scene and world concepts during ideation at the studio, and seeing if and how it could fit into our production pipeline. In parallel, I dove into a series of experiments focused on world manipulation: starting from real captured spaces and transforming them using Marble. I’d already been exploring that idea using img2img diffusion with ControlNet on NeRF renders, real-time video streams, and even mixed reality using headset camera feeds. But Marble brings something different. It generates persistent, spatially cohesive 3D worlds that can be rendered in real time across a wide range of devices. That’s a real shift. Experiment 01: Parallel Realities — The first experiment, Parallel Realities, starts from a volumetric capture of a real location, reconstructed as 3D Gaussian Splats. Using Marble, I generate an alternate version of that same space, something informed by the original architecture: abandoned, nature-reclaimed, alternate era. Then, using Spark (World Labs’ 3D Gaussian Splatting renderer for THREE.js) I make both realities coexist in the same spatial coordinate system. From there, I use a portal UX mechanic to let the user step between the real reconstruction and the Marble-generated version. Experiment 02: Hidden Depth The second experiment, Hidden Depth, does not transform a space as much as expand it. A captured location has a visual boundary (a mural, a doorway, a dark corridor) and Marble generates what exists beyond it. For example: a Montreal alley has a painted mural; step through it and you’re inside a world informed by what is actually depicted there. World Labs showcased part of this work here: And in their Spark 2.0 post: The project page is here: Why this matters to me — Being able to start from a real 3D Gaussian Splat scene and manipulate it with Marble opens up a lot of ideas. The 3DGS pipeline is becoming an increasingly compelling foundation for exploration, experimentation, and storytelling. What matters most to me right now is more control. The more I can steer the generated scene or world, the more useful the tool becomes. I want more features like the already existing multiple input images and Chisel, the blockout-based approach. I would like better local control, the ability to expand a generated world more and more while preserving coherence, and the ability to directly import 3D Gaussian Splat scenes to be used as a starting point. I want more ways to shape the result, not just a “prompt and hope” approach. — It is exciting to see this field moving from research and demos toward actual creative workflows.

Hugues Bruyère

69,960 görüntüleme • 2 ay önce

For generative AI to become an interesting art tool, we need much more control over the output. The slot-machine-like nature of pure text-to-image leaves too much to chance. Using the "Real-time Latent Consistency Model" that I'm using in the example here, is the first time I truly got a glimpse of a future where we'll be able to use our artistic skills and sensibility, to get control over AI image gen. Systems like these will never be able to match the quality or originality of a skilled artist, it won't surprise us in the same way an artist can. Things are a mess in terms of the training data these models are based on, and the questions about copyright concerns and about a time when everything will look the same are very valid. At some point capabilities like these will be embedded in photoshop, and anyone will be able to generate a pretty picture. But to create interesting designs, to tell original stories and to surprise us, we need creatives and artists with something on their mind. We'll be able to create immersive worlds, by making brush-strokes and sculpt marks, without needing to worry about all the dials, plugins, wires of our 3d and 2d tools today. I love to sculpt, I love to draw, and I love to explore new mediums and new ways to create. The Gen AI tools we have today are far from perfect, and things need to be steered in a better direction. For that we need artists to help point the way. Gen AI isn't going away.. it's too powerful and has the potential to allow us to tell stories like never before. Like all other big technological shifts, tech like this will come at a cost, but it will also open up new opportunities and empower a new generation of storytellers. I might be naive, but I believe that human ingenuity and creativity will persevere in this new world ♥️ #art #ai

Martin Nebelong

1,660,426 görüntüleme • 2 yıl önce

If an AI can control 1,000 robots to perform 1 million skills in 1 billion different simulations, then it may "just work" in our real world, which is simply another point in the vast space of possible realities. This is the fundamental principle behind why simulation works so effectively for robotics. Real-world teleoperation data scales linearly with human time (< 24 hrs/robot/day). Sim data scales exponentially with compute. There are 3 big trends for simulators in the near future: 1. Massive parallelization on large clusters. Physics equations are "just" matrix math at their core. I hear GPUs are good at matrix math 🔥. One can run 100K copies of simulation on a single GPU. To put this number in perspective: 1 hour of wallclock compute time gives a robot 10 years (!!) of training experience. That's how Neo was able to learn martial arts in a blink of an eye in the Matrix Dojo. 2. Generative graphics pipeline. Traditionally, simulators require a huge amount of manual effort from artists: 3D assets, textures, scene layouts, etc. But every component in the workflow can be automated: text-to-image, text-to-3D mesh, and LLMs that write Universal Scene Description (USD) files as a coding exercise. RoboCasa is one example of a prior work. 3. End2end neural net that acts as simulator itself. This is still bluesky research and quite far from replacing a graphics pipeline, but we are seeing some exciting signs-of-life based on video gen models: Sora, Veo2, CogVideoX, Hunyuan (text-to-video); and action-driven world models: GameNGen, Oasis, Genie-2, etc. Genesis does great on (1) for certain tasks, shows good promises on (2), and could become a data generation tool for reaching (3). Its sim2real capabilities for locomotion are good, but there's still a long way to go for contact-rich, dexterous manipulation. It shows a bold vision and is on the right path to providing a virtual cradle for embodied AI. It is open-source and puts a streamlined user journey at the front and center. I had the privilege to know Zhou Xian and play a small part in his project since a year ago. Xian has been crunching code non-stop on Genesis with a very small group of core devs. He often replied to my messages at 3 am. Zhenjia Xu from our GEAR team helped with sim2real experiments in his spare time. Genesis is truly a grassroot effort with an intense focus on quality engineering. Nothing gives me more joy than seeing the simulation ecosystem bloom. Robotics should be a moonshot initiative owned by all of humanity. Congratulations.

Jim Fan

157,343 görüntüleme • 1 yıl önce

🎉 new skill unlocked: 20s uninterrupted, unstitched, single render from our new ai video engine: Nami. This is my birb (#7531) from the Moonbirds collection, idling in the library. patent: "Intra-Latent Semantic Injection via Cross-Spatial Encoding and Decoding during Multi-Pass Inference for Generative AI Video Creation" At Scrypted we've been quietly working on an agentic generative AI stack for two years: • integrating and testing w/ partners across the games & entertainment sectors • stealthily building a community of early believers through AVB • showcasing some of what we're doing with amazing projects like H011yw00d Agent. -- about Nami -- Nami is an agentic orchestration layer for AI video models: it unlocks their inner superpowers without making them rely on custom LoRAs or fine-tunings. Instead of throwing raw training power and tens of millions of dollars at training yet another ai video model: we figured out new ways to use what we have. Nami harnesses a multi-agent system to perform the work needed in taking a simple prompt or image and turning it into something bigger - much bigger. The agentic steps are allowed to manipulate latent space, digging into tensors, yet doing so in semantically aware chunks - meaning that Nami inherently supports video generation of arbitrary length, though it's bound to O(n) rendering time. (We do have some cool sharding tech that allows us to cut the generative time in half for a reference pose idle-animation like this demo). It's also fairly agnostic, picking and choosing the right tools for the job, and plays really well with emerging tech like FLUX Kontext, FramePack, or <- without being limited by any of them. -- use cases -- Even just a year or two ago the 20 second render below would cost a company, paying an agency, around $10k start-to-finish. This one cost me $6.25 on our dev hardware in an unoptimized environment. There's something mind-blowing about the state-of-the-art when we reduce costs to 0.0625% - less than 1% - of what we used to pay. It's also empowering. For creators. Game developers. Content influencers: you name it. -- superpowers -- 1. it does the things you ask for, in the order you asked for it 2. consistency is king 3. single-shot text or image-to-video 4. future videos can reference previous ones to seamlessly maintain style 5. semantic stitching: can't wait to showcase this -- gtm -- We think Generative AI Video, like image generation, like text, like games, should be a publicly accessible common good. We believe democratizing access to Nami in web3, via x402 payments proposed by Drew Coffman, or in World's mini-apps, is a bold step forward for digital freedom. Permissionless, decentralized, generative ai video. Naturally, we'll also soon release a web platform for using Nami in a traditionally SaaSy way: bring your own images, videos, or prompts and we'll take care of the rest. In the mid-term, Scrypted is building a stack of agentic skills (we call it AVB) and making them available to projects like H011yw00d Agent on Virtuals Protocol and other platforms. -- long-term vision -- Scrypted's mission is to decentralize the things that can't be decentralized. We participated in a16z crypto's CSX (London 2024) during our pre-seed specifically to research a new consensus protocol for hard things like AI video and AI agents: where there's no "one right answer". When Zero-Knowledge Proofs (ZKP) can't secure it, and Trusted Execution Environments (TEEs) are too small, we've got you covered with our upcoming Inori Network. -- how you can help -- 1. Are you a GPU farm? We're gonna need more flops. 2. Do you represent an L1 or L2? We want to build bridges. 3. Do you represent a Wallet or App creator? Let's get an endpoint exposed. 4. Are you an investor? Let's chat. 5. Like, repost, share! -- team background -- We come from a background of AI in the Video Game industry with each founder having over 20 years of experience at companies like Electronic Arts & Square Enix. -- contact -- DMs are open, reach out if you want to be an early tester for your site, game, collection, or project! -- try it out -- Go anywhere on X and tag H011yw00d Agent with a prompt and she'll give you a free 2 second render. Have fun making cinematic shorts or meme videos! -- thanks -- AWS Startups has been an incredible help scaling our prototypes. Also, shout out to all loyal beans 🫘 in the Autonomous Virtuals Beings (AVB) community. Nami has a very important role in the upcoming XP agent platform, can't wait to show you all. AVbeings

Tim Cotten

12,708 görüntüleme • 1 yıl önce

AGI? One day, but not yet. The only AI that works well right now is the one behind the screen [12-17]. But passing the Turing Test [9] behind a screen is easy compared to Real AI for real robots in the real world. No current AI-driven robot could be certified as a plumber [13-17]. Hence, the Turing Test isn't a good measure of intelligence (and neither is IQ). And AGI without mastery of the physical world is no AGI. That’s why I created the TUM CogBotLab for learning robots in 2004 [5], co-founded a company for AI in the physical world in 2014 [6], and had teams at TUM, IDSIA, and now KAUST work towards baby robots [4,10-11,18]. Such soft robots don't just slavishly imitate humans and they don't work by just downloading the web like LLMs/VLMs. No. Instead, they exploit the principles of Artificial Curiosity to improve their neural World Models (two terms I used back in 1990 [1-4]). These robots work with lots of sensors, but only weak actuators, such that they cannot easily harm themselves [18] when they collect useful data by devising and running their own self-invented experiments. Remarkably, since the 1970s, many have made fun of my old goal to build a self-improving AGI smarter than myself and then retire. Recently, however, many have finally started to take this seriously, and now some of them are suddenly TOO optimistic. These people are often blissfully unaware of the remaining challenges we have to solve to achieve Real AI. My 2024 TED talk [15] summarises some of that. REFERENCES (easy to find on the web): [1] J. Schmidhuber. Making the world differentiable: On using fully recurrent self-supervised neural networks (NNs) for dynamic reinforcement learning and planning in non-stationary environments. TR FKI-126-90, TUM, Feb 1990, revised Nov 1990. This paper also introduced artificial curiosity and intrinsic motivation through generative adversarial networks where a generator NN is fighting a predictor NN in a minimax game. [2] J. S. A possibility for implementing curiosity and boredom in model-building neural controllers. In J. A. Meyer and S. W. Wilson, editors, Proc. of the International Conference on Simulation of Adaptive Behavior: From Animals to Animats, pages 222-227. MIT Press/Bradford Books, 1991. Based on [1]. [3] J.S. AI Blog (2020). 1990: Planning & Reinforcement Learning with Recurrent World Models and Artificial Curiosity. Summarising aspects of [1][2] and lots of later papers including [7][8]. [4] J.S. AI Blog (2021): Artificial Curiosity & Creativity Since 1990. Summarising aspects of [1][2] and lots of later papers including [7][8]. [5] J.S. TU Munich CogBotLab for learning robots (2004-2009) [6] NNAISENSE, founded in 2014, for AI in the physical world [7] J.S. (2015). On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning (RL) Controllers and Recurrent Neural World Models. arXiv 1210.0118. Sec. 5.3 describes an RL prompt engineer which learns to query its model for abstract reasoning and planning and decision making. Today this is called "chain of thought." [8] J.S. (2018). One Big Net For Everything. arXiv 1802.08864. See also patent US11853886B2 and my DeepSeek tweet: DeepSeek uses elements of the 2015 reinforcement learning prompt engineer [7] and its 2018 refinement [8] which collapses the RL machine and world model of [7] into a single net. This uses my neural net distillation procedure of 1991: a distilled chain of thought system. [9] J.S. Turing Oversold. It's not Turing's fault, though. AI Blog (2021, was #1 on Hacker News) [10] J.S. Intelligente Roboter werden vom Leben fasziniert sein. (Intelligent robots will be fascinated by life.) F.A.Z., 2015 [11] J.S. at Falling Walls: The Past, Present and Future of Artificial Intelligence. Scientific American, Observations, 2017. [12] J.S. KI ist eine Riesenchance für Deutschland. (AI is a huge chance for Germany.) F.A.Z., 2018 [13] H. Jones. J.S. Says His Life's Work Won't Lead To Dystopia. Forbes Magazine, 2023. [14] Interview with J.S. Jazzyear, Shanghai, 2024. [15] J.S. TED talk at TED AI Vienna (2024): Why 2042 will be a big year for AI. See the attached video clip. [16] J.S. Baut den KI-gesteuerten Allzweckroboter! (Build the AI-controlled all-purpose robot!) F.A.Z., 2024 [17] J.S. 1995-2025: The Decline of Germany & Japan vs US & China. Can All-Purpose Robots Fuel a Comeback? AI Blog, Jan 2025, based on [16]. [18] M. Alhakami, D. R. Ashley, J. Dunham, Y. Dai, F. Faccio, E. Feron, J. Schmidhuber. Towards an Extremely Robust Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms. Preprint arxiv 2404.08093, 2024.

Jürgen Schmidhuber

72,331 görüntüleme • 1 yıl önce

The disgustingly simple framework behind my best performing thumbnails (1.7M, 1M views) This is the exact method I used for thumbnails on my best performing videos I call it the "analyze-recreate" method: 1. Search YouTube for your video title 2. Find the top performing video 3. Copy their thumbnail image 4. Send it to Claude/GPT with this prompt (steal this): ---- You are a YouTube thumbnail visual analysis expert. Analyze only what is visible in the provided thumbnail. Do not infer identities, brands, context, or any off-frame elements. Output must be plaintext and under 1800 characters. Keep it concise and directly usable for recreating a similar image. Important: Do NOT reference, describe, paraphrase, or characterize any text, typography, logos, watermarks, UI, arrows, icons, or other graphic design overlays. Focus only on the underlying photographic/illustrative scene and its composition. Cover exactly these sections: 1. Overall Style: mood, tone, realism level, and artistic approach (photo, 3D, illustration, composite). 2. Color Palette: 2–3 dominant colors and where they appear (brief; no hex). 3. Composition: camera angle/shot type (close/medium/wide), depth (foreground/mid/background), focal point, balance, and approximate % of frame occupied by each main scene element. 4. Main Subjects: describe visible characters/objects only, pose/expression, size relative to frame, and key environment/background elements. 5. Lighting & Post: lighting type and direction, shadow softness, contrast, depth of field, and noticeable post-processing (very short). 6. AI Prompt Suggestion: a compact generation prompt describing the scene + composition + mood + rendering style; do not list colors; do not include any text/logo/overlay instructions. Write the analysis now: ---- Paste the output into an AI image generator (Nano Banana, etc.) Add your text/graphics manually - I use autotube's built-in thumbnail designer which makes this easy The key is: you're recreating the image from scratch, not stealing. You're analyzing what works and building your own version - while avoiding copyright issues It might take 5-10 generations to nail it. Different AI models handle different styles better, so experiment YouTube is a game of copy/paste, so stick to what has been proven to work by your competitors instead of reinventing the wheel Shout out to the channels BEFORE CITIES, Woke up History, Sleepless Homo, Ancestor Vault among others, who were a big inspiration for this channel The above method was the one I used most of the time for Sleepy Time History, but recently I switched to something crazier: fully AI driven thumbnail generation. AI creates 5 thumbnails based on a "thumbnail design language" based on your channel's 15 best thumbnails So far so good, but this is more experimental, and mostly to save time while I focus on new channels I built autotube to automate the entire YouTube workflow: thumbnails, scripts, rendering, everything. Works well for sleep channels and other ai image/video slideshows. We're gradually opening up access. Link in bio

facelessjourney | YTA & SaaS

13,259 görüntüleme • 5 ay önce

Hyperspace: The Agentic OS Apple Should Have Built On December 19th, 2024, we announced the world’s first Agentic Browser. What followed was a movement — a new category was born which led to many early products in this space and recently the hundreds of people lining up outside the The Agentic Browser Summit in San Francisco underscored that. Silicon Valley instinctively gets it, from students to tech executives, people can feel a revolutionary new change in computing is in the air. Past year taught us why such a product was inevitable, a hard engineering effort, and also the last mover in the entire software world this decade if and when done right. All paths are headed in the same direction: one tool which orchestrates them all. At Hyperspace we showed that path with essays and products we launched in earlier months: from a spatial UI of orchestrating agents, to showcasing transparent activity in how the AI system operates which leads to user trust, to presenting the software end-game, which massively improves human productivity. We also built the world’s largest AI network, drawing participation from people in almost 6000 cities around the world contributing their machines as nodes in the network. Think Uber, but for AI. That is, planetary-scale. And now we are stretching this industry ambition further with our end-to-end vision of the Agentic Supercomputer, the first breakthrough new AI OS, and an effort which spans from AI research to distributed systems to inventing a new UI to inventing a new business model to complement it. All of this together helps us in serving our mission, of delivering “Everyone’s Personal Supercomputer”. While others have built AI-native browsers, no one though has built something agentic from the ground up — with AI as the foundation, not a feature. How do you fundamentally improve the lives’ of billions around the world ? We believe that requires building a native environment for agents to be viewed, created, deployed, executed, discovered and priced in. That is a world where we move on from static apps, to dynamic agents. But, as my 2 year old niece likes to ask: “but why ?” The issue is that the world of software today is fragmented, and everyone is sprinkling on AI as a feature and charging a subscription fees for it. From browser makers, to IDEs, to design and other productivity tools. This leads to a fragmented UX, where people have to learn to use AI in each app, their memory and other context is not shared between all these apps, and they also have to pay separately for compute for each such AI-enhanced app. Each app maker has to figure out basics such as compute, and leads to the issues we saw with Cursor pricing recently. This is not the future. What if AI was the foundation instead of a feature ? What if Apple had built a fundamentally new AI OS from the ground up and what would it have looked like ? At Hyperspace, that is what we did. On July 15th we introduced three breakthrough key pillars of our AI OS: 1. Agentic Browser 2. Agentic Memory 3. Agentic Payments And we didn’t stop there. We also introduced a breakthrough new user interface called the Spatial AI which is inspired both from the spreadsheet and the HyperCard - each card is an agent, with it’s own inputs and outputs, endlessly extensible and pluggable with others, just like cells of a spreadsheet. Update one cell and all the dependents update, like a spreadsheet formula. It goes beyond a static linear workflow to being able to operate in all directions. This revolutionary new interface helps manage all of the below: 1. Multiple websites being browsed in parallel 2. Multiple desktop apps being browsed in parallel 3. Multiple server tools being used in parallel 4. Multiple smartphone apps streamed to your device or opened via an emulator All the software which you need comes together in this one seamless, agent-native interface. This interface provides you access to the largest network of models, vectors, agents and compute on the planet. The Browser. The IDE. The Notepad… they are not separate products: they are all in one, the Agentic Browser. As Steve Jobs famously said at the iPhone announcement, “are you getting it ?” And beneath this UI lies a new intelligence routing layer — leveraging both swarms of specialized models to the Hyperspace Matrix model that recalls thousands of tools in real-time, not by context window hacks, but through retrieval, ranking, and reuse. To many, this will feel like AGI. Not one big system by one big company, but an intelligent network. Now lets talk about privacy… Are you comfortable with one company owning all your memory forever ? I am not. So we have invented Agentic Memory as a new open protocol which provides full power over memory to you, the user. Your memory is yours, encrypted, on your device, and portable if and how you want. Anyone can build on it without our permission, but not without your permission. This protocol, and the decentralized vector database spread out across the world, would enable apps and agents to share context and memory. Think copy-paste, but for the AI world. It doesn’t just remember — it knows what matters. VectorRank helps your AI weigh your life’s most relevant moments over time, just like the way our minds elevate memories. Now each time you use an agent, your experience with other agents will also continuously improve: you don’t have to keep repeating the same things about yourself, while fully preserving your privacy. Agentic Memory is accessible within the Agentic Browser to manage. And there is one more thing… AI as the foundation requires compute to be available at the base layer, but this base layer spans models running on your own device, to cloud APIs, to also running across the peer-to-peer distributed network. Agentic Payments provides a singular interface to all of that compute, running a spot auction clearing marketplace every second to determine the fair price of compute. This results in price transparency, and you as the user paying the lowest possible cost. If you want predictability, you can reserve compute in advance. This end-to-end system provides the most streamlined world for agents to operate in. In order to enable this world and the world of agents being able to pay each other in sub-cent increments millions of times a second, we had to also invent a fundamentally new agentic micropayments blockchain. All of this together would enable a world where you as a user, or the agent itself, can efficiently call and utilize other agents built by others and also pay for content which is unique and useful. This enables a move away from the current AI exploitative economy for bloggers and other content creators, to a web with a fundamental new business model. Earlier we didn’t have the right infrastructure to enable such a world. Now, all the dots connect. The Hyperspace AI OS would give the power of a supercomputer in everyone’s hands. This isn’t a browser, or an IDE or limited to any device or cloud. It’s an entire AI operating system — with a breakthrough new spatial UI, local and distributed compute, agentic memory, agentic payments, and orchestration built into the foundation. As a user, we move the choice back in your hands with an experience you will love and find delightful. You get to choose the level of privacy, cost, and utility you want. And while Apple should have done it, we could not wait, and we feel this just required a new level of passion and DNA which we bring here. We are just getting started. Thank you, Varun Mathur Cofounder and CEO, Hyperspace cc Naval Marc Andreessen 🇺🇸 Vinod Khosla Andrej Karpathy Sam Altman

Varun

169,177 görüntüleme • 1 yıl önce

Today, we’re announcing Kosmos, our newest AI Scientist, available to use now. Users estimate Kosmos does 6 months of work in a single day. One run can read 1,500 papers and write 42,000 lines of code. At least 79% of its findings are reproducible. Kosmos has made 7 discoveries so far, which we are releasing today, in areas ranging from neuroscience to material science and clinical genetics, in collaboration with our academic beta testers. Three of these discoveries reproduced unpublished findings; four are net new, validated contributions to the scientific literature. AI-accelerated science is here. Our core innovation in Kosmos is the use of a structured, continuously-updated world model. As described in our technical report, Kosmos’ world model allows it to process orders of magnitude more information than could fit into the context of even the longest-context language models, allowing it to synthesize more information and pursue coherent goals over longer time horizons than Robin or any of our other prior agents. In this respect, we believe Kosmos is the most compute-intensive language agent released so far in any field, and by far the most capable AI Scientist available today. The use of a persistent world model also enables single Kosmos trajectories to produce highly complex outputs that require multiple significant logical leaps. As with all of our systems, Kosmos is designed with transparency and verifiability in mind: every conclusion in a Kosmos report can be traced through our platform to the specific lines of code or the specific passages in the scientific literature that inspired it, ensuring that Kosmos’ findings are fully auditable at all times. We are also using this opportunity to announce the launch of Edison Scientific, a new commercial spinout of FutureHouse, which will be focused on commercializing our agents and applying them to automate scientific research in drug discovery and beyond. Edison will be taking over management of the FutureHouse platform, where you can access Kosmos alongside our Literature, Molecules, and Precedent agents (previously Crow, Phoenix, and Owl). Edison will continue to offer free tier usage for casual users and academics, while also offering higher rate limits and additional features for users who need them. You can read more about this spinout on our blog, below. A few important notes if you’re going to try Kosmos. Firstly, Kosmos is different from many other AI tools you might have played with, including our other agents. It is more similar to a Deep Research tool than it is to a chatbot: it takes some time to figure out how to prompt it effectively, and we have tried to include guidelines on this to help (see below). It costs $200/run right now (200 credits per run, and $1/credit), with some free tier usage for academics. This is heavily discounted; people who sign up for Founding Subscriptions now can lock in the $1/credit price indefinitely, but the price ultimately will probably be higher. Again, this is less chatbot and more research tool, something you run on high-value targets as needed. Some caveats are also warranted. Firstly, we find that 80% of Kosmos findings are reproducible, which also means 20% are not -- some things it says will be wrong. Also, Kosmos certainly does produce outputs that are the equivalent to several months of human labor, but it also often goes down rabbit holes or chases statistically significant yet scientifically irrelevant findings. We often run Kosmos multiple times on the same objective in order to sample the various research avenues it can take. There are still a bunch of rough edges on the UI and such, which we are working on. Finally, we are aware that the 6 month figure is much greater than estimates by other AI labs, like METR, about the length of tasks that AI Agents can currently perform. You can read discussion about this in our blog post. Huge congratulations to our team that put this together, led by Ludovico Mitchener and Michaela Hinks: Angela Yiu, Benjamin Chang, Sid Narayanan, Edwin Melville-Green, Albert Bou, Arvis Sulovari, Oz Wassie, Jon Laurent. A particular shout out to Michael Skarlinski and his team that rebuilt the platform for this launch, especially Andy Cai Andy Cai, Richard Magness, Remo Storni, Tyler Nadolski Tyler Nadolski, Mayk Caldas Mayk Caldas, Sam Cox Sam Cox and more. This work would not have been possible without significant contributions from academic collaborators Mathieu Bourdenx, Eric Landsness, Dániel Barabási, Nicky Evans, Tonio Buonassisi, Bruna Gomes, Shriya Reddy, Martha Foiani, and Randall Bateman. We also want to thank our numerous supporters, especially Eric Schmidt, who has been a tremendous ally. We will have more to say about our supporters soon!

Sam Rodriques

732,730 görüntüleme • 10 ay önce

Introducing BioCLIP: A Vision Foundation Model for the Tree of Life A foundation model that strongly generalizes on the tree of life (2M+ species), outperforming OpenAI CLIP by 18% in zero-shot classification, and supports open-ended classification over almost the entire tree of life What's the secrete ingredients? > Data: we curate and release TreeOfLife-10M, the largest and most diverse ML-ready dataset of organism images to date. It contains 10.4M images for over 450K taxa, sourced from iNaturalist, BIOSCAN, and Encyclopedia of Life. > Modeling: we creatively repurposes CLIP's multimodal contrastive learning objective for hierarchical image classification. The autoregressive language model naturally encodes the hierarchy of the tree of life taxonomy, which in turn bakes the hierarchical representation into the vision transformer encoder. Key results > Strong zero/few-shot classification for animals/plants/fungi, including rare species, outperforming CLIP by avg 16-18% absolute. > T-sne visualization shows that BioCLIP's vision encoder has captued the fine-grained hierarchical structure of the tree of life > BioCLIP is a kind of universal classifier for the tree of life. Just give it an organism image and it will likely find the correct species (among top 5)! But use it with caution; it's not perfect yet.. Final remarks > AI for Science is really hard but extremely rewarding! It took us a ton of time (1+ year) and frustration trying to find a plausible way to integrate the tree of life taxonomy into foundation model training. But when the "Eureka!" moment came and the idea hit us (by the great Wei-Lun Chao) that CLIP's multimodal contrastive learning objective can be repurposed for that, everything just follows naturally. It was truly a moment of joy and excitement! > BioCLIP is our first attempt at foundation models for biology, but it certainly won't be the last! There's so much more to do at the intersection of one of the oldest scientific disciplines and the young but thriving field of AI. Biological intelligence is the foundation for artificial intelligence, and artificial intelligence will in turn become the most important tool for us to unraval the mysteries of biological intelligence. We are hiring postdocs and PhDs in the NSF Imageomics Institute institute to explore this exciting field! Drop us an email. also happy to chat about it at #NeurIPS2023 with any of Tanya, Wei-Lun Chao, or me. - paper: - project: - demo: - model: - data (TreeOfLife-10M): to be released on Hugging Face soon joint work with the amazing Imageomics Institute team: @samstevens6860 Lisa Wu, Matt Thompson, Elizabeth Campolongo Chan Hee (Luke) Song David Carlyn Li Dong Wasila Dahdul Chuck Stewart, Tanya Berger-Wolf Wei-Lun Chao Yu Su

Yu Su

80,660 görüntüleme • 2 yıl önce

My second book (of now five) was on contextual software at Written to get ahead of Salesforce's strategy (each of our five books were written for the same reason -- to get ahead of the big company strategists -- and each did exactly that). Marc Benioff (Salesforce's CEO) wrote the forward because he liked my work. Today I saw a piece of the future of X. So this post is for Elon Musk. Grok and X needs what I just saw. Two guys from India visited my house this morning to show it to me: Jibin Mathew and Jikku Jose, founders of StoryBrain. The everything app needs to move way past TikTok if it's really going to grow. So, everything here, especially company posts to sell people on things, must become more visual and personalized (companies here are a big deal, I've found 3,600 AI companies already here and mapped them all out at They showed me that their system can generate attractive videos from: 1. An Amazon product page. 2. A PDF. 3. A Walmart product page. Or really any product page. The AI ingests them, finds important things on them to present to sell the product, and generates media for doing just that. Now why does this matter? Well, these "adaptive media" pieces need: 1. To be translated to all languages. India, alone, has more than 20 languages depending on what part of the country you are in. 2. Get the key selling/story points. 3. Learn from users. If everyone buys more on product pages with red backgrounds, it can learn that and change everything to red. 4. Work on every device. Not just new iPhones. Old shitty Android phones (no AR). TV's. Laptops. Tablets. Even watches. 5. Be personalized for culture and better results. Their team is half AI developers and has been working on generative AI to do this for four years. AI generation changes the game. If X is going to be the everything app that Elon says he wants to turn X into, then it must move toward a much more visual interface, and, soon, into 3D as well due to Vision Pro and competitors. In other words, X needs something like this. Why? 30% of the human brain is for visual processing. They process visual pieces like this much faster than the text that GPT generates. There are many millions of people who can't read at all, too. Can someone who is selling, say, jewelry, or travel, or, "everything," easily make these new pieces to push to social media? In an "everything" app it needs to be super simple. Also, by generating these on the fly (with some precaching) you can ensure they are always up to date and topical. This was my second startup of the morning. Starting the week with a bang!

Robert Scoble

21,188 görüntüleme • 2 yıl önce

$AMD Massive Rotation from $NVDA $INTC🧵 Not Financial Advice! DYOR! 5-10 minutes before the bell today, last trading day of May 2026, massive rotation out of $INTC and $NVDA into $AMD. I wrote this thread this morning on what $TSM said on Energy Efficiency is now TOP Priotity and why AMD is the biggest winner. Of course I did not have influence on this rebalancing, I was just pointing out why Dr. Su saw this coming years ago. (Check the picture to understand more). I been talking about Agentic AI for like 3-4 years now. OpenClaw broke the CPU:GPU Ratio 1:4 narrative to 1:1 to 5:1 in late Jan and Feb 2026. I will link various threads where you can understand the full picture from supply chain, to TSMC expansion, and different Wafer Ratio for EPYC Venice and MI455X. Energy efficiency is a structural, long-term driver behind institutional rotation from $NVDA and $INTC into $AMD (with spillover strength in $AVGO for complementary networking/custom silicon). This isn't just short-term rebalancing, it's a massive bet on the shift from AI training (performance-at-any-cost) to inference, deployment, and embodied/agentic systems (where total cost of ownership, power draw, and scalability dominate). Precisely What I been writing about $AMD for years now, probably at least more than 5,000 threads.This is the FOMO from Institutions to own $AMD. Do know that AMD is the least owned Semi Stock among vs Peers. AI infrastructure is moving beyond massive training clusters to widespread inference for Agentic AI (running models 24/7) and embodied AI (robots, autonomous agents, edge devices). These workloads prioritize: ~Tokens-per-watt and performance-per-watt ~Lower total power consumption for data centers facing grid constraints ~Better economics at scale (cost-per-token, TCO) ~Thermal and power efficiency for on-device/robotics use Hyperscalers are now thinking more about Margin, Profitability, and $/M Tokens At $516/share. AMD Fwd PEG Ratio is still 35/100+= 0.35 AKA very cheap IMO for the growth and potential. A. Why institutions rotated out of $NVDA? Because Agentic AI is going to dominated by CPUs for years to come, moving violently to 5-10-20:1 CPU:GPU Ratio as enterprises are demanding more than 10-20 agents to run tasks. Now, that does not mean training is going away, Inference is just going to grow much faster. B. Why instiutitons rotated out of $INTC? Because AMD x86 unit share is only at 30-31% but Revenue share is already at 46.2% according to Mercury Research. And Dr. Su wants 50-60% market share, and that would mean 60-70%+ Revenue share where the CPUs TAM Is now already at $200B in 2026 and projected to be $500B by 2030. C. Why $AMD? Because AMD secured meaningful 2nm Capacity, Advanced Packaging and Memory through 2027-2028. And TSMC is expanding 2 primary 2nm Fabs toward 60-65k WPM each, and speeding up 5 2nm Fabs in Taiwan. With total up to 12 2nm Fabs through 2027/2028. 2nm Capacity is expected to be 140k+ WPM toward end of 2026, and 220-240k WPM by end of 2027. Apple has secured 35-45k WPM. And AMD does not have to worry about allocation competition until late 2027 from $AVGO for $META and $GOOGL(This may change) D. Agentic AI will evolve to 24/7 Autonomous Agent, and that will become the foundational layer for Robotic or Physical AI. Agentic AI (autonomous systems that plan, reason, use tools, self-correct, pursue long-horizon goals, and adapt) provides the high-level cognitive architecture. It turns raw perception and low-level control into useful, general-purpose behavior in the physical world. Physical AI (or Embodied AI) refers to AI that senses, understands, and acts directly in the real world through robots, actuators, and sensors. Agentic capabilities are what make this scalable and useful beyond narrow, scripted tasks. Reactive/programmed machines → To proactive, goal-oriented autonomous agents. How does this work? Autonomous Agent layer is the brain ~Vision-Language-Action models or robotics foundation models. ~Agentic loops: Planning, chain-of-thought reasoning, reflection, tool use (simulators, APIs), multi-step task decomposition. ~Persistent 24/7 operation with Memory, world modeling, continuous learning. Institutions may not like $AMD from 2022-2025, but they cannot stop this evolution and it is inevitable. Part of my main thesis for AMD to get to $5 Trillion Market Cap Long Term. Conclusion: Institutions are rotating capital toward AMD not merely for tactical rebalancing, but because Dr. Lisa Su and her team anticipated this exact inflection years in advance and have been methodically engineering AMD’s platform to dominate it. Dr. Su has long championed the convergence of Agentic AI as the high-level cognitive foundation for Physical AI and robotics. As far back as her 2023/2024 CES keynote and earlier strategic commentary, she described Physical AI (including humanoid robotics and edge autonomy) as “the next big thing”; a natural extension of agentic workflows moving from digital reasoning to real-world action. She emphasized that enabling persistent, 24/7 autonomous agents requires a full-stack approach: high-performance CPUs for orchestration and motion control, dedicated accelerators for real-time vision and multimodal inference, and open software ecosystems for rapid development. This vision aligns precisely with the structural drivers we’ve discussed. As AI shifts from training to massive-scale inference and embodiment, energy efficiency, total cost of ownership, and heterogeneous compute become first-order advantages. AMD’s Instinct MI350/MI355 series, Ryzen AI Embedded processors, and EPYC platforms deliver superior performance-per-watt and balanced CPU + GPU + NPU integration ideal for power-constrained robots that must run sophisticated agentic reasoning loops without excessive thermal or battery drain. Dr. Su has repeatedly highlighted the rising importance of CPUs in agentic systems (moving toward 1:1 or even CPU-heavy ratios with GPUs), positioning AMD’s strengths in orchestration, memory handling, and efficiency as critical for the next phase of growth. AMD is engineered for the deployment realities of embodied agents: scalable, efficient, and deployable at the edge and in physical systems. The institutional flows out of NVDA and INTC into AMD reflect recognition of this prepared leadership. Dr. Su didn’t just see the future of Agentic AI powering robotics, she has spent years building the silicon, software, and partnerships to make it practical and economically viable. This rotation signals confidence that the companies best positioned for the physical, always-on intelligence layer will capture the highest-volume opportunities in the coming decade. Not Financial Advice! DYOR!

Mike

104,109 görüntüleme • 3 ay önce