正在加载视频...

视频加载失败

🚀 Introducing Articulate Anymesh – now open-sourced! An automated framework behind our Genesis simulator, capable of transforming any rigid 3D mesh into its articulated counterpart using an open-vocabulary manner! Given a 3D mesh, our framework uses VLMs + visual prompting to extract rich semantics — enabling part segmentation and...

35,936 次观看 • 1 年前 •via X (Twitter)

6 条评论

MΞntis 🇦🇺 的头像
MΞntis 🇦🇺1 年前

@Scobleizer This is absolutely one step closer to the Holodeck @Scobleizer

Rainmaker 的头像
Rainmaker2 年前

🚀 Discover how Reinforcement Learning can transform trading strategies! Check out my free Substack for full code and backtest of the completed Q-learning algorithm. Don't miss out! 📈✨

Lars Vågnes eu/acc 的头像
Lars Vågnes eu/acc1 年前

So elegant

𝐒𝐇𝐎𝐂𝐊𝐖𝐀𝐕𝐄 ⚡️🌊 的头像
𝐒𝐇𝐎𝐂𝐊𝐖𝐀𝐕𝐄 ⚡️🌊1 年前

Now explain like I'm a human being.

NexGen Cloud 的头像
NexGen Cloud1 年前

Looks incredible! Can't wait to see how the community embraces it 👏

Jose_TechAI 的头像
Jose_TechAI1 年前

How you use?

相关视频

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,821,175 次观看 • 1 年前

We’re thrilled to share that our MERFISH+ preprint is now live on bioRxiv!👉 In this work, the Bintu and Zhu labs (UCSD) developed MERFISH+, a next-generation spatial genomics platform that combines genome-wide RNA and epigenetic imaging over a large field of view. By introducing acrydite-modified probes covalently anchored to hydrogels, MERFISH+ achieves remarkable imaging stability and enables >1,800-gene, multi-modal, and multi-month experiments. With this platform, they, together with the Chi lab at UCSD, profiled a whole developing human heart at 12 post-conception week with merely two slides, resulting in a total of 53 slides, 3.1 million single cells and more than 30 cell types. Building upon our previous 3D reconstruction and modeling framework, Spateo ( we reconstruct the 3D human heart that nicely captures the anatomical structure of the heart, including the intricate vasculature network. Sophisticated analyses provide a holistic view of an entire organ and enable systematic characterization of 3D cellular neighborhoods and transcriptional gradients of substructures such as the descending arteries. Furthermore, using a generative integration framework for spatial multimodal data (Spateo-VI), we harmonized these MERFISH+ transcriptomic and chromatin data to reconstruct a 3D spatially-resolved multi-omics atlas of the developing human heart, shared at and MERFISH+ thus sets a new standard for large-format, multi-omic spatial profiling, enabling holistic, 3D characterization of organs at subcellular resolution. Huge congratulations to first authors Colin Kern, qingquan Zhang, Yifan Lu , and Jacqueline Eschbach, and to all collaborators from the Bintu, Zhu, Chi, and Qiu labs for this amazing team effort. Thanks for your diligence, creativity, and hard work on this project. We’re grateful for support from Arc Institute and our generous donors. Our lab is expanding—if you’re excited about building the next generation of single-cell and spatial genomics techniques and predictive single cell and spatial foundation models, we’re hiring! If you are interested, please reach out to me via direct message or email at [email protected]. We are excited for any potential collaborations along this line of research in Stanford, UCSF and Berkeley and other labs as well.

evo-devo

42,326 次观看 • 10 个月前

In my past research experience, finding or developing an appropriate simulation environment, dataset, and benchmark has always been a challenge. Missing features, limited support, or unexpected bugs often occupied my days and nights. Moreover, current simulation platforms are relatively fragmented—making it challenging to replicate the success of the RT-X dataset in unifying community efforts. Introducing RoboVerse, we provide a unified platform, dataset, and benchmark for scalable and generalizable robot learning. We hope to build a shared foundation to combine the community efforts. RoboVerse includes: MetaSim: We carefully designed a configuration system and a universal interface to align current robotic simulators. With MetaSim, you can use any simulator with the same code—bringing together the community’s diverse efforts under one framework! RoboVerse Dataset and Benchmark: We unify popular simulation environments and benchmarks into a single cohesive system and introduce the RoboVerse dataset—a large-scale, high-quality synthetic dataset. Additionally, we propose a standardized benchmark across both imitation learning and reinforcement learning. A cool feature enabled by our unified framework: Hybrid Simulation! You can now integrate physics engines and renderers from different simulators—e.g., using MuJoCo precise physics with Isaac photorealistic rendering. This not only elevates simulation fidelity but also significantly enhances real-world transfer performance across complex robotic applications. Hopefully, our team’s efforts could serve the robotic community to thrive vibrantly in the years to come. RoboVerse is open-sourced🥳!!! Project Page: Documentation: Github Repo: Paper:

Haoran Geng

84,318 次观看 • 1 年前

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,079,018 次观看 • 3 个月前

🎉 The best way to start the week is to find out that our MedSAM is finally published today in Nature Communications! **Segment anything in medical images** Paper: arXiv: Data & Code: MedSAM is the first promotable foundation model for medical image segmentation. **Highlights**: ⭐ Before its formal publication, we have received 220 citations and 1400+ GitHub stars 🙏🙏❤️‍🔥❤️‍🔥❤️‍🔥 📊 We curated a large-scale medical image dataset with 1,570,263 image-mask pairs, covering 10 imaging modalities and over 30 cancer types. 🚀 Built on top of SAM (AI at Meta ) with transfer learning, we have significantly enhanced its segmentation performance of medical images. 📈 Comprehensive evaluations of 86 internal validation tasks and 60 external validation tasks demonstrate its better accuracy and robustness than modality-wise specialist models. **What is Next? --- Clinical Translation!!** 🍕Our next goal is to make the model deployable on laptops (CPUs) or other edge devices without reliance on GPUs. We have distilled a lightweight model, LiteMedSAM, offering a speed boost of 10x while maintaining accuracy. Plus, we have integrated it into the 3D Slicer plugin, providing an efficient tool for medical image segmentation. 🌐 To further promote developments in this field, we organize a competition on #CVPR2026: Segment Anything in Medical Images on Laptop! An out-of-the-box baseline has been released to reduce the entry barriers. Welcome to join us to push the boundary further: 🙏 Massive thanks to MetaAI AI at Meta for their open-source project SAM and many reviewers/users for their invaluable feedback. A huge shoutout to my postdoc Jun Ma (JunMa) for his leadership on this project!! UHN AI Hub Vector Institute Peter Munk Cardiac Centre AI Department of Laboratory Medicine & Pathobiology U of T Department of Computer Science University of Toronto University Health Network Brad Wouters 🇨🇦 Barry Rubin MD, PhD, FRCSC Shaf Keshavjee

Bo Wang

140,229 次观看 • 2 年前

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 次观看 • 1 年前

Exciting update on PantheonOS: Introducing Pantheon-Notebook & Pantheon-CLI — the first fully open-source, Python-based agentic tools that go beyond Claude Code in the field of data analysis. Pantheon-CLI runs entirely on your computer or server, supports 60+ tools and 50+ databases, and can call any Python, R, or Julia package alongside natural language. Chat with your data directly. It look like python-claude-code, but more appreciate for data analysis. Pantheon-Notebook brings the same agentic framework into Jupyter! Not just for writing code, it can also run and revise code automatically to generate the correct result, and even operate on files and study from website — beyond what any other tool can do! With Pantheon, you mix natural language + programming in one workflow, focusing on discovery instead of syntax barriers. We've applied Pantheon in some real-world cases: finance (customer explore), biology (Seurat, cell segmentation, annotation), sociology (survey analysis), and drug discovery (molecular docking). Pantheon is not just a CLI or a plugin — it's an agentic operating system for science, spanning both terminal and notebook. Why not try it now? We are actively preparing publications from this series of projects. Major contributors will be recognized in our GitHub repository and listed as key authors in these manuscripts. Feel free to reach out for collaborations, research assistant positions, visiting opportunities, rotation project or future PhD projects.

evo-devo

45,652 次观看 • 1 年前

I am happy to be finally able to post what I was able to build over the last few weeks. A full real-time high-frequency state estimation and mapping algorithm completely written line by line from scratch in Rust, which can be used by robots to navigate and reason within the 3D world also in complicated scenarios. TBH this took me longer than expected (which was still super fast :D) but you need to get a lot right: From the sensors over the drivers to their respective estimation pipeline and then fusing everything together - a covariance nightmare - and something that can be refined over years to come (currently using Fisher Information from the real measurements). What you see here is not the output of some structure from motion or Gaussian splatting, these are the points of a tight mesh (high res for the video) that a robot can use in real time to plan a path using any open-source planner. The flight you experience through the world is the actual state estimate of the scanner which is published at IMU rate. Yes, currently we have some artefacts of filtered-out humans (GDPR compliant of course :) ) and moving cars and there is still some calibration that could be improved. Offline refinement with SFM and Gaussian splats is possible as well but currently not on the road map. What is on the road map is an exciting step of now being able to collect data from customers at construction sites and in warehouses (currently handheld in the near future with a robot). This data can then be used by our physical agents to reason within this world and automate any customer’s task related to 3D data. If you have anyone who wastes time manually looking 👀 through 3D data, or cannot collect enough 3D data and interpret: Tell me how to reach them!

Benedikt Seidel

16,671 次观看 • 4 个月前

Your our history shouldn’t be taught in a boring way 🫠 I’ve always loved history, but I could never imagine what life actually looked like back then. So I built Empire Atlas, a 3D interactive explorer of 8 historical empires using Three.js and Kimi.ai K3 🔥 It lets you explore how people lived, what their homes looked like, their maps, daily life, interiors, and more. Properly researched. But the craziest part is that this was near one shot vibe coded with Kimi K3 🤯 When I previously built a 3D anatomy app with GPT 5.6 sol, I had to iterate on performance and optimization. With Kimi, the moment I handed over the 3D assets (generated using Tripo), prompt and design (by GPT Image 2.0), it created an 11 step engineering plan to build the entire thing. The very first step it did was optimizing the assets. It took nearly 500MB of 3D assets and brought them down to just 17.8MB using mesh simplification, Draco compression, and 1024px WebP textures. Absolutely nuts. It also generated 56 historical images across the 8 empires showing daily life, maps, interiors, and more using its image plugin with batch processing. Those were converted to WebP too, bringing the total image size to around 10MB. That’s a huge reason the experience loads so fast on website. It's engineering workflow or intelligence has really impressed me so far. The only downside is that it took more than 5 hours, though 😅 Anyway, back to history. In Empire Atlas, you can explore 8 different empires and see how people and our ancestors lived at that time. I really love those textures I was able to create using Tripo. You can explore their homes in 3D, and there’s so much more we could do with this. We could extend these houses into fully explorable interiors and create increasingly realistic reconstructions of what life actually looked like. And maybe create fun education games too. I genuinely think this can make history education so much more immersive. Much more than showing black and white images in boring textbooks. Go explore your history now 👇 Live: Code:

The Bugged Dev

120,448 次观看 • 1 个月前

If an AI can control 1,000 robots to perform 1 million skills in 1 billion different simulations, then it may "just work" in our real world, which is simply another point in the vast space of possible realities. This is the fundamental principle behind why simulation works so effectively for robotics. Real-world teleoperation data scales linearly with human time (< 24 hrs/robot/day). Sim data scales exponentially with compute. There are 3 big trends for simulators in the near future: 1. Massive parallelization on large clusters. Physics equations are "just" matrix math at their core. I hear GPUs are good at matrix math 🔥. One can run 100K copies of simulation on a single GPU. To put this number in perspective: 1 hour of wallclock compute time gives a robot 10 years (!!) of training experience. That's how Neo was able to learn martial arts in a blink of an eye in the Matrix Dojo. 2. Generative graphics pipeline. Traditionally, simulators require a huge amount of manual effort from artists: 3D assets, textures, scene layouts, etc. But every component in the workflow can be automated: text-to-image, text-to-3D mesh, and LLMs that write Universal Scene Description (USD) files as a coding exercise. RoboCasa is one example of a prior work. 3. End2end neural net that acts as simulator itself. This is still bluesky research and quite far from replacing a graphics pipeline, but we are seeing some exciting signs-of-life based on video gen models: Sora, Veo2, CogVideoX, Hunyuan (text-to-video); and action-driven world models: GameNGen, Oasis, Genie-2, etc. Genesis does great on (1) for certain tasks, shows good promises on (2), and could become a data generation tool for reaching (3). Its sim2real capabilities for locomotion are good, but there's still a long way to go for contact-rich, dexterous manipulation. It shows a bold vision and is on the right path to providing a virtual cradle for embodied AI. It is open-source and puts a streamlined user journey at the front and center. I had the privilege to know Zhou Xian and play a small part in his project since a year ago. Xian has been crunching code non-stop on Genesis with a very small group of core devs. He often replied to my messages at 3 am. Zhenjia Xu from our GEAR team helped with sim2real experiments in his spare time. Genesis is truly a grassroot effort with an intense focus on quality engineering. Nothing gives me more joy than seeing the simulation ecosystem bloom. Robotics should be a moonshot initiative owned by all of humanity. Congratulations.

Jim Fan

157,343 次观看 • 1 年前