🚨 Meta introduces Segment Anything Model 2 (SAM 2)... — the first unified model for real-time, promptable object segmentation in images & videos. 🚀SAM 2 is available today under Apache 2.0 so that anyone can use it to build their own experiences ➡️show more

Zhengzhong Tu
181,296 views • 2 years ago
Today we're releasing the Segment Anything Model (SAM) —... a step toward the first foundation model for image segmentation. SAM is capable of one-click segmentation of any object from any photo or video + zero-shot transfer to other segmentation tasks ➡️show more

AI at Meta
3,571,190 views • 3 years ago
Here are the steps to create this animation: 1.... Design and model your own mech in Plasticity. 2. Import the model into the animation software 3ds Max. 3. Create the animation on your own. 4. Render and export the monochrome model animation. 5. Make the final render for reference. 6. Use Seedance 2 to re-render your monochrome model animation according to the reference render you provided. Summary: In this workflow, Seedance 2 acts as a renderer that saves you hours of rendering time. All other work has to be done manually by yourself. Don’t blindly believe that AI can fully produce mech animations for you — it simply cannot do that.show more

LAS91214
47,667 views • 3 months ago
🚀 The Segment Anything Model (SAM) has been upgraded... to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearchshow more

Bo Wang
178,579 views • 2 years ago
Player stats UI animation with GPT Image 2 and... MiniMax H3. This model is amazing for UI animations. Give GPT Image 2 your character image and ask it to create a stats UI. Then use that UI as a reference for MiniMax H3. You can check the prompt in the replies.show more

Kōda
41,985 views • 1 month ago
Just observed that Gemini 3.7 Flash in Antigravity can... now analyze videos natively.. strange that it wasn't explicitly mentioned.. Anyways, a very great addition.. earlier the model had to use tools like ffmpeg or py scripts to extract frames and view individual images do build an context about the given video.. Now it's just so clean...show more

Bee
20,105 views • 9 days ago
DeepSeek R1 is *the* best model available right now.... It's at the level of o1, but you can use it for free, and it's much faster. A huge leap forward that nobody saw coming. No wonder so many people are throwing tantrums online trying to discredit the Chinese students who built this. You can use DeepSeek in Visual Studio Code right now: 1. Install the Qodo Gen AI extension 2. Select DeepSeek R1 from their list of models The Qodo team is hosting DeepSeek on their servers, so none of your data will go to China. I've been building a Tetris game using DeepSeek, and this is the most impressive model I've seen so far.show more

Santiago
1,224,381 views • 1 year ago
Very proud to see the FourCastNet model that I... helped build at NVIDIA in collaboration with Berkeley Lab and other universities be highlighted in Jensen's keynote at #GTC24 - FourCastNet was the very first high-resolution #AI model to show competitive performance for weather forecasting and is part of Earth-2 services. It is now running at ECMWF and you can get 2-week forecast there in realtime. Hear Bill Collins from Berkeley Lab leverage FourCastNet for studying extreme weather events, offering accurate and efficient results from huge ensembles at a lower computational cost.show more

Prof. Anima Anandkumar
26,434 views • 2 years ago
AI hooks are f*cking insane you can create some... insanely realistic hooks with seedance 2.0 that grab attention right away and subtly market your product in the background this is a great way to market any product and you can use literally anyone for these videos in any setting or scenario... capturing attention is literally all that matters. once you have that you can build anything around it prompt below ⬇️show more

Miko
30,843 views • 4 months ago
We are sharing a major update to our General... World Models efforts. GWM Worlds 2 can generate full, interactive, real-time video simulations in one continuous 720p stream at 24 fps and audio at 48000 Hz, responding to your inputs as you explore. It is the first world model that generalizes to arbitrary actions rather than a fixed set of actions or just navigation. Dynamic actions are what actually matter for building gaming and interactive experiences, and for training agents in these simulations. We also introduced WorldPrompt, a way to tell GWM-2 that a world can be split into two kinds of state: what persists and what changes over time. For example: persistent gravity, physics, and light, alongside actions that describe movement, gestures, object interactions, speech, and sound.show more

Cristóbal Valenzuela
247,890 views • 7 days ago
ARB8i 💛 Lab - is LIVE 🚀 This started... as a speed tool for my @Doodles edits, I turned it into an editor anyone can use in seconds! Live with my packs + 2 overlays by artdood. Now opening to community artists - more to come! 👀 Link in the first reply 👇show more

ARB8i 💛
40,246 views • 1 year ago
seedance 2.0 + my v2 AI UGC prompting system... is giving insane results i spent the last 24 hours generating over 200 seedance 2.0 videos to figure out the best prompting framework system for AI UGC this video was made with 1 prompt and 1 tool, no editing was done to the video this was just a prompt to a video this is by far the best model i've ever used and the craziest part is that it can be fully automated this is the first time we can actually automate high quality ai ugc at this level bytedance owns tiktok so this model is trained on millions of high quality ugc videos. you just need to know how to extract that and call it in your prompt. we are so early... it's insaneshow more

Miko
81,462 views • 7 months ago
Probably the first of a kind... a Substrate node... executing a Plutus script to validate an UTxO transaction. This is a project by TxPipe called Griffin, a Substrate-based framework for building app-chains using the eUTxO model and the Plutus language. Cardano devs can use it to build partnerchains that leverage their exiting expertise around UTxO programmability.show more

Santiago Carmuega
10,281 views • 1 year ago
We just launched Claude Sonnet 4.5 and “Imagine with... Claude”, a 5-day preview that creates dynamic experiences for you in real-time (and the first product I’ve worked on at Anthropic!). It pioneers the concept of “model-as-backend”, using a model to not only generate interfaces on the fly but also power all the functionality behind it – all made possible by our newest model, Claude Sonnet 4.5. Here’s a "choose your own adventure" version of my founder journey, as retold by Imagine with Claude:show more

Sean Strong
115,693 views • 11 months ago
This week is already so hot. 🔥 Massive release... from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."show more

Rohan Paul
12,761 views • 7 months ago
Meta just launched its 1st image model after Mark... Zuckerberg’s AI shake-up. Muse Image is Meta Superintelligence Labs' first image generator after Muse Spark. They said the model will power new editing features in Meta’s Instagram photo app and be added to its marketer tools for creating platform ads. Meta had relied on Midjourney and Black Forest Labs for generation inside Meta AI. Muse Image brings that layer in-house, so Meta controls quality, cost, and product timing. Consumers get free access through Meta AI, WhatsApp chats, and Instagram Stories. Power users need Meta One or another monthly plan when free limits run out. Users can start from prompts, add photos, annotate edits, or sketch changes directly. Meta says it can erase photobombers, create QR codes, and keep visual text readable. Advertisers get variants through Advantage+ creative, with edits, style swaps, and brand-matched versions. Meta says internal tests trail GPT Image 2 but beat Nano Banana 2 on editing. This is another move in Meta's effort in trying to convert AI infrastructure spending into revenue beyond social ads.show more

Rohan Paul
20,377 views • 2 months ago