SAM3 video tracking is so good yesterday: collect data,... train custom object detector, use tracker to estimate object motion - days today: track anything with text prompt - secondsshow more

SkalskiP
708,630 views • 9 months ago
Adding keyframes to the 3D feature tracking shows that... the object in the Pentagon's PR043 UFO video is consistent with a ~13 cm diameter object moving in level flight at 20 knots. Or any number of faster, larger objects. So maybe a bird, or a balloon.show more

Mick West
11,275 views • 24 days ago
🚨 Meta introduces Segment Anything Model 2 (SAM 2)... — the first unified model for real-time, promptable object segmentation in images & videos. 🚀SAM 2 is available today under Apache 2.0 so that anyone can use it to build their own experiences ➡️show more

Zhengzhong Tu
181,295 views • 2 years ago
PhD Students - How to detect AI text in... your writing? We often use ChatGPT for writing. However, this leads to AI-plagiarized text. This can be problematic in many scenarios. For example, if you use AI text in your papers. Your research paper can get desk rejected. 🍁How to detect if there is AI text in your writing? 1. Go to and log in. 2. Click on 𝐴𝐼 𝑑𝑒𝑡𝑒𝑐𝑡𝑜𝑟 from the left menu 3. Insert your text and click on 𝐴𝑛𝑎𝑙𝑦𝑧𝑒. 4. will generate AI detection report This report shows the following. → Percentage of AI generated text → Options for converting AI text into non-AI text 🍁How good is this AI detector? SciSpace conducted a benchmarking study. In this study, the detection capability was compared with other AI-detectors. SciSpace AI detector was tested with 4000 samples. It showed an accuracy of 96%. This means it can detect AI-generated text with 96% accuracy. The study showed that SciSpace AI detector has outclassed AI detectors like GPTZero, ZeroGPT, and Grammarly. 🔴Anything you'd like to add?show more

Faheem Ullah
13,102 views • 10 months ago
LTX-2.3 is now live on OpenArt. 🎬 The most... capable open video model just got a major upgrade and you can use it right now. What's new in 2.3: → Sharper fine detail: Hair, textures, text, edges. All of it. → Tighter prompt adherence: Complex multi-subject prompts? Handle it. → Stronger image-to-video: less freezing, less Ken Burns drift, more actual motion. → Cleaner audio: fewer artifacts, tighter sync across text-to-video and audio workflows. → Native portrait: up to 1080×1920, trained on vertical data.show more

OpenArt
2,151,122 views • 4 months ago
🔥 AI video creation is finally evolving beyond basic... motion prompts. Try out Depth Map Control right here: Historically, the hardest part of the process has been maintaining scene consistency: • camera movement • object placement • spatial relationships Thanks to PixVerse Depth Map Control, you can now grab depth data from a reference video and merge it with a fresh image to explore highly directed video generation. The final outcome? Much smoother camera tracking. Highly cohesive visual scenes. It is a powerful workflow that hands creators actual authority over how their AI visuals shift and flow over time. 🎬show more

Elara Quinn
12,202 views • 1 month ago
MiniMax H3 is now 50% OFF on Magnific for... 2K video, only until September 1. 🔥 I’ve been trying MiniMax H3 on Magnific, and it feels like a big upgrade for AI video creation. It’s not just about turning text into videos. You can use text, images, videos, and audio together in one prompt, giving you more control over the final video. Here’s what makes it stand out: - Multimodal: Use text, images, video, and audio in one prompt. - Multiple references: Add up to 9 images, 3 videos, and 3 audio files. - 2K video: Create videos up to 15 seconds long. - Built-in sound: Generate voice, music, and sound effects with the video. - Easy editing: Remove objects or transfer motion easily. - More control: Control the camera, characters, and voice. You can use it to turn posters into videos, moodboards into short films, and product images into ads. It also helps bring your ideas to life with realistic movement, lighting, reflections, and sound. The workflow is simple: give it your references → generate → edit → refine. Try MiniMax H3 on Magnific:show more

Markandey Sharma
96,964 views • 18 days ago
Introducing Sora, our text-to-video model. Sora can create videos... of up to 60 seconds featuring highly detailed scenes, complex camera motion, and multiple characters with vibrant emotions. Prompt: “Beautiful, snowy Tokyo city is bustling. The camera moves through the bustling city street, following several people enjoying the beautiful snowy weather and shopping at nearby stalls. Gorgeous sakura petals are flying through the wind along with snowflakes.”show more

OpenAI
98,186,645 views • 2 years ago
🤯 I didn’t expect Depth to track rooftop parkour... this cleanly! Used a fast climbing clip as the motion reference, then swapped in a student with twin tails and a backpack on a coastal school rooftop. The running path, wall contact, jumps, landings, and full-body continuity all hold together surprisingly well. 🌟 Workflow: 1. Pick a reference clip under 15 seconds 2. Convert it to Depth with Depth Anything V2 3. Generate a new character + scene 4. Feed everything into Seedance with the prompt below 🌟 Depth video conversion: You can build a local Depth converter with Codex — prompt in the comments. Depth opens up a lot more possibilities for parkour, climbing, martial arts, dance, and other complex full-body motion. Workflow + Prompt below 👇show more

Larus Canus
24,613 views • 1 month ago
Yesterday someone suggested I create a fight between Alucard... and Drolta Tzuentes using Kling AI 2.6. I usually work with Text to Video, but in this case I had to use Image to Video. The base image is a fusion of two images created in Niji and combined with GPT Image 1.5. Then I added the right prompt to give the fight speed and dynamism. The result isn’t perfect, but it’s pretty cool.show more

OscarAI
18,533 views • 8 months ago
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,992 views • 4 months ago
Gemini Omni's motion control is f*cking cracked i just... figured out how to turn 1 reference video into 50+ AI videos with the exact same movements... you have a video of someone eating, dancing, using a product, doing whatever complex motion you need. you feed it to Gemini Omni and it recreates that exact motion with a completely new AI character in literally one prompt i've tested this against Kling motion control and it's not even close. Kling falls apart the moment you try anything complex. eating scenes look weird, hand movements get mangled, anything multi-step breaks down completely. Gemini Omni handles all of it if you're still using kling motion control or paying creators to split test your videos, this replaces that entire workflow here's the thing though. you can't just prompt this out of the box. if you try to do motion transfer with default prompting you're going to get errors or the motion won't transfer properly. there's a specific prompting method that makes it work every time so i packaged up the whole system.. here's what you're getting: > full step by step video breakdown > how to find the best reference videos to use > the exact prompting system that allows for motion control transfer so you never get errors > the workflow for batching this out at scale (1 video → 50+) RT + reply "MOTION" and i'll send it over (must follow so i can dm)show more

Miko
60,571 views • 1 month ago
Kling 2.6 Motion Capture is so good. It's a... huge leap forward. We've been able to do motion capture with AI for a while, but the quality bump finally made this approach useful. Here are the steps: - Get a reference video - record it or get a stock video - Grab the start frame of that video - Edit it using Nano Banana Pro. Ask to replace character and background. - Select the Kling 2.6 Motion Capture model in the video generator on Freepik (now Magnific) - Upload reference video + set the edited start frame - Video prompt can be simple e.g. "Gandalf dancing"show more

Martin LeBlanc
35,274 views • 8 months ago
gemini omniflash is actually f*cking cracked. you can animate/edit... any video with a text prompt. character swaps, object transforms, full environment changes without regenerating/rotoscoping. everyone using AI to to animate and edit videos right now hits the same wall. the clip comes out 90% right and you regenerate from scratch hoping the 10% fixes itself. it never does. the fix is using your video as the input. omniflash edits what's already there instead of rolling the dice again. here's what's in the system: > the two-layer premiere trick: generate the same shot twice (one with background removed), stack them, cut at one frame, instant scene change > character swap with a single reference image (plus the one line you need or the model keeps the original's features) > object transforms that leave the rest of the frame untouched: stone into glowing sphere, candles into flowers > style transfer from an image reference instead of text, way more accurate > why stacking edits in one prompt breaks everything and the exact step order that doesn't > the audio limitation nobody mentions and how to work around it i packaged every prompt, the edit sequence, and the premiere layering setup. RT + reply "OMNI" and i'll send it over.show more

Sulfur
36,571 views • 2 months ago
Introducing Attio Objects 🚀 We know how hard... it is to find a CRM that fits your unique business model. That's why we built Attio Objects – our powerful data model with custom objects that gives you complete flexibility to structure your CRM exactly how you need it. Along with custom objects, we've also introduced new standard objects: - Workspaces and Users objects for PLG businesses. - A robust Deals object for sales-driven companies. This is the culmination of a 4-year effort, with 3 years of work put in even before launching Attio. Since day one, we've been determined to solve the fundamental problem in the CRM space: the trade-off between power and time-to-value. If you wanted power and flexibility, your CRM would take forever to build and not work well with your stack. If you wanted speed, you'd need to use highly opinionated, inflexible software that doesn't really work for your business. That ends today. With Attio, you no longer have to compromise. Build your CRM your way, fast. Iterate as you grow. High-growth startups like Replicate, , and Modal and more are already using Attio's object architecture to perfectly match their businesses and accelerate their growth. To get all the details, check out our blog post 👇 show more

Attio
26,821 views • 2 years ago
GPT-5.6 Sol is unbelievably good at creating and editing... videos. It can do motion design, product demos, and animations like this one I made by simply giving it a screen recording. GPT 5.6 has the best design taste and significantly outperforms Fable, which relies heavily on repetitive design patterns. To help you experiment with video editing on it, we just launched a collection of 100 ready-to-use skills that show what’s possible and help you get started with video editing using GPT-5.6. These skills can create anything from motion graphics launch videos for your product to a 3B1B-style science explainer video. You can also use them to edit existing videos: add captions, generate motion graphics, create voiceovers, redesign visual styles, translate into new languages, and much more. If you want access to the full library, comment “VIDEO SKILLS” and I’ll share it with you. (You'll have to follow me so I can DM you.)show more

Akash Anand
515,838 views • 1 month ago
This is some quietly impressive work on making video... world models actually controllable in 4D space. VerseCrafter lets you take an input image, use something like Blender to animate the 3D camera path and object trajectories, then uses that to condition generation. Scribbling in 2D feels so crude in comparison. The authors represent everything in a shared 4D world state - static background as a point cloud, moving objects as 3D gaussian trajectories. The gaussians are an interesting choice because they capture position, shape, and orientation probabilistically rather than forcing rigid bounding boxes or category specific models like SMPL-X for human bodies. They bolt this onto frozen Wan2.1 with a lightweight adapter, so they get a strong video prior. They also built a pipeline to auto extract 4D annotations from real world videos to train this puppy. It doesn't look sexy yet, but IMO this is the interface video world models need - actual 3D authoring tools to exert control rather than crude scribbles and prompt incantations.show more

Bilawal Sidhu
26,017 views • 7 months ago
I tried MiniMax Design (H3) to see how it... handles real content creation. The workflow is simple. You just write a prompt or drop in an image, and it turns that into a dynamic video with motion, framing, and scene depth. No timeline to manage. No editing setup. No back and forth. What stood out to me: • Text to video and image to video both feel smooth. • It handles motion, camera angles, and flow on its own. • Output is fast, usually within seconds. • Works well for reels, quick ads, storytelling, and idea testing. It removes the hardest part: starting from scratch and turns your ideas into content in minutes. Instead of thinking, “How do I make this video?” You start with, “What do I want to create?” That shift alone makes it worth exploring. Try it here: #Hailuoshow more

Manish Kumar Shah
27,680 views • 5 months ago
I've been working a lot with SAM3 and the... Momentum Human Rig (MHR). I finally integrated it into the data I'm working with Rerun. The progression I've taken looks as follows SAM3 + SAM3D-body on 1. a single image 2. a set of multiple images 3. a single video 4. A multiview video capture I took inspiration from the SAM3D-body paper and built a multiview fitting optimization pipeline. This pipeline involves using the 2D keypoints from the single-view pipeline, triangulating them, and employing an L1 loss between the 2D/3D keypoints. The temporal stability isn't great, so that's the next portion I'm going to focus on. One really frustrating thing about SAM3D-body is the lack of per-joint confidence values. It makes it harder to deal with occlusions. I'm probably going to need to use a separate model, or maybe add a confidence head.show more

Pablo Vela
42,267 views • 7 months ago
LinkedIn wrapped is here! And it's made with Rive... Stoked to be invited by BUCK to join the project. I wasn’t actually animating the thing though I had more of a technical assignment and we had quite a few things to solve. There is 9 chapters total (but not everyone sees all the chapters or even all the slides in given chapter) in 3 languages, the data being sent to Rive is raw so we had to cover for many variables: - A progress bar aware what chapters are included for given user - A progress bar checking what slides are included for and logging their duration time to progress bar - A custom slider solution (the default Rive slider is not enough if you turn contents on and off) - A State Machine logic aware of chapters or slides skipped from playback - A custom slider behavior for when there is auto-slide (instant) and tap-slide (slide animation) - A custom text-length tracker for to dynamically change font-size for long strings - A custom screen ratio tracker to scale down the UI for small screens - Handling dynamically handle big numbers (10,000 becomes 10k etc.) - Handling translations (3 languages) - Super complex behaviour for when if some data-point is missing we don’t leave a blank space, but it is being replaced with other data point instead (so some user se A-B-C, but some will see only B-C as if A never existed without leaving blank space) - Rage click prevention Amazing experience, thanks for having me. People at BUCK are absolutely goated, kudos to everyone who was helping me and to everyone who was actually doing the animations, stunning. Is it the biggest Rive exposure yet? Maybe Riveshow more

Bartek Radziejewski
24,978 views • 8 months ago