Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

SAM3 video tracking is so good yesterday: collect data, train custom object detector, use tracker to estimate object motion - days today: track anything with text prompt - seconds

708,222 görüntüleme • 8 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

LinkedIn wrapped is here! And it's made with Rive Stoked to be invited by BUCK to join the project. I wasn’t actually animating the thing though I had more of a technical assignment and we had quite a few things to solve. There is 9 chapters total (but not everyone sees all the chapters or even all the slides in given chapter) in 3 languages, the data being sent to Rive is raw so we had to cover for many variables: - A progress bar aware what chapters are included for given user - A progress bar checking what slides are included for and logging their duration time to progress bar - A custom slider solution (the default Rive slider is not enough if you turn contents on and off) - A State Machine logic aware of chapters or slides skipped from playback - A custom slider behavior for when there is auto-slide (instant) and tap-slide (slide animation) - A custom text-length tracker for to dynamically change font-size for long strings - A custom screen ratio tracker to scale down the UI for small screens - Handling dynamically handle big numbers (10,000 becomes 10k etc.) - Handling translations (3 languages) - Super complex behaviour for when if some data-point is missing we don’t leave a blank space, but it is being replaced with other data point instead (so some user se A-B-C, but some will see only B-C as if A never existed without leaving blank space) - Rage click prevention Amazing experience, thanks for having me. People at BUCK are absolutely goated, kudos to everyone who was helping me and to everyone who was actually doing the animations, stunning. Is it the biggest Rive exposure yet? Maybe Rive

Bartek Radziejewski

24,978 görüntüleme • 7 ay önce

If you need McKinsey-style slides, try this prompt: I asked Kimi to conduct a comprehensive analysis of the GenAI video model market, focusing on leading players (e.g., Seedance 2.0, Sora, Kling, Veo, Luma). My prompt: Conduct a comprehensive analysis of the GenAI video model market, focusing on leading players (e.g., Seedance 2.0, Kling, Veo, Luma). Compare their core architectures, temporal consistency, and prompt adherence to identify current industry benchmarks. Use the latest information Requirement: A professional, high-density consulting presentation slide, designed in the style of a top-tier strategy firm (McKinsey/BCG) blended with high-end editorial aesthetics. Core Content & Layout: 1. Rich Data Visualization: The slide is populated with complex, precise charts (stacked bar charts, waterfall charts, or line graphs) and detailed data tables with rows and columns. 2. Structured Frameworks: Includes strategic diagrams or 2x2 matrices constructed with thin, clean lines. 3. High Information Density: The layout is sophisticated and multi-column, mimicking an actual business analysis deck, not just an empty cover page. Visual Style: 1. Aesthetic: Tech-minimalist but information-heavy. Clean, sharp, and authoritative. 2. Typography: Serif fonts (like Times New Roman) for the main headlines to give a premium financial report feel; clean Sans-serif for chart labels and data numbers. 3. Color Palette: Clean white background. Text is sharp black. Charts and graphical accents use Deep Royal Blue and distinct shades of grey for data hierarchy. 4. Graphics: Use fine hairline borders for tables and precise vector lines for graphs.

Crystal

704,925 görüntüleme • 5 ay önce

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 görüntüleme • 11 ay önce