Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Excellent new fine-grained tracking from DeepMind: TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement arxiv: project: tldr: TapNet for localization then PIPs-style refinement; outperforms everything!

203,971 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Adam W. Harley
Adam W. Harleyvor 3 Jahren

TAPIR not only outperforms PIPs and TAP-Net by a wide margin, but also beats the concurrent (and beautiful) "Tracking Everything Everywhere All at Once" (aka Omnimotion). This shows the power of (1) a well-designed model and (2) large-scale training on synthetic data.

Profilbild von Jared Schnelle
Jared Schnellevor 3 Jahren

I’m very excited to show this to my wife who is an equine vet. There we every expensive systems that can help find lameness in a horse’s gait, but most of them just tape cards on the body and watch for asymmetry. Super cool!

Profilbild von zak
zakvor 3 Jahren

@samjstudios

Profilbild von The A - Z of
The A - Z ofvor 3 Jahren

This would be amazing to analyse opponents in sports 👌

Profilbild von Ramon Sanromà Aragonés
Ramon Sanromà Aragonésvor 3 Jahren

🤔 That will help a lot in sports! To detect anomaly trajectories, muscles... Imagine in a martial arts combat plus eye tracking!

Profilbild von Dr Sly
Dr Slyvor 3 Jahren

May I ask how this supersedes or complement classical methods of optical flow analysis in computer vision, which have been around for decades and use a bajillion times less parameters? Like, with OpenCV? Honest question, because if there is a value added, I'd sure like to know.

Profilbild von Adam W. Harley
Adam W. Harleyvor 3 Jahren

The hope here is to track through occlusions, which flow cannot do. Notice the rhino example, where the trajectories follow the rhino behind the tree.

Profilbild von Mahmoud Mohajer
Mahmoud Mohajervor 3 Jahren

I should say it's a good step to teach AI to sense the motion, and based on that, AI predict where that object is moving. The above feature will allow driverless cars to have better awareness.

Profilbild von Mikko Rantalainen
Mikko Rantalainenvor 3 Jahren

@bayraitt Look great! Everything else seemed to be already spot-on except the rotating wheels of the car seemed to cause problems. I guess that's expected because the wheels look nearly identical again after rotating just 1/5 or 1/6 or 1/7 of a 360° rotation.

Profilbild von Hemal🦉Naik હેમલ નાયક
Hemal🦉Naik હેમલ નાયકvor 3 Jahren

@AlexHHChan

Ähnliche Videos

Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation paper page: Recent advances in generative modeling have led to promising progress on synthesizing 3D human motion from text, with methods that can generate character animations from short prompts and specified durations. However, using a single text prompt as input lacks the fine-grained control needed by animators, such as composing multiple actions and defining precise durations for parts of the motion. To address this, we introduce the new problem of timeline control for text-driven motion synthesis, which provides an intuitive, yet fine-grained, input interface for users. Instead of a single prompt, users can specify a multi-track timeline of multiple prompts organized in temporal intervals that may overlap. This enables specifying the exact timings of each action and composing multiple actions in sequence or at overlapping intervals. To generate composite animations from a multi-track timeline, we propose a new test-time denoising method. This method can be integrated with any pre-trained motion diffusion model to synthesize realistic motions that accurately reflect the timeline. At every step of denoising, our method processes each timeline interval (text prompt) individually, subsequently aggregating the predictions with consideration for the specific body parts engaged in each action. Experimental comparisons and ablations validate that our method produces realistic motions that respect the semantics and timing of given text prompts.

AK

126,612 Aufrufe • vor 2 Jahren

EllesmereUI's Biggest Patch since Raid Frames is live! Aug Evokers since Blizzard won't give you an active state on your CDM for Ebon Might, I decided to give you one. check out the video! This feature also allows trinkets/pots/racials to get custom active states Players who share profiles: you can now include your full spell layout (which spells go where) and per-spell settings for CDM when exporting profiles! ----------- Full patch notes with new features and bugfixes: **Profiles:** - **NEW:** Profile exports can now carry your entire Cooldown Manager setup - which spells sit on which bars plus every per-spell setting - for the specs you choose, so importing a profile recreates your CDM layout instantly. **CDM:** - **NEW:** Give any trinket, potion, racial, or custom spell an Active State that adds its own glow and color while active, plus a Cooldown State Effect that changes its look based on whether it's ready. - **NEW:** Sync your trinkets, potions, and racials across specs with one button so you only set them up once. - Sound pickers for Focus Cast Sound and per-buff Audio Effect now have a search box. - Pandemic glow's Apply to All now also syncs tracking bars and keeps the same glow style everywhere. - A trinket, potion, or racial already on another bar now auto-moves to the new bar instead of being grayed out. - A buff saved under two spell IDs now shows as a single icon in the buff bar preview. - Cooldowns you remove from Blizzard's Cooldown Manager now disappear from your previews, while trinkets, racials, and custom spells are kept. - Added tracking for the Nightborne racial (Arcane Pulse), which was missing from the racial list. **Tracking Bars:** - **NEW:** Grouped bars now pack together with no blank gaps, always filling the next available slot. - Eclipse (Solar) and Eclipse (Lunar) now each drive their own bar. - Switching specs while the page is open now refreshes the selected bar correctly. - Pandemic glow now fires for Lifebloom on you or a group member. **Resource Bars:** - **NEW:** A new GCD Bar fills over your global cooldown, with full control over size, position, color, and look. - Expand Power Bar if No Resource now also expands when the class resource is toggled off or disabled for the spec. - Fixed the threshold color not showing with Enhance 5 Bar Style. - The cast bar latency overlay now reads live latency, so spell queueing no longer stops it showing. **Raid Frames:** - **NEW:** New healer tools including heal-absorb text, a crowd-control glow on debuffs, and more ways to position and highlight dispellable debuffs. - The Auto Resize toggle is now a dropdown that scales Indicators & Auras and Tracked Buffs independently with frame size. - New Show Over Dispels toggle lifts the heal-absorb overlay above the dispel gradient. - Fixed custom (non-20-player) sizes loading at the wrong position after login. **Unit Frames:** - **NEW:** A non-tank threat border shadows your player frame when you pull or hold aggro, with Has Aggro and Close to Aggro colors. - **NEW:** Player health text gains Heal Absorb Amount and Heal Absorb Short options. - **NEW:** Mini frames and Boss frames gain a per-frame Bar Texture dropdown. - **NEW:** Boss frames gain a Hover Borders control with its own mouseover and target colors. - **NEW:** Independent Spacing X and Spacing Y sliders for buff and debuff icons. - **NEW:** Absorb and heal-absorb style dropdowns now include your installed SharedMedia textures (also on Raid Frames and Nameplates). - A new Hover Borders control lets you turn the mouseover highlight border on or off per frame. - New Show 2 for Boss option adds a second decimal to boss frame health text. - Buff and debuff Offset X/Y sliders now reach plus or minus 1500. - A new Cast Bar Position cog adds an Offset Y slider for the boss cast bar. - Player, target, and focus cast bars now show above other frames instead of behind them. - Cast bar timer text now has room so it no longer cuts off early. **Action Bars:** - **NEW:** A new When Not Dragonriding visibility mode hides a bar while skyriding and shows it the rest of the time. - The When Dragonriding option now also shows the bar in Druid Flight Form. - The stance bar now shows its GCD swipe even for spells that don't change form. - Bars are briefly forced visible while Myslot's window is open to stop a stall during import/export. **Minimap:** - **NEW:** Hovering the calendar button now shows your raid and dungeon lockouts with boss progress, server time, and time until the weekly reset. - The Omnium Folio button no longer goes missing or drifts after a loading screen, and its position and scale now persist. **Mythic+ Timer:** - **NEW:** A new Show Time Remaining toggle adds an MM:SS countdown to the +2/+3 threshold row that reddens as time runs out. - Fixed timer and detail text cutting off after a font swap. **Colors:** - **NEW:** A new Global Colors section lets you share one profile's custom colors across all profiles or give each profile its own. **Localization:** - **NEW:** Full Russian language support. **Auras, Buffs & Consumables:** - Warrior stance reminders now read the stance bar, so each spec is reminded of its correct stance and clears the moment it's active. - The Inky Black Potion reminder now clears after you drink the potion and reappears on cancel, expiry, or death. - The last-used flask, food, and weapon-enchant preference now saves correctly. **Nameplates:** - A new sync icon on the Pandemic Glow Style row applies the nameplate's pandemic glow to all CDM and tracking bars at once. **Chat:** - The Whisper Sound dropdown gained a search box. **Bags:** - Mythic Keystone dungeon abbreviations now split on hyphens (Nexus-Point Xenas shows as NPX) and handle localized names.

Ellesmere

53,738 Aufrufe • vor 2 Monaten

Hello everyone, Here is a personal project that builds on what SkalskiP did with his computer vision project, I tweaked it to be able to run full games handling dead balls, free throws, camera cuts, etc. and offload all the data into a CSV where I store: - Shot locations with closest defender distance and name. - Distance Traveled Total throughout the game. - Velocity Information of players - Max and Average Speed - Time of Possession The cleanup is done post-processing where we fill in players information based on known information from our tracking with the NBA API to get key points to ignore free throws polluting the data along with filling in unknown players. Nothing is hardcoded besides team names, all roster and game information is gathered from reading the scoreboard to find the unique game and then it downloads roster information and provides a safety net for tracking from the NBA API. With being able to process a full game, one implementation I did was use Dynamic Time Warping to run two clips from separate games of similar looking SLOB's to see if it could identify plays which in the below sample has worked. Using this, a team could grab clips of teams running plays or sets and index them, points of attack, personnel, first and second options, frequency and success rate information to build out scouting reports and tendencies. The next step is to create play diagrams automatically based on the clip FastDraw-style for easy scouting insertion to coaching staff and players documents. This is part one of a greater project that will be showcased over the next few days where its incorporated into an all purpose basketball operations artificial intelligence. The clip below is the spliced together vision tracking analyzing both play segments but a full game scan is possible. Would love to hear feedback and see any questions you might have. Again, huge thanks to SkalskiP his resources were great to springboard into something compelling to myself.

Suge Nach

17,329 Aufrufe • vor 3 Monaten

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing abs: paper page: present the content deformation field CoDeF as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis.Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline.We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video.With such a design, CoDeF naturally supports lifting image algorithms for video processing, in the sense that one can apply an image algorithm to the canonical image and effortlessly propagate the outcomes to the entire video with the aid of the temporal deformation field.We experimentally show that CoDeF is able to lift image-to-image translation to video-to-video translation and lift keypoint detection to keypoint tracking without any training.More importantly, thanks to our lifting strategy that deploys the algorithms on only one image, we achieve superior cross-frame consistency in processed videos compared to existing video-to-video translation approaches, and even manage to track non-rigid objects like water and smog.

AK

153,305 Aufrufe • vor 3 Jahren

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,438 Aufrufe • vor 9 Monaten

AI TENNIS ANALYSIS. A FULL COMPUTER VISION SYSTEM. BUILT ON YOLO, PYTORCH, AND KEYPOINT EXTRACTION. Take any tennis match broadcast, any camera angle, any resolution. Feed it into the pipeline. YOLO detects both players and the tennis ball frame by frame. No manual labeling, no pre-annotated dataset. A fine-tuned YOLOv5 model trained on a Roboflow tennis ball dataset handles the ball - the hardest object to track in any sport. Tiny, fast, constantly occluded. The model finds it anyway. Trackers maintain identity across frames so Player 1 stays Player 1 from the first serve to match point. But detection is just the start. A ResNet50 CNN trained in PyTorch predicts court keypoints from every frame - the corners, service lines, baselines, net posts. Fourteen points that define the entire playing surface geometry. From those keypoints the system builds a homography matrix and warps the broadcast perspective into a top-down mini court with real coordinates. Now every player has a position in real space, not pixel space. Every frame becomes a measurement. Every rally becomes a dataset. Player movement speed - calculated from position deltas between frames, converted to meters per second through the homography. Ball shot speed - measured from the ball trajectory across consecutive detections. Number of shots per rally - counted automatically through ball direction changes. All of this rendered live on the video as an overlay. A mini court in the corner showing both players as dots moving in real time. Stats updating after every point. OpenCV handles the rendering. Pandas handles the math. PyTorch handles the intelligence. YOLO handles the eyes. No Hawkeye subscription, no court-embedded sensors, no tracking chips in the ball. A Python script, a trained model, and a GPU. The full code is on GitHub. The tutorial walks through every module - from ball detector training to court keypoint extraction to the final statistical overlay. Professional teams used to need broadcast deals and proprietary hardware for this kind of analysis. Now you build it in an afternoon with open-source tools. Trading here: Computer vision didn't just enter tennis. It made the expensive stuff free.

zostaff

120,370 Aufrufe • vor 4 Monaten

9 repos that mass replace a $150,000/year NBA analytics department. all free. all open source. -> replaces Second Spectrum and SportVU YOLO tracks every player and ball from any broadcast. assigns teams by jersey color. court keypoint detection builds a tactical top-down map. speed, distance, passes - all from a TV feed. no sensors. -> replaces paid NBA prediction services ($100/mo) XGBoost + Neural Net. moneyline and totals. Kelly Criterion sizing. 69% accuracy. pulls odds from FanDuel/DraftKings automatically. the most starred NBA betting repo on GitHub. -> replaces an entire quant sports desk 5-model ensemble: XGBoost + PyTorch MLP + Ridge + Lasso + baseline. Optuna-tuned hyperparameters. SQLite database with box scores, play-by-play, betting lines, injury reports. production-grade. -> replaces manual daily prediction workflows XGBoost/LightGBM with GitHub Actions automation. scrapes new data, retrains models, outputs daily win probabilities. set it and forget it. -> replaces ELO subscription services custom ELO + Ridge + XGBoost + Neural Networks ensemble. full data scraping pipeline. comprehensive visualizations. FiveThirtyEight-style ratings from scratch. -> replaces Four Factors analytics dashboards ELO rating system + Four Factors + PCA dimensionality reduction. detailed comparison of 10+ models. honest 65.3% accuracy - because that's what real NBA prediction looks like. -> replaces computer vision analytics platforms ($500/mo) YOLO player/ball tracking. automatic team assignment. court keypoints. pass and interception detection. speed and distance. full tactical view. modular architecture. -> replaces shot tracking hardware YOLOv8 detects ball and hoop in real-time. linear regression predicts trajectory. registers makes and misses automatically. works on any video feed. -> replaces paid sports data subscriptions ($300/mo) official Python client for NBA. com API. box scores, play-by-play, shot charts, player tracking. 40+ years of data. zero cost. the foundation every NBA ML project is built on. like + bookmark you'll need this when you build your first NBA prediction bot

zostaff

103,649 Aufrufe • vor 4 Monaten

Introducing: A native HyperEVM explorer with analytics and dev tools for CoreWriter, precompiles, HIP-3, HIP-4, and more. HyperCore and the HyperEVM share the same state. This reads both sides of it. Real-Time: every page updates the moment a block lands. Home, blocks, transactions, CoreWriter, any address. You never need to refresh. HyperCore Support: on any transaction that calls CoreWriter, what HyperCore did with the action: executed with the fill, rejected with Core's own reason string, or still pending. Dev Tools: 16 of them. A CoreWriter composer, a read-precompile playground, a dry run of any call with its state diff, and more. HIP-3 and HIP-4: builder DEX indices resolved to real dex names against Hyperliquid's own registry, plus outcomeOp split, merge and negate, settled value, and outcome asset ids. CoreWriter Tracking: every EVM to HyperCore write decoded by action, with the callers ranked and sendAsset broken out by token. 1,057,498 actions since launch. CoreWriter/Precompiles Growth Tracking: the share of HyperEVM transactions that read HyperCore state 10xed in six months. See which precompiles the EVM reads, and which contracts read them. Query and Export: filter any transaction, transfer, internal call or user operation by address, block range and amount, then take the result as CSV. Protocol Analytics: follow a protocol's money, and tell deposits apart from price. Money Flows: capital between protocols, wallets and HyperCore across 580 pairs, netted per pair, with the tokens behind every edge and the counterparties on both ends. Bridge: HYPE and every token crossing in both directions. Liquidations: every lending liquidation decoded, per market, ranked by debt repaid, with the collateral seized. Protocol Growth: month-one retention and cohort triangles per protocol. Concentration: the share of each protocol's transactions sent by its busiest 1% of addresses, which runs from a fifth to nine tenths. Also: name.hl resolution, verified sources, address labels, 4-byte method names, token approvals with one-click revoke, dual-lane gas, and a directory of 59 HyperEVM builder tools. Powered by Quicknode Hyperliquid.

HL Eco

33,779 Aufrufe • vor 1 Monat