Loading video...

Video Failed to Load

Go Home

SLAM just got a serious speed boost. Efficient LoFTR is now integrated into the Hugging Face Transformers library. It’s 2.5× faster than the original LoFTR and can even outperform the SuperPoint + LightGlue pipeline. Image matching finds correspondences between two images taken from different angles, lighting, or scales. It’s...

26,154 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

A team tested Pi0, Pi0 Fast, Gr00t, and ACT on real robot arms in manufacturing tasks. (🔖 Bookmark this for later!) The task was precise: place thin rectangular frames from a messy stack into a holder. The team fine-tuned each model on 100 real trajectories and compared training time, inference speed, motion quality, and success rates. ⬇️ Here’s a breakdown of what they found Pi0 (Original) ✅ Strongest overall performance in precise pick-and-place ✅ High success rate even in edge cases ✅ Longest training time (~11 hours, ~$30 per run) ✅ Inference time of 80 ms causes short pauses between actions Despite delays, it handles complex scenarios well… solid for high-precision tasks, but slow to train. Gr00t ✅ Trains fast (~2 hours, ~$5 per run) ✅ Performs almost as well as Pi0 on large-object tasks ✅ Struggles with fine precision; random movement in some trials ✅ More training didn’t fix jitter or random offsets Best suited for tasks where exact precision isn’t critical. Not ready for manufacturing-grade accuracy without more tuning. Pi0 Fast ✅ Promised faster training, but results were underwhelming ✅ Training at 6 hours still showed low success rates ✅ Inference was slower than expected ✅ Not reliable for generalizing even slightly new tasks Currently too unstable for real-world deployment. Doesn’t live up to the “Fast” name yet. ACT (Baseline) ✅ 200MB model—lightweight, but limited ✅ Struggles with stacked objects or ambiguous scenes ✅ Success rates around 70% in best-case setups ✅ Can’t match newer models on precision or generalization Still a solid baseline, but clearly a generation behind in robustness. 🚨 Extra Notes All newer models share a common issue: •Inference takes longer than a frame (80 ms vs 33 ms), so robots “pause” between chunks. •This results in jittery movements, but not a dealbreaker unless tasks are time-sensitive. Language-conditioned tasks also fell short: after training on two labeled tasks, the model couldn’t generalize to a third unseen combination using only text prompts. ✅ The good news? These models adapt well to new robot arms with quick fine-tuning. ❌ The bad news? There’s still no plug-and-play solution for improving performance after deployment. Reinforcement learning or DAgger-style data collection during real-world operation may be the next big step, something many teams in robotics are actively working on.

Ilir Aliu

21,844 views • 1 year ago

Wow. Recreating the Shawshank Redemption prison in 3D from a single video, in real time (!) Just read the MASt3R-SLAM paper and it's pretty neat. These folks basically built a real-time dense SLAM system on top of MASt3R, which is a transformer-based neural network that can do 3d reconstruction and localization from uncalibrated image pairs. The cool part is they don't need a fixed camera model -- it just works with arbitrary cameras -- think different focal lengths, sensor sizes, even handling zooming in video (FMV drone video anyone?!). If you've done photogrammetry or played with NeRFs you know that is a HUGE deal. They've solved some tricky problems like efficient point matching and tracking, plus they've figured out how to fuse point clouds and handle loop closures in real-time. Their system runs at about 15 FPS on a 4090 and produces both camera poses and dense geometry. When they know the camera calibration, they get SOTA results across several benchmarks, but even without calibration, they still perform well. What's interesting is the approach -- most recent SLAM work has built on DROID-SLAM's architecture, but these folks went a different direction by leveraging a strong 3D reconstruction prior. Seems to give them more coherent geometry, which makes sense since that's what MASt3R was designed for. For anyone who cares about monocular SLAM and 3D reconstruction, this feels like a significant step toward plug-and-play dense SLAM without calibration headaches -- perfect for drones, robots, AR/VR -- the works!

Bilawal Sidhu

704,124 views • 1 year ago

🔥 Pi Network Has Officially Entered Beast Mode – Global Payment Giants Now Support It! While others doubted, we built. While the world watched, Pi Network quietly integrated with the biggest financial platforms on Earth. And now… we’re ready. 💪🌎 💥 The Walls Are Down. Pi Is Borderless. The Open Mainnet is live, the infrastructure is in place, and now the final pieces are falling into place — Pi Network is supported by a massive alliance of global payment powerhouses, including: 🛡️ Exchange Titans & Fiat Gateways: •✅ Binance P2P •✅ Binance Connect •✅ Transak •✅ Sardine •✅ Topper •✅ UTORG •✅ Paybis •✅ Onmeta •✅ Onramp Money •✅ •✅ TransFi •✅ DFX •✅ Alchemy Pay •✅ Banxa •✅ BTC Direct •✅ Coinify •✅ MoonPay •✅ Fonbnk •✅ GateConnect •✅ Unlimit •✅ Guardarian •✅ Koywe •✅ LocalRamp •✅ Yellow Card 💳 World-Class Fintechs Now Support Pi: •✅ Stripe •✅ Skrill Crypto •✅ Revolut 🌐 This Is Bigger Than Just Crypto Stripe and Skrill aren’t just crypto services — they’re global payment kings. And now they’re helping Pi bridge traditional finance with the new digital economy. Revolut is a top fintech unicorn, and it’s already supporting Pi. Binance is the biggest crypto exchange in the world — and it’s ready. Let that sink in. 🔥 🚀 Why This Changes Everything •🌍 Anyone can access Pi, anywhere in the world •🏦 Buy and sell Pi directly with bank cards, local fiat, Apple Pay, and more •💼 Businesses can now prepare to integrate Pi into payments •🧠 Investors now see the foundation for real-world use and massive growth This is what true mass adoption looks like — infrastructure before hype. Utility before listing. Power before price. 🔮 The World’s Not Ready, But We Are. While other coins chased listings, Pi Network built alliances. While others pumped and dumped, Pi built its own economy. And now? The rocket is fully fueled. We’re not waiting for the future — we’re building it. “You don’t have to be first. You just have to be the one who finishes the race prepared. And Pi Network is ready to dominate.” 📢 Pioneers, Stand Tall Your patience, your faith, your mining, your contribution — it’s all paying off. With 40+ global financial platforms now supporting Pi, we’ve gone from an idea… to an unstoppable force. 🎯 Pi Network is not just another crypto project. 🚀 Pi is the people’s currency — and now the world is opening its gates to it. 📣 Share this. Shout it. Post it. Let every Pioneer know — Pi is ready. Let every skeptic watch — we told you so. Let every investor understand — the game has changed. Pi is no longer coming. Pi is here. 🔥 And the world is about to feel it. Pi Network Nicolas Kokkalis Chengdiao Fan ✅✅✅✅🚀🚀🚀🚀🚀🚀🚀🚀

Mr Spock 𝛑

42,315 views • 1 year ago

99% of iPhone users don’t know this: Their phone already has a free music production powerhouse hidden inside. It’s one of the treasures left behind from the Steve Jobs era. Millions of people own an iPhone, iPad, or MacBook — but have never truly opened it. I asked an AI Agent to create a completely new piece of music from scratch. The result surprised me. It wasn’t just generating audio. It could be edited. Sounds, channels, tracks — everything was adjustable. But I didn’t ask AI to generate an MP3. Instead, I asked it to create a: structured, editable, and extendable music project. The model had to understand: 🎹 Main melody 🎻 Harmony relationships 🥁 Rhythm design 🎸 Instrument arrangement 🎼 Multi-track structure The final output was MIDI, which was imported directly into GarageBand. And this reveals a key difference: Traditional AI music tools like SUNO often generate the final audio directly. You get an MP3. It sounds good, but it’s difficult to modify. MIDI is like the Python of the music world. And this is exactly where Ling-3.0-flash Ant Ling shines. It works like a high-speed execution engine in an AI workflow: ⚡ Low latency 💰 Cost-efficient 🔧 Stable tool calling It can transform complex plans into real outputs quickly. In this music creation workflow, it can independently handle: •Which instruments should play •Which notes should be combined •How rhythms should evolve •How different tracks should interact After importing into GarageBand, you can: ✅ Edit the melody ✅ Replace instruments ✅ Adjust tempo ✅ Re-arrange the composition ✅ Add your own creativity I was genuinely impressed by the result. Try it.

实践哥MinLi

139,763 views • 1 month ago

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

413,002 views • 11 months ago

New model: your robot can now pack your suitcase 🧳 Xiaomi has released a new robot foundation model. Called Xiaomi-Robotics-1, it is designed to have a robot pick things up and move them around. But first, DEFINITIONS: - Mixture-of-Transformers (MoT): An architecture where separate transformer "experts" (e.g., one for vision-language, one for actions) share a single attention stream, so each modality gets specialized parameters without losing joint reasoning. - Vision-language model (VLM): A model that jointly understands images and text. - Diffusion transformer: A transformer trained to turn noise into structured outputs by iterative denoising, here generating robot actions rather than images. - Action chunks: Short sequences of future actions (e.g., the next ~50 motor commands) predicted in one shot instead of one step at a time. - Flow matching: A faster version of diffusion. The model learns a straight-line velocity field from noise to the target action, so it needs only a few integration steps instead of many denoising ones. Its peculiarity comes from its two stage training: 1. 100,000 hours of video shot through a UMI rig: a handheld 3D-printed gripper with a camera, worn by humans doing ordinary tasks in homes, shops, factories and offices. 2. Adapt to actual robot bodies with ~10,000 hours of real-robot data. It replaces the standard approach of teleoperating a real robot for every hour of training data. Its architecture is a Mixture-of-Transformers pairing a pre-trained Qwen3-VL vision-language model with a diffusion transformer that emits action chunks via flow matching, released in 2.6B, 5.1B and 10.5B parameter variants. However, if you read the entire paper ("Scaling VLA Models with over 100K Hours"), you realize that all of the scaling experiments on 20k hours. Therefore the headline "out-of-the-box success climbing 26% → 75% as pre-training data grows" tops out at 100% of 20k hours! What the full corpus does to that curve is never shown -> and this where things would become interesting! Xiaomi's own conclusion is that model size has stopped mattering and data is the binding constraint. The performance gap among different model sizes are less pronounced than those observed across different data scales. This result suggests that model capacity at the billions-parameter scale may already be sufficient to capture the current dataset's distribution. Which further asks the same question: why not use the 100k video hours? Anyway, I would definitely love to have a couple robots at home that can cooperate to pack my suitcase with items relevant to my next destination:

Léo

15,662 views • 26 days ago

I’m thrilled to announce that we just released GraspGen, a multi-year project we have been cooking at NVIDIA Robotics 🚀 GraspGen: A Diffusion-Based Framework for 6-DOF Grasping Grasping is a foundational challenge in robotics 🤖 — whether for industrial picking or general-purpose humanoids. VLA + real data collection is all the rage now but is expensive and scales poorly for this task. For every new gripper and/or scene, you’ll have to recollect the dataset in this paradigm for the best perf. 💡Key Idea: Since grasping is such a well-defined task in simulation - why can’t we just scale synthetic data generation and train a generative model for grasping? By embracing modularity and standardized grasp formats, we can make this a turnkey technology that works zero-shot for multiple settings. GraspGen is a modular framework for diffusion-based 6-DOF grasp generation that scales across embodiment types, observability conditions, clutter, task complexity. Key Features: ✅ Multi-embodiment support: suction, parallel-jaw, and multi-fingered grippers ✅ Generalization to partial + complete 3D point clouds ✅ Generalization to single-objects + cluttered scenes ✅ Modular design uses other robotics modules and foundation models (SAM2, cuRobo, FoundationStereo, FoundationPose). This allows GraspGen to focus on only one thing - grasp generation ✅ Training recipe: grasp discriminator is trained with On-Generator data from the diffusion model - so that it learns to correct the mistakes (if any) of the diffusion generator ✅ Real-time performance (~20 Hz) before any GPU acceleration; low memory footprint 📊 Results: • SOTA on the FetchBench [Han et al. CoRL 2024] benchmark • Zero-shot sim-to-real transfer on unknown objects and cluttered scenes • Dataset of 53M simulated grasps across 8K objects from Objaverse 📄 arXiv: 🌐 Website: 💻 Code: A huge thank you to everyone involved in this journey — excited to see what the community builds on top of it! Joint work with Clemens Eppner , Balakumar Sundaralingam , Yu-Wei, Jun Yamada Wentao Yuan and other collaborators #robotics #diffusionmodels #physicalAI #simtoreal

Adithya Murali

24,106 views • 1 year ago

Private transactions between wallets now possible on Solana using ZERAs Private Cash Addresses. A closer look at last week’s MVP drop: Private P2P. Most "privacy" on Solana still falls back to withdraw-to-address (recipient + metadata leaks) or multi-step send flows that leave trails. Our P2P is different: one atomic in-pool transaction, the sender’s note is nullified, and two new encrypted notes are created, all inside the vault. No stepping out. No re-deposit choreography. Why it’s so hard to deanonymize: ✅ Your Private Cash Address has zero link to your Solana wallet: it’s an X25519 keypair derived from a wallet signature, and the public key becomes your private address. ✅ Notes are encrypted with NaCl box using a fresh ephemeral key every send, meaning even two payments to the same person looks unrelated. ✅ Recipients discover incoming notes via trial decryption, the chain never learns "who owns what." While we have immense respect for those building privacy on Solana, we are proud to be the first to achieve one-shot, in-pool P2P with full privacy. To our knowledge, this remains an industry first. This is the difference between hiding balances inside a pool and moving value privately between people. Note: Right now, the sender’s wallet still signs the private transaction, so on-chain you can see that a wallet performed a P2P action, but the amount and recipient remain private. The upcoming P2P Relayer removes that footprint too, completing the "cash-style" flow with no on-chain sender trace. What’s coming next: ✅ More assets: At least SOL + ZERA alongside USDC. ✅ P2P Relayer: A decentralized signing/relaying network that can submit transactions for you — enabling withdrawals with minimal linkage after the initial deposit, and removing the sender’s on-chain P2P footprint. Try it out in the ZERA Dashboard below ⤵️

ZERA

36,550 views • 5 months ago

OpenClaw has 186K GitHub stars and 1.5M compromised API keys. I needed a secure alternative. So, I built it with n8n and Claude Opus 4.6. It can already: - Reply to your Telegram messages - Access selected folders from your laptop - Access Gmail, Drive, Notion, Linear, etc. - Install new local tools in a sandbox - Run autonomously for hours - Create multiple subagents - Learn from experience - Wake up regularly But, unlike OpenClaw, it: - Can't access your API keys - Can't modify its environment - Can't access folders you haven't shared - Can't access tools you haven't approved - Must get your confirmation, e.g., when sending emails These aren’t prompt instructions. They’re hard architectural boundaries — Docker isolation, mounted folder permissions, n8n’s tool approval system. Key components: ✅ The VPS on Hostinger hosts n8n and a sandbox container. Agents can also connect to my laptop's sandbox via a Claudeflare tunnel + Desktop Commander MCP. ✅ The Manager agent is the brain. It plans, decides, delegates, and talks to the user. It never touches files. It never runs scripts. It works entirely from executor summaries. ✅ The Executor agents are the hands. Each receives a task (what to do + why it matters), decides how to execute it, and reports back. They can install new tools and execute code only in their dedicated sandboxes. ✅ Data Tables in n8n store both memories and sessions — no external database, no vector store, no infrastructure. Just rows in a table. Turns out, that's enough. Two memory types: - Manager memory: user preferences, facts, corrections, relationship, skills, context - Executor memory: what tools are installed, what’s broken, workarounds ✅ Sessions are short-term state for multi-step tasks. Original request, plan, assumptions, and what happened so far. When the Manager loops with fresh context, the session is all it gets. That's a Ralph Wiggum loop. I've been using it for 5 days. And already can't imagine not having it on my phone. What's next: - Heartbeat via Cron (a scheduled prompt) - Civic Nexus governance + MCPs - Supermemory integration - WhatsApp as an additional surface - Hardening The architecture supports all of it. OpenClaw proved people want personal AI agents. It also proved that 'just trust the prompt' isn't a security model. Docker isolation, mounted folder permissions, tool approval — none of this is new technology. It's just discipline. You can easily do this even with n8n — no coding required. --- Want to try it or read more? More, what I learned, and a setup guide: productcompass[.]pm

Paweł Huryn

54,038 views • 6 months ago

I started using Blender through MCP about two weeks ago, and I quickly realized that you can build almost anything with AI. This model was created using Blender, Hunyuan3D, Gemini, and ChatGPT. Here’s how I did it: I opened Gemini, uploaded an image of the Gundam model, and asked it to generate a clean front-view image. I uploaded that front view to ChatGPT and asked it to generate two additional angles: a back view and a 45-degree front-left view. I went to: I signed up with my email and translated the page into English. Then I opened Image to 3D and selected the multi-image option. You’ll see a diagram of a whale from several angles. Upload each reference image in its corresponding position, such as front, 45-degree front-left, and back. I selected the 1.5M-face option. This produces a very high-poly model, but don’t worry, we’ll fix that next. Once the generation is complete, download the model as a GLB file. From the Hunyuan homepage, open 3D Studio using one of the dropdown menus. Select the topology or retopology tool and upload your GLB. I chose the High setting to preserve as much detail as possible. After a few seconds, the model was retopologized. It kept most of its visual detail while using far fewer polygons. The original head and rifle didn’t look very good, so I generated them separately. I returned to ChatGPT and created dedicated reference images for the Gundam’s head and rifle. I generated each part individually in Hunyuan at the 1.5M setting, then ran both through the same retopology process. Next came the textures. Open the texture section, select the multi-image option, and upload the same reference images according to the whale orientation indicators. However, instead of using the original clean textures, I asked ChatGPT to recreate them with wear, rust stains, scratches, and other surface damage. This gave the Gundam a much older and more authentic appearance. Once every part was textured, I imported everything into Blender. You can ask Codex through MCP to remove the original head and rifle, or you can do it manually. Select the main model and press Tab to enter Edit Mode. Press 3 to enable face selection, then press C to activate Circle Select. Paint over the faces belonging to the original helmet or rifle. You can press X to delete those faces or P to separate them into another object. Then position the newly generated, more detailed head and rifle in their place. I demonstrate this process in one of my older videos. And that’s it. You now have a very cool 3D Gundam model! Afterward, I created the cockpit, separated the model into movable sections, and rigged everything for use in my Three.js game. That process deserves its own tutorial, though. If anyone wants to see it, let me know. Or just ask your AI, I guess. They seem to know everything these days. xD

Spectro

79,661 views • 1 month ago

The Future of Commerce Is Being Rewritten Right Now with Interlink and LinkersMap 🚀 Digital commerce is changing rapidly, and the connection between crypto payments, real-world businesses, and global consumers is becoming more important than ever. 🌐 InterLink Labs 👤 + 🌐 is taking a major step forward by expanding crypto payment utility across 50+ global brands, helping bring digital assets closer to everyday use. This is an important move towards practical adoption, where crypto is not only held or traded, but also used in real economic activity. 💳✨ From shopping 🛍️ and entertainment 🎮 to travel ✈️, technology 💻, hospitality 🏨, education 🎓, services 🧰, and global retail 🌍, digital payments are becoming part of the modern customer experience. The future of commerce is not just online. It is global, connected, mobile, and increasingly powered by digital assets. 🔥 💜 LinkersMap: Building Real-World Utility for the InterLink Ecosystem LinkersMap is emerging as a powerful marketplace and geo-commerce platform designed to connect businesses, entrepreneurs, communities, and customers in one trusted digital environment. It is built to support verified InterLink Stores, product listings, payment points, and real-world services, giving users a simple way to discover where digital commerce is active. For businesses, LinkersMap is more than just a listing platform. It is a gateway to greater visibility, stronger customer access, and future-ready participation in the growing digital economy. 🏪🌍 🏪 What LinkersMap Offers to Businesses Businesses can use LinkersMap to create a stronger digital presence and connect with customers beyond traditional local limits. Through the platform, businesses can: ✅ Register and display their store on the map ✅ Showcase products and services to a wider audience ✅ Improve brand visibility within the InterLink community ✅ Reach both local and global customers ✅ Build trust through verified store listings ✅ Position themselves for future crypto payment adoption ✅ Connect directly with users looking for real-world businesses ✅ Participate in a growing digital commerce network Whether it is a restaurant, retail shop, hotel, service provider, education centre, wellness brand, tech business, or Web3 project, LinkersMap gives businesses a practical way to become more discoverable. 📍✨ 🌍 Why This Matters for Real-World Adoption One of the biggest challenges in crypto has always been real-world utility. 💵 People need places to spend. 👪 Businesses need customers. 🏪 Communities need trusted platforms. 💳 Ecosystems need practical use cases. This is where LinkersMap plays an important role. By connecting physical businesses, digital listings, products, and payment visibility, LinkersMap helps bridge the gap between crypto users and real-world commerce. 💳🏪 As the InterLink Network grows, more businesses can join, more users can discover services, and more value can move through the ecosystem. This creates a stronger cycle of adoption: 🔁 More businesses join 🔁 More users discover them 🔁 More transactions become possible 🔁 More utility is created 🔁 The ecosystem becomes stronger 💳 Strengthening Utility for $ITL The expansion of InterLink payment utility helps support the long-term usefulness of $ITL within real-world commerce. As more stores, brands, and services become connected to the ecosystem, $ITL gains more practical relevance. This is important because strong digital assets are supported not only by community interest, but also by actual use. LinkersMap helps make this utility more visible by giving users a place to discover businesses, products, and services connected to the InterLink Network. 🌐💜 ✨ Key Benefits for the Community The growth of LinkersMap creates value for different parts of the ecosystem. For Businesses 🏪 ✅ Increased online and map-based visibility ✅ Access to a growing Web3-focused community ✅ Opportunity to attract new customers ✅ Stronger digital presence ✅ Future-ready payment positioning ✅ Better exposure for products and services For Customers 🛒 ✅ Easier discovery of verified businesses ✅ Access to InterLink-connected stores and services ✅ More real-world use cases for digital assets ✅ Simple map-based search experience ✅ Greater confidence through store verification For the InterLink Ecosystem 🌐 ✅ Stronger real-world utility for $ITL ✅ More payment points and business listings ✅ Better connection between online and offline commerce ✅ Community-driven growth ✅ Increased merchant adoption ✅ A clearer path towards mainstream use 🚀 A New Chapter for Digital Commerce The next generation of commerce is not only about payments. It is about creating a complete ecosystem where businesses can be discovered, customers can connect, communities can grow, and digital assets can be used in practical ways. LinkersMap supports this vision by making real-world business participation easier, more visible, and more accessible. It gives entrepreneurs, merchants, and service providers a platform to be part of the digital economy while helping users find businesses that are ready for the future. 🌎✨ 💼 Register Your Business Today If you are a business owner, entrepreneur, merchant, or service provider, now is the time to join the growing LinkersMap ecosystem. 📍 Add your store 🛍️ Showcase your products 🌍 Reach more customers 💳 Prepare for crypto-powered commerce 💜 Connect with the InterLink community 🌐 Register here: 💜 One Ecosystem. Unlimited Opportunities. 🌎 Global Reach. Real Utility. Future Growth. 🚀 LinkersMap is helping connect businesses, communities, and digital commerce for the future. LinkersMap Dr Altcoin ✝️ Arif Ahmed Core Ambassador | Interlink Labs InterLink Labs 👤 + 🌐 KV InterLink Foundation ITLX Wallet #Interlink #ITLG #ITL #Linkersmap #business #ecommerce #globalgrowth #onlinestore

Tekkaus® | InterLink • MOD • T2 Community Builder

10,912 views • 2 months ago

AI TENNIS ANALYSIS. A FULL COMPUTER VISION SYSTEM. BUILT ON YOLO, PYTORCH, AND KEYPOINT EXTRACTION. Take any tennis match broadcast, any camera angle, any resolution. Feed it into the pipeline. YOLO detects both players and the tennis ball frame by frame. No manual labeling, no pre-annotated dataset. A fine-tuned YOLOv5 model trained on a Roboflow tennis ball dataset handles the ball - the hardest object to track in any sport. Tiny, fast, constantly occluded. The model finds it anyway. Trackers maintain identity across frames so Player 1 stays Player 1 from the first serve to match point. But detection is just the start. A ResNet50 CNN trained in PyTorch predicts court keypoints from every frame - the corners, service lines, baselines, net posts. Fourteen points that define the entire playing surface geometry. From those keypoints the system builds a homography matrix and warps the broadcast perspective into a top-down mini court with real coordinates. Now every player has a position in real space, not pixel space. Every frame becomes a measurement. Every rally becomes a dataset. Player movement speed - calculated from position deltas between frames, converted to meters per second through the homography. Ball shot speed - measured from the ball trajectory across consecutive detections. Number of shots per rally - counted automatically through ball direction changes. All of this rendered live on the video as an overlay. A mini court in the corner showing both players as dots moving in real time. Stats updating after every point. OpenCV handles the rendering. Pandas handles the math. PyTorch handles the intelligence. YOLO handles the eyes. No Hawkeye subscription, no court-embedded sensors, no tracking chips in the ball. A Python script, a trained model, and a GPU. The full code is on GitHub. The tutorial walks through every module - from ball detector training to court keypoint extraction to the final statistical overlay. Professional teams used to need broadcast deals and proprietary hardware for this kind of analysis. Now you build it in an afternoon with open-source tools. Trading here: Computer vision didn't just enter tennis. It made the expensive stuff free.

zostaff

120,370 views • 4 months ago

NEW RESEARCH: You can now create a new robot optimized for any given task! I love this new project by Huy Ha, Shuran Song, and others. Called "Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design", it generates a robot's physical design and its controller together from a task spec. DEFINITIONS: - Reward function: A scoring rule that assigns a number to how well a behavior achieves the task. Here, it is the objective the generated design is pushed to maximize (e.g., track the target motion with low error). - Tokenizing: dividing continuous or structured data (a robot's links, joints, motor specs, states, actions) into a discrete vocabulary of symbols a transformer can process, the same step that turned pixels and audio into "language" for these models. - Diffusion transformer (DiT): A transformer trained to turn random noise into structured output through iterative denoising. Here, it generates robot bodies and trajectories instead of images. - MuJoCo: The standard fast physics simulator for robotics research (DeepMind-maintained). The Menagerie is its curated zoo of ready-to-use robot models. - CMA-ES: Covariance Matrix Adaptation Evolution Strategy, the workhorse black-box optimizer: it evolves a population of candidate designs, keeps the best, and needs thousands of simulator rollouts. - Bimanual multi-trajectory optimization: Finding one design/controller that performs well across several target motions for a two-armed robot at once, harder than optimizing for a single arm and a single motion. - BERT/MAE masked-modeling trick: Train one model to fill in whatever parts of the input you hide (words for BERT, image patches for MAE); at inference, choosing what to mask chooses the task, so masking the body makes it a designer and masking the actions makes it a controller. In practice, you give it a target end-effector motion and a reward function, and it outputs a complete embodiment (link, joint, motor, and inertial property), as well as a controller to drive it. It works by tokenizing both the body (links/joints/motors) and the dynamics (states/actions) into a compact scheme called RoboTokens, training a diffusion transformer (DiT) over them. The same model predicts dynamics using those predictions ("Dynamics Self-Guidance") to push generated designs toward higher reward at inference time. Masking different token types (using the BERT/MAE masked-modeling trick) lets the one model do three jobs: generate an embodiment, control an arbitrary embodiment, or design one conditioned on a motion. It is trained on 11 robots from the MuJoCo Menagerie (0.65 kg hand to 67.5 kg quadruped, 6–35 joints), and validated in sim and on a physical ALOHA doing cloth flinging. I like the fact that this approach inverts the entire recent robotics ideas: designing a policy for a fixed robot -> designing the robot for a fixed task. Every other approach assumes the body is given and learns a controller. Transformer Transformer takes the task (target motion + reward), then generates the body and controller jointly. In practice, it is a ~180× speedup over the standard optimizer at equal-or-better quality. It reaches "CMA-ES-level quality in seconds" and finishes bimanual multi-trajectory optimization in that is worth underlining nowadays! Also worth mentioning: this is the lab behind UMI and Handroid, that I mentioned here previously! The team seems extremely creative, i love these out-of-the-box approaches. Enjoy watching the demo of robot optimization in 3D, data acquisition, then real-life testing:

Léo

26,186 views • 25 days ago

If you are running local LLMs without N-gram speculative decoding, you are wasting massive amounts of compute. Whether your AI is editing a document, outputting structured JSON, or rewriting boilerplate templates, a huge chunk of the text it generates is highly repetitive or already exists right there in the prompt. Standard decoding wastes expensive GPU compute cycles "re thinking" every single token. By adding one hidden flag in llama.cpp, you can instantly fast forward through the repetition. Zero draft models. Zero extra VRAM. And virtually zero compute overhead. Google Colab hands you an enterprise grade NVIDIA Tesla T4 GPU with 16GB of VRAM for free. It’s the perfect Ubuntu Linux sandbox to build a bleeding edge inference engine from scratch. Recently, I showed you how to double your local speeds using MTP (Multi Token Prediction). But MTP requires a secondary neural network draft model. That eats into your precious VRAM (slightly though) and burns extra compute for every guess it makes. N-gram Speculative Decoding gives you a massive speed boost for exactly 0 memory cost and minimal compute. And it's faster than MTP when it works. Here is how it actually works under the hood: Standard autoregressive decoding is slow because it predicts one token at a time. If you ask an agent to format a long JSON object or update one line in an HTML file, it runs heavy matrix multiplications to calculate the probability of every single bracket, space, and letter from scratch. N-gram changes the game. It acts as a lightweight caching system. Instead of running heavy neural network math to guess the next word, it uses a simple hash table. Whenever the LLM starts outputting a sequence of tokens that already exists anywhere in its context window, N-gram instantly recognizes the pattern. Because it is just doing lightning fast string matching, the compute cost is practically zero. It "fast forwards" through the text, drafting the boilerplate instantly from memory, and the main model just verifies it in parallel. Pure speed. Using quantized GGUFs from Unsloth via HuggingFace, I spun up DeepMind’s massive Gemma 4 26B A4B QAT MoE on a free Colab instance to test this. Just look at the raw benchmark data on code editing task: Without N-gram: [ Prompt: 638.6 t/s | Generation: 45.9 t/s ] With N-gram: [ Prompt: 601.9 t/s | Generation: 107.1 t/s ] Here is the exact llama.cpp CLI command to activate it. Notice we don't even need the --model-draft flag: ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -cnv -n 6000 -c 12000 -ngl 99 -fa on --spec-type ngram-mod Stop waiting for your GPU to re calculate words it already knows. I’ve built a free, interactive, cell by cell Google Colab notebook that lets you test this live in your browser. You can literally chat with the model and watch the text generation speed absolutely fly on the second turn when you ask it to edit a file. There are additional parameters for ngram-mod that you can tune once you get it working with the single flag. Link to the free Colab Notebook is in the comments below. It walks you through the entire stack: pulling pre built llama.cpp CUDA binaries for Linux, fetching GGUFs from HuggingFace, and spinning up the inference engine with ngram-mod from scratch. Let me know if you have already tried ngram-mod

Alok

31,765 views • 1 month ago