Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

High fidelity visuomotor data from handheld grippers or data capture gloves provide the most valuable robot learning data. This data collection modality achieves the best balance of trade-offs compared to teleop, master–slave devices (e.g., ALOHA), and egocentric cameras. They’re portable, intuitive, low-cost/scalable, and directly transferable to robots. They don’t...

39,082 Aufrufe • vor 7 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

🚀 Introducing EgoExo Forge - built on top of Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tuned

Pablo Vela

32,085 Aufrufe • vor 1 Jahr

🚨 SCIENTISTS JUST BUILT A CHIP THAT CAN SEE, THINK, AND REMEMBER ALL AT THE SAME TIME. And it works more like a biological brain than a traditional computer. Researchers at RMIT University have created a neuromorphic vision chip that mimics the human eye and brain. Unlike conventional systems that capture images and send data to external processors, this chip performs sensing, processing, and memory storage directly where the light hits. The active layer is thousands of times thinner than a human hair. It uses doped indium oxide to detect light, process the information on-chip, and retain what it sees over time without constant electrical refreshing. Why this matters: • It dramatically cuts energy use and latency by eliminating data transfer to separate processors • Enables much faster real-time decision making for autonomous systems • Works more like biological vision than traditional machine vision • Could power the next generation of efficient edge AI in vehicles, robots, and remote sensors The deeper implication: For decades, we’ve built vision systems by bolting cameras, processors, and memory together like separate organs. This chip collapses those functions into one biological-style unit. It’s a step toward machines that don’t just “see” but actually perceive and remember in a more efficient, brain-like way. If scaled successfully, it could become a foundational component for autonomous systems that need to operate intelligently with minimal power and minimal delay. We’re moving from cameras that take pictures to chips that truly see. How do you think neuromorphic vision chips like this will change what’s possible for self-driving cars and autonomous robots? Follow for more frontier neuromorphic computing, AI hardware, and brain-inspired technology.

TheNewPhysics

23,196 Aufrufe • vor 2 Monaten

🚀Just launched: Amazon Q, the most capable GenAI-powered assistant is generally available today: Customers are using Q to transform how their teams get work done. When employees chat with Amazon Q, it provides immediate, relevant information and advice to help streamline tasks, speedup decision-making, and help spark creativity and innovation at work. . Early indications signal Amazon Q could help our customers’ employees become more than 80% more productive at their jobs; and with the new features we’re planning on introducing in the future, we think this will only continue to grow. 🟠 Amazon Q Developer allows developers to spend more time coding and less time on maintenance and performing other tedious, repetitive tasks. Q assists developers and IT professionals (IT pros) with all of their tasks—from coding, testing, and upgrading applications, to troubleshooting, performing security scanning and fixes, and optimizing AWS resources. Q also comes with Q Developer Agents which can autonomously perform range of tasks and we expect it to be the state of the art accuracy in benchmarks like SWE-Bench. 🟠 Amazon Q Business empowers employees to be more data-driven, and helps customers make better, faster decisions using company knowledge and data. Q Business is a generative AI–powered assistant that can answer questions, provide summaries, generate content, and securely complete tasks based on data and information in enterprise systems 🟠 Amazon Q Apps, a new and powerful capability of Amazon Q Business, enables employees to use natural language to quickly and securely build their own generative AI applications to automate daily tasks without requiring any prior coding experience. Employees simply describe the type of app they want, in natural language, and Q Apps will quickly generate an app that accomplishes their desired task, helping them streamline and automate their daily work with ease and efficiency.

Swami Sivasubramanian

25,216 Aufrufe • vor 2 Jahren

A Letter to Our Community: The Road Ahead for Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷

Axis Robotics

27,858 Aufrufe • vor 7 Monaten

This reported breakthrough apparently used a dataset that’s open for researchers. It’s the work of Eddy Xu, a teenager who was one of the first to get in on the video training data gold rush. He dropped out of Columbia last year to launch Build AI, which has raised around $22 million so far. Build AI’s Egocentric-1M dataset reportedly includes 1 million hours of data recorded using the startup’s self-developed devices across factories in Southeast Asia. A lot of it is from India. Xu has said he’s moved his team to Bengaluru, dedicating $10 million to get data from Indian factories. India has become one of the prime locations for collecting this kind of data. While enrolled at Columbia Engineering, Xu went viral in January 2025 after showing Meta Ray-Ban smart glasseshe modified to cheat at chess. The student, then 17, connected the device’s camera to a chess engine that calculated the best move and relayed it in real-time. The tech reached a wider audience thanks to popular streamer and chess master Alex Botez publicly tested them. Before college, the Long Island-raised Xu won DECA’s global business championship and sold an edtech startup that reached more than 178,000 users in 90 days. He also launched a startup called Omega Robotics in middle school, raising about $120,000 to run an independent, coach-free competitive robotics team out of a basement. Xu and co-founder Jonathan Jia, who serves as CTO, moved to San Francisco to build the first recording devices with a small team. They quickly moved operations to Shenzhen to quickly iterate and scale production. Build previously offered smaller datasets with 10,000 and 100,000 on Hugging Face but the 1M dataset requires emailing Xu directly. I’m sure he’s flooded with requests now.

Mike Kalil

13,025 Aufrufe • vor 14 Tagen

𝗗𝗼𝗻'𝘁 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗲 𝗿𝗼𝗯𝗼𝘁 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹𝘀. 𝗦𝘁𝗲𝗲𝗿 𝘁𝗵𝗲𝗺 𝘄𝗶𝘁𝗵 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘀𝘁𝗲𝗮𝗱, 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗯𝗮𝘀𝗲 𝗽𝗼𝗹𝗶𝗰𝘆 Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging. A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe. In our new paper, 𝗙𝗹𝗼𝘄𝗗𝗔𝗴𝗴𝗲𝗿, we take a different approach: 𝗜𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹, 𝘄𝗲 𝗹𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗲𝗲𝗿 𝗶𝘁 𝗳𝗿𝗼𝗺 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀. The key idea is 𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘃𝗲𝗿𝘀𝗶𝗼𝗻: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities. Across simulation and real robots, FlowDAgger: 📈 Learns from only 5–20 human intervention episodes 🏆 Outperforms supervised fine-tuning and latent-space reinforcement learning 🤖 Works across VLAs, diffusion policies, and world-action models ✔️ Provides reliable improvements without modifying the pretrained policy We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections. 📄 Paper: 🌐 Project: 💻 Code: This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and Andrey Kolobov at Microsoft Research and Maya Cakmak at University of Washington

Oier Mees

13,032 Aufrufe • vor 1 Monat

This work makes a humanoid robot do simple parkour moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.

Rohan Paul

37,121 Aufrufe • vor 6 Monaten

We all remember. We all remember when blockchain was pitched as the next big thing. And today, we feel like we’ve been waiting and waiting. Until recently, Blockchain was too expensive, slow under load, and hard to integrate for most businesses. So enterprises ignored it. It didn’t solve their business problems. That’s changed. Why blockchain, why now? Businesses don’t care about the tech, they care about cost and performance. They’d ask a simple question “Does it save or make me more money?” For a long time, blockchain didn’t clearly do this. That’s no longer true. Blockchain is proving real business cases, especially on Avalanche. On Avalanche, transactions cost fractions of a cent. settle in about a second. And instead of forcing everything onto one shared chain, businesses can launch their own Avalanche L1s with their own rules. To understand this let’s identify the problem and then provide the solution in a way that's easy to understand. Where Businesses Lose Money Most large industries lose money due to operational inefficiencies. Data lives in different systems. Teams spend hours reconciling records that should already match. Intermediaries sit in the middle, taking fees to coordinate all of it. Individually, each step looks small. Together, they create real cost: > Labor spent on manual processes > Capital locked up during settlement delays > Fees paid to intermediaries > Risk introduced by time gaps and mismatched data This is where businesses actually lose money. Not in big, obvious ways. In constant, compounding friction. Take Private Credit, for Example Private credit is loans held outside of traditional banks. It’s a multi-trillion dollar market, and much of it still runs on spreadsheets and weekly reconciliation processes. Loan data is tracked across systems. Teams manually process requests. Funds move on traditional rails, often on delayed cycles. It doesn’t have to be this way Entire teams exist just to keep systems in sync. Now move that system onto Avalanche. Loan data updates in real time. Transactions settle in about a second. Every participant sees the same state instantly. Reconciliation isn’t a separate step because the system itself is the source of truth. The impact is straightforward. > Reduced manual work > Shortened settlement cycles > Fewer layers of coordination between parties Avalanche is Infrastructure for Real Businesses Avalanche is designed to match how businesses actually operate. Instead of sharing a single chain, they can launch their own Avalanche L1s with custom rules, built-in compliance, and predictable performance. They control the system. Avalanche’s Moment For the longest time, blockchain naysayers said this could all be done better with spreadsheets or existing systems. They were right. That’s what the technology allowed. Now it’s changed. Avalanche can replace many of those systems with real-time settlement, shared data, and automated execution. For the first time, the economics work. Built for business. 🔺

Avalanche🔺

13,104 Aufrufe • vor 4 Monaten

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 Aufrufe • vor 7 Monaten

𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more details

Oier Mees

12,379 Aufrufe • vor 2 Monaten

Heres an actual way to make $10k a month from TikTok + organic affiliate TT slides is probably the best way to make AI content for a few reasons > easy and not time consuming to make > much harder to detect images are AI > slideshows rarely get the AI label from TT (especially if you do what I’m gonna show you) >slideshows require less engagement to go viral (more consistent virality) >CTA can be more natural this AI slideshow format is going insanely viral consistently on brand new accounts and no one is even detecting it’s AI they’re super easy to make and use a story telling format that’s really smart you can take the exact same formula to promote sweeps offers from Glitchy and make $10k a month pretty easily here’s the blueprint >Content creation to create the slides realistic you can simply take a photo from Pinterest and put it into Gemini or Chat GPT and ask it to give you an EXACT JSON to recreate the image add in any details you want to add like “make her hair blonde” “make her eyes blue” take that JSON and put it into Nana Banana If you want her in certain backgrounds do the same process and add it in with the JSON of the girl you’ve created or just describe it Now clear the meta data from said image to stop getting the AI label If you still get it go to a meta data analysis site and put the data in ChatGPT Ask if their is anything in the meta data that is signalling this to TT >Writing scripts The reason these go so well it’s because they have a negative scroll stopping hook that instantly make you want to know what the slides going to say next “got fired from my job” “failed my exams” Along with the music that sets the emotion of the video It wouldn’t work as well if they had a random viral song that’s up beat plus the cherry on top is that the image correlates with what’s being said in the hook it’s self acting as a visual hook It wouldn’t work aswell if it was just an image of a girl on her bedroom >how to interpret sweeps Choose an sweeps offer from Glitchy for a retail store (Walmart/target) follow the same format of scroll stopping negative hook Example: “broke my arm” then you would tell a story that paints a bad working environment, maybe she got fired for breaking her arm and is now exposing secrets and your CTA would then be “they don’t promote this but they have a secret feedback program” this is just to give you inspiration but there is literally countless ways >account set up - US proxy / US sim - Download TT with US proxy on - Buy aged account (to help with account trust and getting banned due to proxy issues) - warm up for 2 days (scroll vids all the way through, like, comment authentic things relating to video) Then you post This is a very good way to at least reach a couple K a month but I’d be surprised if you don’t reach $10k beyond

Pounds

10,886 Aufrufe • vor 6 Monaten

For the first time in the history of automation, the people most likely to be replaced are the ones building their own replacement, frame by frame, for a few dollars an hour. Across India, Nigeria, China, and Argentina, workers are strapping cameras to their heads and recording every fold of laundry, every stitch, every washed dish, and that footage is training the robots designed to do those exact jobs. This is documented, not rumor. No jokes! Garment workers in Tamil Nadu, India have been filmed wearing head-mounted cameras on the factory floor, sending point-of-view footage to data firms whose clients include Fortune 500 companies. One US company alone has hired thousands of workers across more than 50 countries to record themselves cooking, cleaning, and folding clothes. More than 6 billion dollars poured into humanoid robots last year, and the one ingredient every maker is starved for is precisely this: real human hands doing real human work. The endpoint is stated plainly by the buyers. In China, one supplier said his pitch to factories is to let workers wear the cameras now, because trained robots will eventually work there instead. The quiet part is the exchange itself. The worker is paid for the hour and keeps nothing after it. No share, no royalty, no ownership of the movements their own body is teaching the machine. The skill leaves their hands and becomes someone else's product, and almost no one along the chain sees the full shape of the trade, not always the person filming, not the millions who watch the clip and scroll on. One scene holds all of it. A humanoid robot spent an hour folding three shirts while a human housekeeper, hired to guide it, quietly finished the rest of the chores. Every automation before this arrived from the outside. A machine showed up and took the job. This one is being built from the inside, by the workers themselves, handing over the last thing they had left to sell. UBI ? or something totally else should pave the way in the future? Thoughts?

Shanaka Anslem Perera ⚡

95,549 Aufrufe • vor 1 Monat

THE DEPTH MAP TRICK THAT FIXED DANCE ACCURACY IN SEEDANCE 2.0 Feed the model a video of someone dancing and it tries to interpret everything- the person, the clothes, the lighting, the room, and somewhere in there, the movement. Feed it a depth map and there's nothing left to interpret but the motion. Most creators trying to transfer a dance to a character reference the source footage directly, then wonder why the choreography drifts. The problem isn't the model - it's that you handed it ten variables when you only wanted one. Here's the workflow 1. Lock the character reference in GPT Image 2 first -face, build, costume, so identity holds independently of whatever motion gets applied to it 2. Convert the source dance footage into a depth map instead of using the raw video -this strips out the original performer's appearance, clothing, and environment entirely 3. Feed the depth map as the motion reference and the character sheet as the identity reference- two separate inputs doing two separate jobs, not one input trying to do both 5. Let the depth map carry only spatial movement -the model receives body position and momentum with no competing information about who's moving or what they look like 6. Keep the character and motion inputs isolated throughout - the moment you mix appearance data into the motion reference, the model starts negotiating between two identities Why this works • Raw footage passes the model everything at once- performer, wardrobe, room, lighting -and the choreography competes with all of it for attention • A depth map is pure spatial information, so the only thing left to transfer is movement • Separating identity from motion means the character can stay locked while the dance stays accurate - normally you're trading one for the other • The accuracy gain isn't the model getting better, it's the model getting fewer decisions to make Use cases: ⁃ Dance and choreography transfer onto original characters ⁃ Motion capture-style workflows without motion capture ⁃ Any sequence where a specific movement needs to survive intact ⁃ Character showcase content built on existing performance footage The character sheet answers who's dancing. The depth map answers how - and keeping those two questions separate is the whole trick.

Nexlow

85,394 Aufrufe • vor 1 Monat

We continue to analyze the technical details of the war with Iran, and we would like to note the Iranian novelty - subsonic barraging anti-aircraft missiles Missile 358, equipped with compact turbojet engines. To some extent unexpectedly, they showed good effectiveness against Israeli reconnaissance and strike UAVs Hermes-900. The key role is played by the thermal homing head: it is able to reliably detect and track a wide range of heat-contrasting targets, including UAVs with various flight profiles. An additional advantage is the command and telemetry channel, which ensures data transmission and allows for radio correction of the trajectory. This is especially important in situations when the target attempts to disrupt the capture with infrared decoys or other means of counteraction. According to the stated parameters, the range of application of Missile 358 reaches about 100 km. At the same time, the maximum interception altitude is about 8.5 km, and the speed is up to 700 km/h, which expands the capabilities of the complex in covering objects and intercepting medium-altitude UAVs. Of course, this is a niche tool, and compared to solid-fuel missiles, the turbojet engine provides exponentially higher flight energy. This allows to dramatically increase the range at low speed, and the mass of the main units of the ammunition, its warhead and control system. On the other hand, this is a solution of necessity, because classic anti-aircraft missiles perfectly hit such high-altitude and slow-moving targets. But for such ammunition, no radar, complex and expensive beam installations are required, which greatly improves its survivability under constant air strikes. And in its niche of targets, there are enough of them, as such UAVs of Israel and the USA are the basis of UCAV, and are constantly over the territory of Iran. So with the 358th ammunition, you can score quite a lot of frags, and significantly complicate air strikes on Iran for the Epstein coalition. Russian Engineer -

𝐃𝐚𝐯𝐢𝐝 𝐙 🇷🇺🇮🇪

17,874 Aufrufe • vor 5 Monaten

LLM Artifacts Connected to Andrej Karpathy's LLM Knowledge base idea, I've been building out a fun way to generate dynamic artifacts from these knowledge bases with the goal of discovering and revealing meaningful and deeper insights. LLM KBs are hard to consume for humans, as I think they are more built for agents. So the question is, what form would be useful for humans to take actions and make important decisions? That's what I am trying to figure out with these artifacts. The artifact example shows a pulse on HN discussions around AI-related stories. The insights can go deeper, of course, but this is already super fun and thought-provoking, like some of my favorite podcasts. The format and depth matter a lot. The aggregation skills of agents are outstanding if you tune the prompts and skill carefully. I built this artifact generator in a few minutes through an agent skill, but I feel like there are so many ways that LLM-generated information can be used and consumed. Like generating deeper insights and analysis, and things that are just not feasible for humans today. The generated artifact (including its data and design) serves as reusable templates or can be updated in real-time via auomations, which is something I am also working on. It is truly an insane way to monitor and track information. Better than a newsletter. Better than newspapers. There is something about this that gets me really excited about the future of AI agents for knowledge generation and discovery. Lots of hidden gems everywhere just waiting to be discovered and acted on if the information is presented correctly. This is not perfect. The format, style/prose can be improved, but this is easy to customize via skill. You can personalize it to your liking. I feel like these dynamic artifacts are going to emerge as a strong new medium to stay on the cutting edge of things, both for agents and humans. My target is research, of course. This was just a basic example. Besides animation, I am also targeting other components like voice, videos, images, slides, etc. This space is full of opportunities to explore. Skill for this coming soon.

elvis

31,295 Aufrufe • vor 4 Monaten

The Chinese are flying 4 sixth-generation prototypes, but what does that mean? While the West keeps debating wars that seem never-ending, huh, China is flying low – or rather, high! – improving their 6th generation fighter prototypes, like the J-36/J-50, with total focus on advanced integration. This gives a huge strategic advantage, with emphasis on long-range missiles and multiple guidance to dominate global scenarios. China already has about 4 6th generation prototypes and plans to reach 8, selecting the most adapted one. All this under the General Concept: Indestructible Flying Brain: 6th generation fighters go way beyond just a slightly improved stealth; they are central platforms that command a global war web via AI, drones, and varied weapons, making previous fighters obsolete in connectivity and limiting them to very local operations. This omnipresence redefines air superiority, with the fighter surviving as a resilient node in the first hours of conflicts and being able to operate with speed. Kill Web: The Global War Web: The fighter acts as the central node of a real-time network, connecting submarines, satellites, ships, drones, and troops worldwide. It allows omnipresence, receiving data from a destroyer thousands of km away and attacking as if it were right nearby, with AI assisting the pilot in analysis and target acquisition. That's why the Chinese focus on missiles with ranges of thousands of km, with multiple guidance, turning the 6th generation pilot into a tactical manager very different from today's. Being a 6th generation fighter pilot is going to demand a lot. Command of Drone Swarms (CCA/Loyal Wingman) The fighter controls 6-20 drones simultaneously for reconnaissance, jamming, or suicide attacks. It transforms the pilot (or AI) into a "maestro" of a robotic orchestra, or quarterback of Collaborative Combat Aircraft (CCAs), which carry extra weapons, expanding offensive power without exposing the main fighter. Like, a controlled symphony of destruction! Superior Multi-Spectral Stealth (Stealth++) Not limited to radar, it covers infrared, acoustic, visual, and electromagnetic. It uses advanced materials, tailless designs, and minimal thermal signature to penetrate dense A2/AD defenses, making it extremely hard to detect and essential for operations in contested environments. Extreme Range, Autonomy, and New Generation Weapons Combat radius of 1,800-2,500 km without refueling, with sustained supercruise (Mach 1.5-2.0) without afterburner, thanks to adaptive cycle engines and huge internal tanks. There's talk of including lasers, but so far, what's really there are internal hypersonic missiles and 2-3x greater armament capacity than the F-35, all while maintaining total stealth. Artificial Intelligence, Integrated Sensors, Resilience, and Open Architecture AI as co-pilot or main, processing data in real time and making tactical decisions to reduce human load; optionally manned mode: piloted, remote, or autonomous flight; virtual cockpit via helmet visor. Multifunctional sensors combine radar, electronic warfare, communications, and non-kinetic effects, with total data fusion transforming the fighter into a flying data center. Network resistant to jamming and GPS loss via quantum-resistant communications, mesh networks, and inertial/computer vision navigation. Modular architecture allows quick upgrades (90% by software), avoiding high costs like in the F-35; in "Decision Centric Warfare," AI decides in milliseconds, with the human as an optional bottleneck, including cyber warfare and active defense. In another article, I'll talk about what I think of this in terms of costs and demand and if such an investment is really worth it.

Patricia Marins

60,403 Aufrufe • vor 9 Monaten

In 2025, demand for blockchain applications with genuine real-world utility has collided with a technical barrier that leaves developers questioning what they can realistically build. Anyone building things like tokenized assets, supply chains, AI agents, or prediction markets still juggle a mess of middleware, and somehow end up spending more time stitching than innovating. How so? Every: - Bridges to move assets, - oracles to fetch data, - indexers to make that data searchable, - relayers and bots to keep everything on schedule— is necessary, but each layer also adds cost, latency, and new risks. The end result is an application that’s expensive to run, fragile under stress, and slower than the Web2 software it’s trying to replace. This is the problem Rialo says it wants to solve. Built by Subzero Labs and backed by $20 million from investors like Pantera Capital and Coinbase Ventures 🛡️, Rialo’s pitch is simple: instead of accepting the middleware tower as an unavoidable cost of doing business, compress it into the base chain itself. But Rialo doesn’t describe itself as another Layer 1, its very name, Rialo Isn’t a Layer One, makes that clear. The team frames it instead as a unified real-world network: a protocol rebuilt from the ground up with the assumption that external connectivity is not an afterthought but a core design principle. To understand what this means, consider how today’s dApps are typically assembled. A typical RWA dApp stack involves: - Oracle providers (Chainlink, Pyth, Band) for asset pricing and event settlement - Bridges (Wormhole, Multichain, custodians) for cross-chain asset movement - Indexers (The Graph, Aleph, Stacks API) for querying and preprocessing chain data - Schedulers/relayers for automated tasks and monitoring - Web2 integrations via cloud services, centralized APIs, and off-chain pipelines Each of these steps adds another vendor, another trust boundary, and another operational layer to monitor. By the time the application is live, it resembles a patchwork of loosely coupled services, each carrying its own risks. You don’t have to look far for proof: - Base went dark for 29-43 minutes in August 2025 when its sequencer misfired, freezing every DeFi app on it. - A few months earlier, an AWS outage rippled through Binance and KuCoin, stalling withdrawals because even “decentralized” systems leaned on centralized middleware. - When Infura has faltered, Ethereum dApps have gone offline in sync, not because Ethereum broke, but because the middleware holding it together did. What should feel like building an application instead feels like maintaining a fragile machine. Rialo architecture embeds the primitives that normally live in middleware directly into the protocol. Smart contracts on Rialo can: - be event-driven, able to respond not just to blockchain state changes but also to external events through built-in webhook and API triggers. - fetch data from the web natively, without relying on external oracles or relayers. - include privacy and identity management—KYC hooks and two-factor authentication, at the protocol level rather than as add-ons. - handle cross-chain communication without wrapped assets or third-party bridges. - run on a virtual machine that is compatible with ecosystems like Solana but extended with RISC-V to support modern programming concepts such as async/await and event loops. If these features work as intended, the implications are significant. Today, much of a team’s energy goes into building and maintaining infrastructure: fullnodes, indexers, monitoring scripts, oracle integrations, relayer logic, bridge infrastructure. Each requires engineering headcount and ongoing maintenance. With Rialo, much of this is absorbed by the protocol, freeing developers to concentrate on business logic. Projects can deliver production-grade dApps with smaller, leaner groups focused directly on product design and execution. Operational costs also shrink: indexing and oracle services can run into thousands of dollars a month; collapsing those into built-in functions reduces recurring expenses while simplifying onboarding for new developers. But folding middleware into the chain doesn’t erase complexity, it reshapes it. Some of the problems to be encountered include: - Scale and complexity: Rialo’s validators won’t just be securing transactions; they’ll also be securing APIs, cross-chain data, and scheduled triggers. Any failure in one subsystem could ripple across the entire network. - Performance vs. decentralization: Richer indexing, scheduling, and data ingress could make nodes heavier to run, narrowing who can realistically participate as a validator. That risks reducing the decentralization blockchains depend on for resilience. - Governance pressures: Disputes or failures involving real-world data feeds, external APIs, or cross-chain actions will arise more often, requiring not just technical fixes but robust social infrastructure, clear rules for voting, transparent arbitration, and mechanisms for community trust. Without them, Rialo risks re-centralizing decision-making around a handful of operators. Where, then, does this model make the most sense? That would be in sectors where external connectivity is indispensable and middleware bloat has consistently been a blocker: - Real-world assets: settling tokenized securities or commodities against off-chain events. - Supply chains: triggering a payment the moment a shipment clears customs, without relying on a third-party oracle. - Agent systems: AI agents interacting with real-world APIs and on-chain contracts simultaneously. - Real-time markets: prediction markets or insurance contracts that must resolve immediately against external data. For purely on-chain domains like DeFi primitives or NFTs, where composability matters more than external triggers, the advantages may be less pronounced. This shift is familiar to anyone who remembers the rise of Web2 platform services. Just as Heroku and Firebase abstracted away server maintenance so developers could focus on building products, Rialo is betting that a unified real-world network can let blockchain developers do the same. Adoption will ultimately depend on: - whether its protocol primitives mature quickly, - whether the ecosystem builds out SDKs and tooling that make them usable, - whether compliance features can adapt to changing regulations, - and whether governance proves resilient under adversarial conditions. The first applications will be the test case. If they show that Rialo can replace a fragile patchwork of middleware with a secure, auditable, and cost-effective base layer, it could set a new standard for real-world connectivity in blockchains. If not, it risks simply moving complexity from one part of the stack to another. But at a minimum, Rialo has forced the question: should real-world connectivity in blockchains continue to depend on layers of external vendors, or should it be built into the chain itself? That’s the question Rialo has put on the table — and it’s why I got interested in Rialo .

Jen

12,178 Aufrufe • vor 10 Monaten

Though this video might not look like much, it shows something pretty remarkable. This footage, which we captured this past September on the Kabetogama Peninsula, is the first confirmed observation of lynx kittens in Voyageurs National Park (and one of only a few observations of known lynx reproduction in the Greater Voyageurs Ecosystem). Although lynx have been observed/documented in the park sporadically for several decades, there has never been evidence of kittens in the park (i.e., a breeding population of lynx). As a result, most lynx in and around Voyageurs are likely transitory individuals (predominantly males) roaming large area. From 2000 to 2004, biologists with the National Park Service made a concerted effort to study lynx in the park. They only documented the presence of one male and one female, and concluded “lynx were either transient or present at a low density”. Another study from 2007-2008 in and around Voyageurs National Park using trail cameras and snow tracking did not detect the presence of lynx. However, they noted that there were some sightings of lynx outside of the park with one unconfirmed report of a female lynx with a kitten observed to the west of the park. The researchers concluded that “even though patches of high-density snowshoe hare habitat exist in the Voyageurs National Park area, the low density of snowshoe hares at the landscape level would not support resident lynx”. Based on this research and some subsequent work, researchers with the park concluded in 2015 “it does not appear that there are currently resident lynx”. In fact, the researchers in 2012 concluded: ‘only three confirmed observations of adult lynx have occurred within the boundaries of VNP since 2001”. Obviously, studying these elusive animals in places like Voyageurs has historically been difficult because trail camera technology was not what it is today (or did not exist). However, trail cameras have provided a great tool to observe and understand lynx in places like Voyageurs. Skip forward to today: we now routinely get numerous observations of lynx each year in Voyageurs National Park as well as outside the park. Many observations are likely of the same few individuals wandering around, though, based on physical appearances, it seems we are capturing observations of more than 1 or 2 lynx. That said, in general, our trail camera data generally supports the conclusions of the Voyageurs National Park researchers. Most lynx likely are transient individuals, and certainly lynx are at low densities in and around Voyageurs. Nonetheless, this observation of a lynx with two kittens shows that lynx can reproduce in this area, though such occurrences are likely rare. One observation of reproduction is not evidence of a self-sustaining resident population but it suggests it might be possible. Anyway, this finding just highlights how our trail cameras, which are deployed to study wolves, simultaneously capture rare and valuable data on other wildlife species that have been traditionally difficult to study. Sources: Route et al. 2009. Status of Canada lynx in Voyageurs National Park, Minnesota, 2000-2004. National Park Service Report. Moen et al. 2010. Lynx habitat suitability in and near Voyageurs National Park. Natural Areas Journal. Moen and Windels. 2015. Lynx Habitat Suitability. Voyageurs National Park Website.

Voyageurs Wolf Project

24,233 Aufrufe • vor 7 Monaten