Want better MiniMax-H3 LoRAs? DiffSynth-Studio just open-sourced two things... worth checking out. 🔥 ⚙️ The MiniMax-H3 Training Adapter, a rank 64 LoRA. Combined with differential training, it helps you train higher quality LoRAs: 🔗 📦 The self-generated dataset behind the adapter, fully open as well: 🔗 One more thing: the adapter will soon power MiniMax-H3 LoRA training on ModelScope Civision. Stay tuned. 🚀show more

ModelScope
10,197 просмотров • 14 дней назад
Hailuo / MiniMax H3 is here! 🚀 Just ran... my first experiment with the new H3 model using an idea I previously tested on the 2.3 version — and the results are surprisingly impressive. This was a true one-shot generation: first attempt, no retries, no selecting the best take. More experiments coming soon as I explore what H3 can really do. 🔥 Diving deeper into Hailuo AI (MiniMax) ✨ #MiniMaxH3show more

Anissa
14,609 просмотров • 1 месяц назад
Introducing MiniMax H3 on Argil Some days, the future... decides to accelerate. Two days ago, we launched the new version of Argil. Today, I’m extremely happy to announce our partnership with MiniMax (official): H3 is live in Argil. We tested it internally. The quality is really solid. An intelligence that reads text, images, video and audio as one language. That captures creative intent at a glance and already delivers production-ready content. MiniMax H3 will soon be open source. And that’s a real game changer. For the occasion, Golden Ticket: 40% off Pro & Business yearly during MiniMax Hive until tonight. Comment "argil" I’ll send you the code in DM. To those who create: the field just got wider. Link 👇show more

Brivael Le Pogam
23,030 просмотров • 1 месяц назад
I've been testing MiniMax H3 in early access, and... it's one of the most capable AI video models I've used recently. What stood out to me most is the level of creative control. Instead of relying on a single prompt, I can combine text, images, videos, and audio to guide the final result much more precisely. A few features I've been enjoying: ✅ Native multimodal understanding and generation across text, images, video, and audio ✅ Omni Reference for more consistent characters, motion, and camera movement ✅ Precise editing that follows detailed instructions surprisingly well ✅ Commercial-grade output built for ads, branding videos, UI demos, game cinematics, and music videos ✅ Cinematic visuals with impressive audio quality I'm still exploring everything H3 can do, but my first impression has been really positive. Looking forward to sharing more examples as I continue testing. What would you create with MiniMax H3? MiniMax Design (H3) #MiniMaxH3 #HailuoAIshow more

Manish Kumar Shah
22,942 просмотров • 1 месяц назад
MiniMax H3 is now 50% OFF on Magnific for... 2K video, only until September 1. 🔥 I’ve been trying MiniMax H3 on Magnific, and it feels like a big upgrade for AI video creation. It’s not just about turning text into videos. You can use text, images, videos, and audio together in one prompt, giving you more control over the final video. Here’s what makes it stand out: - Multimodal: Use text, images, video, and audio in one prompt. - Multiple references: Add up to 9 images, 3 videos, and 3 audio files. - 2K video: Create videos up to 15 seconds long. - Built-in sound: Generate voice, music, and sound effects with the video. - Easy editing: Remove objects or transfer motion easily. - More control: Control the camera, characters, and voice. You can use it to turn posters into videos, moodboards into short films, and product images into ads. It also helps bring your ideas to life with realistic movement, lighting, reflections, and sound. The workflow is simple: give it your references → generate → edit → refine. Try MiniMax H3 on Magnific:show more

Markandey Sharma
96,964 просмотров • 22 дней назад
Minimax H3, same 4x3090 rig, same clip. Yesterday it... rendered in 11:21. Today it renders in 3:45 🤯 The hardware did not change. Two things happened overnight. An AI agent rewrote the attention CUDA kernel and hit a half speed trap inside GeForce tensor cores that most people never heard of. Then a stranger on HuggingFace dropped a LoRA that cuts 20 sampling steps down to 4. One of these two mattered way more than the other. Breakdown below. I packed all of it into ready ComfyUI workflows for my rig. If you want them, ask in the replies and I'll share.show more

Alexey Fateev
70,024 просмотров • 1 месяц назад
The new H3 model from MiniMax Design (H3) is... a beast! At 7000 character context adherence limits, and native 2k out of the box.. this model destroys Seedance 2 in my opinion.. and the reason for that is this is rollout level capabilities, far beyond what Seedance can do presently, and what they rolled out with. Meaning, updates to this model over time will also exceed that of Seedance. Not only that, ITS CHEAPER and NO CHARACTER/FACE RESTRICTIONS 🔥🔥🔥🔥🔥🔥🔥🔥🔥 And the CONSISTENCY IS OFF THE CHARTS ❤️ and the adherence is so good I have t been able to break it yet! 🙃 Just to show you, my previous teaser was text-2-video, and these two are the first two out of the gate with one image ref!show more

Dustin Hollywood
11,745 просмотров • 1 месяц назад
Yeaah with Kartel.ai and Jesse Wellens ! So training... Lora for character on Flux and also for Wan 2.1 14B ( using dataset video ) make a Huge difference. So here you are seeing a volumetric capture of Jesse Wellens that we did in at the spatial studio of Kartel.ai We use a a World Labs gaussian splatting to create a simple environment that we put with the volumetric capture in our own webgl viewer ( out soon for everyone ). After that, We use after this output in a ComfyUI workflow with Wan Fun control ( a bit similar Vace but its working for Wan 14B) with the double loras , the first to generate the first frame with flux, and the second one to guide the generation of Wan 2.1 Control. So we keep very good consistency of character and position, and creating amazing worlds :) ! Hope you will find it cool !show more

Lovis Odin
23,066 просмотров • 1 год назад
I see a lot of online posts about 🍌... Nano Banana 🍌 making ANYTHING POSSIBLE, and all you do is CLICK & BOOM, but it's not that simple. There are lots of ways to get incredible results, but nothing is ever as simple as it seems, so... In the coming weeks, I'm going to break down some of the key ways to get the most out of Nano Banana, like head/face swapping, product placement (and yes, I mean your obscure product loaded with small print text), scene blocking, and more. Connect with me, here on X to stay in the loop on upcoming posts, and if you want to get more serious 1:1 AI training, feel free to schedule some AI consultation here: Freepik Hailuo AI (MiniMax) #NanoBananashow more

Jordan Daniel Chesney
27,760 просмотров • 1 год назад
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
55,281 просмотров • 26 дней назад
Introducing ml-intern, the agent that just automated the post-training... team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.show more

Aksel
1,267,561 просмотров • 4 месяцев назад
Let’s try a fashion campaign in Rome with the... new model MiniMax H3 of MiniMax Design (H3) 😎 The result is really stylish and realistic😍 Don’t you think!? #MiniMaxH3 Video prompt: A 15-second hyper-realistic cinematic fashion campaign video. Blonde woman, fair glowing skin, sharp features, long blonde hair in sleek high ponytail, small dark red rectangular sunglasses, pearl drop earrings. Wearing open black leather jacket over white cropped t-shirt with “Keor” print, shiny red mini skirt, matching red pointed-toe heels, and white smooth leather hobo bag with gold dome studs hanging on her arm in every scene. Scene 1 · 3s Full-body wide shot. She stands on a Rome cobblestone street near the Spanish Steps, adjusting her sunglasses. Warm golden-hour light, real people and city chaos blurred around her, classic Roman buildings and Spanish Steps in background. Scene 2 · 3s Full-body shot. She strides boldly across a zebra crossing on Via del Corso, iced coffee in one hand, bag swinging on the other, mouth slightly open as if singing. Confident walk, traffic blurred behind. Scene 3 · 3s Cinematic wide shot. She walks tiny in the center of a grand Rome boulevard with the Colosseum massive behind her. Real crowd around, bag clearly visible. Scene 4 · 3s Medium shot. She stops near Piazza Navona and does a spontaneous carefree dance move, arms slightly out, bag swinging. Street musician blurred in background, joyful expression, golden light on her face. Scene 5 · 3s Medium close-up. Inside a Rome Metro carriage holding the pole, she slowly pushes her red sunglasses up with one finger and gives a slight smirk. Bag visible on her arm, moody warm metro lighting, passengers naturally blurred. Smooth seamless transitions. Mood: candid chaotic Roman city energy, luxury streetwear editorial, golden-hour warmth. Shot on Canon EOS R5 35mm f/1.4, Kodak Portra 400 film tone, natural grain, handheld movement, real people in every background. Hyper-realistic, no plastic skin, no robotic movement, no stiff poses, no AI artifacts. Avoid: cartoon, CGI, plastic skin, robotic/stiff movement, blurry face, overexposed, watermark, texshow more

KeorUnreal
14,334 просмотров • 1 месяц назад
The future of housework just leaked on GitHub and... nobody is talking about it. knox byte just open sourced a framework that coordinates swarms of Unitree G1 humanoid robots to clean your entire house on their own. It's called ARGOS. You tell it "clean the bedroom" in plain English and 2+ G1 robots split the room into zones, sweep in parallel, and sync up for the tasks that need four hands like making the bed or moving furniture. The Claude API decomposes your sentence into a task graph. An auction system makes every robot bid on every task based on distance, battery, and current load. The cheapest robot wins. Cooperative jobs go to the cheapest team. Here's what makes this different from every demo video Boston Dynamics keeps teasing: → 12 cleaning tasks baked in sweeping, mopping, wiping, vacuuming, taking out trash, making the bed, changing sheets, moving furniture, sorting items → 3 policy architectures running underneath OpenVLA-7B for language tasks, Diffusion Policy for floor coverage, ACT for dexterous bimanual work → Train it on your own footage record yourself cleaning, run one command, it extracts poses, builds a LeRobot dataset, and LoRA fine-tunes the policy → PEFA protocol for cooperative work Propose, Execute, Feedback, Adjust. If one robot fails halfway through making the bed, the team replans and retries → Full MuJoCo simulation so you test policies before pushing them to real hardware → Silver and cyan terminal dashboard that shows live fleet status, zone maps, task queues, and battery levels in real time The G1 robots talk to each other over CycloneDDS mesh using Unitree's native SDK. No cloud. No middleware. The whole thing runs on a Jetson Orin inside each robot. The wildest part is the training pipeline. Drop cleaning videos into a folder, run argos train ingest, and the framework does the entire pipeline frame extraction, pose estimation, action labeling, HDF5 dataset, fine-tune, evaluate in sim, deploy to robot. One command per stage. Unitree G1s already exist. The framework to make them clean your house just hit GitHub. 52 stars. MIT License. 100% Opensource.show more

Guri Singh
27,404 просмотров • 3 месяцев назад
You do not need to be in the gym... six days a week. Doing so will reward you with worse progress, not more. Here's why: Fatigue accumulates. Not just from session to session. From week to week. From month to month. The hard leg session on Monday leaves a tax that has not been fully paid by the time you're back under a bar on Tuesday. Multiply that across six sessions and you are walking into every workout pre-fatigued. The first casualty is your high threshold motor units. The fibres with the most growth potential. The ones recruited only when the muscle is fresh enough to call on them with real load. Train under accumulated fatigue and those fibres never get touched. You're left stimulating the lower threshold fibres you maxed out years ago. Two rest days a week should be the bare minimum for anyone training hard. Three is better. Four sessions of about an hour each is not the maintenance protocol the high-volume crowd will tell you it is. It is the optimal dose for a natural lifter who wants to actually grow. Four hours of lifting a week. Four. That's it. The best physiques in your gym are almost always built on less work than the worst. The volume bros sweat five hours a week for a body the four-hour man built while still having time to read a book. More is not the answer. Recovered is.show more

Sama Hoole
147,964 просмотров • 3 месяцев назад
Training glutes still doesn’t get the love it should... from men. Aesthetically, sure, your focus will be more on the V taper frame, or strong arms. But as the largest and strongest muscle in the body, the glutes have the potential to really boost your transformation to a top 1% physique. It’s the same concept as investing money into the market or a cash flowing real estate deal. You put in the capital… And then the asset pays you for your ownership. When you have strong glutes, you now have the biggest muscle in the human body paying you metabolic “dividends” around the clock. Burning more calories just by existing. As a result, you get leaner while having MORE wiggle room diet wise. Which makes your hormones improve, and all of the sudden you have this flywheel of positive change. It’s a compounding effect. Not to mention, it helps keeps your back and hips healthy and flexible. You retain more of your ability to move well as you get older. Strong glutes = absolutely one of your biggest ASSets (da dum, tssssk 😂) And the whole point of owning assets is that they work on your behalf, instead of you staying on the treadmill of the same effort over and over and over again just to get average results. If you want to get a top 1% physique without the suffering and sacrifice promoted all throughout the fitness industry, apply to partner with me 1-1 at the link in my bio.show more

Coach Paul
22,773 просмотров • 1 год назад
Grip a club in your lead hand. Observe how... twisting your forearm changes the angle of the face. This is largely what controls where your ball flies. Small differences in this angle lead to huge changes in shot direction. This precision is not being trained in the gym with “golf exercises”. Poor control of this face angle is what leads to a whole host of swing issues. More so than physical limitations. For example, it is far more likely you come over the top in an effort to stop the ball going right because your face is wide open and you’ve seen the ball go right the last 5000 times than it is you lack hip or shoulder mobility. It would be beneficial for me as a trainer to say otherwise, but I just don’t believe it to be true. Use the gym to work on your physical capabilities. Improve your mobility, strength, and power. These are the things you are losing as you age. Leave the swing mechanics for your practice and golf lessons. If you try to “fix your swing in the gym” you will likely do a poor job of training your swing and your body.show more

Fit For Golf - Mike Carroll
227,270 просмотров • 7 месяцев назад
more frontend vibecoding tips (results below): WHY YOUR VIBECODED... FRONTENDS ALL LOOK THE SAME AND SUCK: when asked to make a frontend, the agent/llm will default to the center/average of its training data (in a very loose sense). through the training process, the model essentially converges on some default UI style. it's very capable of doing things that are different from this style, but you have to ask! for instance, ChatGPT tends to reply in the same tone for all users untill you interact with it and instruct it differently ("be sassy", "eli5"). the second reason is that most of us are not good at coming up with designs and describing them precisely (see my tweet on a crash course in common components, which i'll link below). treat frontend generation just like any other eng task! you need to provide a good detailed spec. TIPS: 1. give ur agent screenshots of designs you like (you may not know the right words to describe them but the agent will! a pic = 1000 words) where to find ui inspo? Behance, Dribbble, Mobbin (Mobbin is paid but worth it!) 2. ask ur agent for proposals, this helps "seed" different directions so the final frontend stands out. don't be afraid to go back and forth. 3. ban certain tendencies: no Inter/Roboto, no shadcn (controversial), no gradients, no emojis 4. encourage the agent to be extreme and make bold decisions, not safe ones. i think that the underlying models tend to get taught during RL/fine-tuning to make conservative choices that produce reasonable but boring frontends 5. give ur agent Figma MCP. the best results will come if you mockup your vision in Figma first. 6. Ideally choose an agent with vision capabilities TLDR: Most people are tremendously underusing agents for frontend design. They are much better than you might expect.show more

andrew gao
64,712 просмотров • 6 месяцев назад
AI token usage is up 10x in 7 months,... compounding 40%/MONTH! There is NO BUBBLE when demand is STILL accelerating And this is just OpenRouter, it doesn't count the labs direct token usage and APIs But here's what's interesting about these numbers, the demand is coming from everywhere at once US models (OpenAI, Anthropic, Google) keep growing, while Chinese open weight models (DeepSeek, Tencent, Xiaomi, Minimax) grew even faster and now drive over 60% of usage on OpenRouter Closed source and open source both compounding at the same time. This is literally the best case scenario for AI Infra investors It means both frontier model tokens and cheaper tokens have product market fit. This means the application layer is finding ways to use both and generate ROI with both types Demand for tokens IS demand for compute. This is why SpaceX is looking to build 10GW of compute by next year, because the demand is clearly here Now combine this demand set up, with NVIDIA yesterday announcing financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilize over $500 billion of third party capital for AI infrastructure And Jensen has said publicly he expects $3 to $4 TRILLION of AI infrastructure spend by 2030 The build out will have to continue for a lot longer than the market is expecting, that is very clear to me. Don't let this consolidation period in AI infra stocks shake you out, they will have their moment again and take their next leg higher p.s. if you want to see how im investing in this, you can track my real-time portfolio and the research of all 5 Milk Road PRO analysts with live trade notifications, and it's just $1 to try it out (insane price just to check it out). Learn more here: Good luck out there!show more

Kyle Reidhead | Milk Road
28,320 просмотров • 27 дней назад
🚀 Introducing EgoExo Forge - built on top of... Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tunedshow more

Pablo Vela
35,626 просмотров • 1 год назад
A Letter to Our Community: The Road Ahead for... Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷show more

Axis Robotics
28,096 просмотров • 8 месяцев назад
Are you safer with LIDAR, or are you safer... with vision? This is a false dichotomy. The more pertinent question today is "do you have something, or do you have nothing?" As you can see from the clips below, vision based systems avoid countless potential collisions every day. The difference between a crash and no crash isn't what sensor suite you chose — it's whether you have any AI on your car at all. Even if we concede that LIDAR may help prevent some additional crashes, we are really debating whether it is 1% of crashes or 0.00001% of crashes. Not all crashes are super complex and require lasers to detect. Most are simple, routine, and can easily be prevented by today's vision based AI. In fact, evidence is mounting that computer vision based systems can actually outperform more traditional approaches to self-driving. Why? Because the low cost of cameras enables you to create a much larger, more varied, and more diverse dataset. If you want to have expensive custom cars that's fine, but you're going to get fewer vehicles for the same budget. Seeing what's in front of you now is actually less important than predicting what's going to happen next — and the large scale datasets used to train pure vision systems are the best for predicting what's next. Counter-intuitively, the simpler and lower cost sensor actually has properties that make it better suited for training advanced AI. Computer vision based self-driving is often framed by LIDAR proponents as "cheaping out" on the sensor suite to save money. But it's not about being cheap, it's about bringing the technology to everyone. 1.2 million people die on the road every year around the world. That's around 39 million people who've died on the roads around the world since I was born — equivalent to a city the size of Tokyo or New Delhi getting wiped off the map. The status quo is simply unacceptable, and something has to be done to fix it as soon as possible. Of the 1.2 million people that will die on the roads this year, about 40,000 will be Americans. That's about 3%. So if we moved entirely to self-driving cars in America and brought crashes down to 0, 97% of the world's crash fatalities would still be taking place as usual. Deploying a $200,000+ retrofitted self-driving car may work in a few American cities, but it is not going to make sense in most places around the world where fares are much cheaper. Most often, the choice is not between LIDAR and vision. It's between vision or nothing. The best system is the system that's there running on my car when I need it to save my life. To say that all self-driving cars must have LIDAR is to sentence most of the world to death. We can't write off computer vision if we want to make a serious dent in this problem. It's going to be a key piece of the solution. Let LIDAR based players build the best self-driving car they can, and let vision based players do the same. We need to be trying everythingshow more

Whole Mars Catalog
45,801 просмотров • 1 год назад