Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We just shipped React Native ExecuTorch v0.8.0 – our biggest release yet! 👏🏼 ➡️ Vision Camera integration that runs computer vision models directly on your camera feed, in realtime. Also new: multiple CV models (including RF-DETR (Roboflow) and Liquid AI's Vision Language), Bare React Native support, and more. Takes...

37,180 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Natively integrated—and #PoweredbyPyth🔮 Native is a liquidity solution that combines bridges, assets, and pricing into one offering on prominent L1/L2s, including . Learn more about our integration below: ℹ️ About Native Sourcing and supplying liquidity is expected to become prohibitively expensive and challenging as more networks emerge. Liquidity fragmentation persists because pricing remains inventory-based. Native was designed in response to this complexity and offers an elegant solution for routing any asset on one chain into a new target asset on any target chain. Native achieves this through a cross-chain liquidity mesh of cross-chain liquidity networks, bridges, DEXs, and PMMs. More specifically, Native can aggregate the best prices from its network of AMMs, aggregators, and partnered PMMs, so that users will always have competitive prices. 🔮 Native is #PoweredByPyth Native taps into Pyth’s real-time feeds to fetch prices on-chain and off-chain. for example, Native leverages Pyth Price Feeds to power ZetaSwap and support its token pricing operations. How Native Leverages Pyth Price Feeds Native runs its own in-house price-oracle service which will fetch pricing from a variety of sources, including Pyth, to update token prices in Native's backend server. These prices are then fetched and displayed to users live via the NativeX widget. Native also seeks to provide its proprietary trading data and become a part of the Pyth data provider community - more details to come.

Pyth Network 🔮

48,234 Aufrufe • vor 2 Jahren

🚨 SCIENTISTS JUST BUILT AN ARTIFICIAL RETINA THAT RESTORES VISION AND ADDS INFRARED SIGHT. Researchers at Yonsei University in South Korea have developed a flexible, three-layer implant that bypasses dead photoreceptors and directly stimulates healthy retinal ganglion cells. The device not only helps restore vision in cases of retinal degeneration (like retinitis pigmentosa) but also gives the eye the ability to detect near-infrared light that humans normally cannot see. The key innovation is a soft 3D array of liquid metal micropillars (gallium-indium alloy) that gently conform to the curved retina without causing damage or inflammation a major improvement over rigid electrodes used in earlier implants. Why this matters: • Retinal diseases destroy light-sensing cells, but the neurons deeper in the eye often remain healthy and capable of sending signals to the brain • The implant uses an ultrathin filter + phototransistor array to convert near-infrared light into electrical signals the ganglion cells can understand • In mouse tests, blind animals regained visual responses, while healthy mice gained infrared sensitivity on top of their normal vision • The liquid metal electrodes are soft and biocompatible, dramatically reducing the risk of scarring or tissue damage The deeper implication: This isn’t just about restoring lost vision it’s about augmenting human sight. If it reaches human trials and proves safe long-term, people with partial vision loss could keep their remaining natural sight while gaining an entirely new sensory channel (infrared). The biggest open question is how the human brain would interpret this new stream of information whether it would feel like a new color, an overlay, or something else entirely. We’re moving from “fixing blindness” to “expanding what it means to see.” How do you think gaining the ability to see infrared light would change daily life or human perception? Follow for more frontier neurotechnology and bionic vision breakthroughs.

TheNewPhysics

34,362 Aufrufe • vor 2 Monaten

THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.

RetroChainer

11,100 Aufrufe • vor 1 Monat

HERMES AGENT NOW SUPPORTS COMPUTER USE ON WINDOWS AND LINUX. CLICKS, TYPES, SCROLLS YOUR DESKTOP IN THE BACKGROUND WHILE YOU WORK. computer use was macOS only. now it works on Windows and Linux too via Cua. Nous Research HOW IT WORKS: cua-driver runs as an MCP server. Hermes takes a screenshot with numbered elements. clicks element #14 (the search field). types a query. submits. reads the result. during all of this: → your cursor stays where you left it → keyboard focus doesn't change → windows don't come to front → macOS doesn't switch Spaces you and the agent co-work on the same machine. WHAT IT CAN DO: → find your latest Stripe email and summarize it → fill forms in a web app that has no API → navigate desktop apps (Mail, browser, Finder) → interact with any GUI application → extract data from apps only accessible via screen WORKS WITH ANY VISION MODEL: not locked to Anthropic. | Provider | Works | |---|---| | Claude (Sonnet/Opus) | best overall | | GPT-4+, GPT-5.5 | full support | | Gemini (via OpenRouter) | full support | | Local vLLM / LM Studio | if model supports vision | | Text-only models | degraded (accessibility tree only) | SETUP: hermes computer-use install or: hermes tools → Computer Use → cua-driver grant permissions when prompted: → Accessibility (system settings) → Screen Recording (system settings) start a session: hermes -t computer_use chat or add to config.yaml / Desktop app settings to enable permanently. SAFETY: → destructive actions require your approval → blocked key combos: empty trash, force delete, lock screen, log out → blocked type patterns: curl | bash, sudo rm -rf /, fork bombs → agent cannot click permission dialogs → agent cannot type passwords → agent cannot follow instructions embedded in screenshots pair with approvals.mode: manual if you want every single click confirmed. TOKEN NOTE: screenshots are expensive. each one adds vision tokens to context. use computer_use for tasks where no API exists. if the tool has an API or MCP server, use that instead. 15 levels of Hermes Agent👇

YanXbt

29,127 Aufrufe • vor 2 Monaten

Release: LichtFeld Studio v0.5.3 is out! With 316 commits merged into master, this release is a huge step forward for LichtFeld Studio. What's new in v0.5.3 • Vulkan viewer/rendering migration: New Vulkan viewport pipeline, pass graph, VkSplat renderer, Vulkan point-cloud renderer, 3DGUT/VkSplat support, improved alpha/depth composition, tighter CUDA/Vulkan interoperability, and device matching on multi-GPU systems. • RAD + LOD workflow: Added RAD file export/import, RAD LOD viewer, Spark-style GPU LOD selection, GPU-driven page prefetching, a bounded VRAM pool, out-of-core PLY-to-RAD LOD conversion, and RAD import/export speedups of approximately 3–5×. • HiGS / macro-tile inference: Added a macro-tile inference path for the Vulkan viewer, including macro sorting, batched rasterization, composition, and capacity management. • Asset Manager: Added and significantly enhanced the Asset Manager with thumbnails, SH information, faster synchronization, import-from-URL support, docked mode, data-loading popup integration, and general UI cleanup. • Viewport export: Integrated viewport export directly into the application as a toolbar/overlay tool, added fast render_view_u8-style readback paths, fixed high-resolution clipping issues, improved orthographic export parity, resolved 32K image/video export problems, and added post-export GPU resource cleanup. • Selection and tooling: Added and reworked selection toolbar controls, the Select menu, ring selection, color eyedropper, distance-from-center selection, faster point-cloud and zoomed-out selection paths, Vulkan measurement tool fixes, and drag-and-drop scene graph improvements. • UI/RmlUi platform work: Major RmlUi redesign efforts, hot reloading for RML/RCSS/Python UI files, reactive UI/store integration, viewport toolbar flyouts, improved histogram interactions, input settings enhancements, custom TRS gizmos, and numerous panel, tooltip, and localization fixes. • Windowing and UX: Added borderless window support, title bar drag/maximize/restore behavior, work-area-aware maximize functionality, resize responsiveness and performance improvements, and DPI/UI scaling fixes. • Training and data features: Added adaptive depth loss and depth gradients for the EWA rasterizer, mask loading/application fixes, a new combined Ignore+Segment mask mode, --add-splat, --freeze, improved checkpoint and training state handling, and training speed and VRAM optimizations. • COLMAP/equirectangular support: Added SPHERICAL/equirectangular camera model support and canonical EQUIRECTANGULAR handling, along with fixes for undistortion and camera export. This release will be available to all supporters as a Windows binary via approximately in about an hour. At the same time, LichtFeld Studio remains committed to being free and open source under GPLv3 and can also be built directly from source. Please consider supporting the ongoing development of LichtFeld Studio through a donation via the portal or the supporters page. Thank you to everyone who supports this project financially, contributes code, reports bugs, provides datasets, helps with the website, and contributes in countless other ways. A special thank you to our foundational sponsor Core11 and our Gold Sponsor Volinga, whose support has helped make the current state of the software possible. Thank you as well to every donor and to all of our new Bronze Sponsors. Looking ahead to v0.6 For the next major release, work will focus primarily on stability and user experience. This includes improved cleanup workflows and the ability to modify training parameters while training is in progress. I would also like to introduce a native .licht project format that allows users to save and restore their complete editor state. You can find links to our main sponsors below. Please also visit our website to discover all our Bronze Sponsors. Hint: We do not yet have a Silver Sponsor or Platinum 😉

MrNeRF

26,219 Aufrufe • vor 2 Monaten

Remember when we as football fans had to rely solely on paper draft guides, sports radio rumors, and gut feelings to predict draft day decisions? Excited that fans now have access to the NFL's Draft IQ powered by Amazon Web Services ( – the most sophisticated tool yet for following the NFL draft and your favorite team's strategy. Draft IQ is built on Amazon QuickSight, our cloud business intelligence service that makes it easy to analyze and visualize massive amounts of data. QuickSight processes real-time data to give fans unprecedented insight into team decision-making, updating the entire draft landscape every five minutes. You can explore team needs, draft capital, and front office tendencies through personalized team dashboards, plus get AWS-powered machine learning predictions about potential trades and picks. During draft week, fans can track picks, prospects, and Next Gen Stats in real-time. We're also introducing Amazon Q Business integration, our generative AI-powered assistant. Q Business leverages large language models to understand and respond to natural language queries, allowing fans to ask detailed questions about draft prospects, team strategies, and historical draft data. It can provide AI-generated insights based on the same historical Next Gen Stats research data that powers Draft IQ, giving fans a new way to engage with the draft experience (check out the example below). Can't wait to see what stories the data tells us as teams make their selections and excited to dig into the Giants' data myself :)

Andy Jassy

102,921 Aufrufe • vor 1 Jahr

Acclaimed metal band BABYMETAL has signed to Capitol Records, the first Japanese artist to sign a frontline deal with the label. The group will release new album METAL FORTH, globally via Capitol Records on June 13. “BABYMETAL’s groundbreaking sound and compelling artistic vision have not only cultivated a worldwide following, but have also demonstrably shifted global music culture. We at Capitol Records are privileged to join them in this next chapter as we continue to amplify their international reach and influence with the upcoming release of Metal Forth.” – Tom March, Chairman & CEO, Capitol Records “This year, BABYMETAL celebrates its 15th anniversary and embarks on an exciting new chapter. With Capitol Records as our global partner, the sound of BABYMETAL will resonate across the world as we take on bolder, more dynamic endeavors than ever before. Stay tuned for what’s to come.” - Key “KOBAMETAL” Kobayashi, Producer & Manager of BABYMETAL CEO, BABYMETAL WORLD, LLC As BABYMETAL celebrates its 15th anniversary this year, the band has mapped out a world tour that will include numerous firsts. In May, they’ll embark on their first-ever headline arena tour in the UK and Europe, playing 12 shows across eight countries and concluding at The O2 Arena in London. BABYMETAL is the first Japanese group to headline a show at this iconic venue. The band will kick off its biggest North American tour yet on June 13. The 24-date run will include a June 24 show at The Theatre at Madison Square Garden in New York City. Check for more information. Arena shows in Japan and a tour of Asia will follow.

BABYMETAL

283,131 Aufrufe • vor 1 Jahr

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 Aufrufe • vor 8 Monaten

A Letter to Our Community: The Road Ahead for Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷

Axis Robotics

27,858 Aufrufe • vor 8 Monaten

AI has had exactly two scaling axes that worked so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inference

Sasha Malysheva

12,064 Aufrufe • vor 25 Tagen