Загрузка видео...

Не удалось загрузить видео

На главную

llama3 8B (not quantized) running on an heterogeneous home cluster made of: - iPhone 15 Pro Max - iPad Pro (not sure which version XD) - MacBook Pro ( M1 Max ) - NVIDIA GeForce 3080 (not visible in video) - 2x NVIDIA Titan X Pascal Very soon also...

304,072 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 9

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

I also want to take this chance to announce that I'm about to open a Discord where communities from different projects on my github can meet and possibly start some interesting knowledge exchange, so stay tuned!

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

Just pushed a couple of commits that entirely decoupled the Cake system from the specific model. This in theory can now be used with any generative model ... nice things ahead

Фото профиля Alberto Fuentes (e/acc)
Alberto Fuentes (e/acc)2 лет назад

evilsocket doing AI inference code wasnt in my 2024 prediction list; the AI community earned one of the best in cyber back in the days 🙌🙌

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

thank you! still here trying to work on new stuff :D

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

If anybody wants to help, the iOS app source code is available here -> it's a super simple app (the logic is implemented in rust elsewhere, the app is just a launcher), it would be nice to make it prettier and possibly show in the main view what the library is logging to stdout ... it should be simple enough, i just don't know swift/swiftui :D

Фото профиля Yohei
Yohei2 лет назад

This is a cool project! Curious to see where you take it.

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

thank you!

Фото профиля ColorCoded (Road to Broke Arc)
ColorCoded (Road to Broke Arc)2 лет назад

I have no idea what any of those words mean but I shared

Фото профиля Simone Margaritelli
Simone Margaritelli2 лет назад

appreciate the sentiment buddy!

Похожие видео

Nvidia has just announced Alpamayo 2 Super, an open 34 billion parameter reasoning vision-language-action model designed to accelerate the development of autonomous vehicles. This new model combines the NVIDIA Cosmos 3 Super reasoning model with a 2 billion parameter diffusion-based action expert model, and is post trained with reinforcement learning. The model can return multiple outputs: future trajectory plans, reasoning traces, grounded answers to questions about the scenes, and auto label generation. The model weights are now available for anyone to download on Hugging Face, and the inference code has been posted to GitHub. Distilled models can be deployed commercially without any further permission from Nvidia, and model outputs carry no license conditions. Automakers can distill down a compact version of this model that can run on the Nvidia computer in the car. Major kudos to Nvidia and Jensen Huang for advancing the state of the industry by releasing this as an open model with permissive licensing. Jensen isn't just paying lip service to the idea of open models, Nvidia is actually contributing to the ecosystem — and it's great for their business, because it helps sell more Thor computers that go in the car. Anyone can go download the model and play with it. If you do, let me know what you think. Personally I think it's so cool that we have open weights models that are this advanced, for anyone to download.

Whole Mars Catalog

45,595 просмотров • 1 месяц назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 просмотров • 2 месяцев назад

This is big. NVIDIA and Apple just unlocked the next level for Vision Pro with CloudXR. Here’s what you need to know: It streams from a PC or the cloud directly to your Vision Pro. No cables. Up to 4K at 120fps. The technology is called dynamic foveated streaming. Your eyes only see in full resolution at the center of your gaze. CloudXR tracks exactly where you’re looking and delivers maximum resolution there. Everything in your periphery gets optimized. The stream stays efficient without you ever noticing a difference. Why does this matter? Vision Pro is already the most advanced spatial computer ever made. But standalone processing has a ceiling. There are workflows that need more compute than any headset can carry. CloudXR removes that ceiling by connecting Vision Pro to the full power of NVIDIA RTX in real time. This is not a workaround. It is a native visionOS integration. Also worth knowing. Gaze data never leaves your device. Not to the app. Not to the server. Developers get the full performance benefit of foveation without ever touching your raw eye tracking data. Privacy built into the architecture. Three industries are already on this: Kia, Rivian, and Volvo are running 1:1 scale design reviews with photorealistic accuracy. Full size vehicle models evaluated in spatial computing before a physical prototype exists. That is a fundamental shift in how design works. Foxconn is walking factory floors digitally, optimizing facilities before construction begins. Switch is managing data center infrastructure through a full digital twin. Companies using this approach are reporting up to 30% improvement in development processes. And sim enthusiasts finally get to cut the cord. iRacing and X-Plane 12 are the first titles. Full GeForce RTX power, streamed wirelessly to Vision Pro, inside your physical space. For developers: one Xcode template, one codebase, deployed across Vision Pro, iPhone, and iPad. And multiple headsets can share the same streamed environment at once. Some fully immersed, others on a tablet. That collaborative layer is what enterprise has been asking for. Coming this spring with visionOS 26.4. It’s wild to think about what this unlocks. The future of spatial computing is incredibly bright. Live from GTC. More coming soon.

Justin Ryan

43,685 просмотров • 5 месяцев назад

Etched came out of stealth at $800M and by lunch X had NVIDIA in the ground We do this every few months. A chip launches, the deck says killer, the timeline holds a funeral, and NVIDIA closes green anyway Etched hardwires the transformer into silicon. That is where the speed comes from, nearly the whole die on one job instead of the ~30% a GPU uses. It is also the trap. The day that chip tapes out is the best it will ever be. You cannot patch it. You burned progress into a wafer and now pray the field stops moving NVIDIA made the opposite bet. Same board, faster every quarter in software. Dynamo is pulling more tokens per watt out of the same rack, on version 1.0 The depreciation risk the bears aimed at NVIDIA for two years does not live at NVIDIA. It lives here, on the chip built to bury it Etched is not a fraud. It is a niche tool priced like a general one, and $800M is not enough to run a frontier supply chain. The rest get bought on the next down cycle Bury the lead, not the leader. Full case with Jack Farley and Max Wiethe on MTS And special thanks for Baseten for the cool T Shirt! Chapters 00:00 Switching from bonds to semis 00:33 What Etched actually is 01:31 Faster and cheaper, but how much HBM 02:56 Maturing market, not an NVIDIA killer 05:01 $1B in contracts and a Taiwan factory 05:20 Why these startups all get absorbed 06:48 Tiered inference and the obsolescence trap 09:39 Etched vs TPUs and Trainium 12:17 Is the CUDA moat weakening 13:21 Co-design, squeezing every token per watt 14:37 NVIDIA is a software company that sells a chip 14:58 Who is NVIDIA's most dangerous competitor 16:23 The NVIDIA killers, ranked 18:19 A rich man's game 18:43 AMD's MI500 vs Rubin Ultra 20:19 The neocloud business decision 22:46 Lightning round, Rambus the toll on HBM 25:11 The CXL run-up on Astera, Marvell, Credo 26:22 Use AI less, go to the booth 28:10 EDA is not dead

Ben Pouladian

29,969 просмотров • 2 месяцев назад

This soldiering training is the most impressive and immersive I've ever tried. It is 10x better than a YouTube tutorial video. It really allowed me to see the procedure from all the points of view, and even get super close to get the details. It is a collaboration between @gracia_vr and Imperial College London: they recorded a soldering session with Gaussian Splatting Videos (4DGS), so that you can enjoy it from your VR headset. You can see the action happening in front of you; you can pause and re-watch what you need, change the point of view, get closer, get more distant. And the quality with which it has been shot is impressive: I enjoyed this experience with my DELL Pro Max Tower T2 with NVIDIA Pro RTX 6000, and I could really see all the small details of the PCB that was being soldered! I was really impressed by it. But I also noticed some drawbacks. First of all, some scenes have artifacts that make seeing the details of the PCB hard. I think when it comes to training involving small details, the capture and reproduction of the splat should be flawless. Then, as much as I loved it as a passive experience, I would have liked to have also some sort of practice session in VR. The power of VR is to let you learn by doing in full safety, and this kind of training does not exploit it. But maybe the best way to enjoy it would be in MR, where you have this training video close to a real workbench where you do the actual soldering while following the tutorial. Gaussian Splat can really revolutionize training: this kind of recording is much better than any flatscreen experience, and more accurate than any 3D CGI recostruction. I suggest you give it a go at this experience in the Gracia app (it's free). Then let me know your impressions! #VirtualReality #training #DellProPrecision #GaussianSplat #technology [Disclaimer: I'm a DELL Pro Precision Ambassador, and this is why I mentioned the model of my PC. I have been given a PC to do cool tests and share my results on social media. No monetary compensation or sales affiliation is part of the collaboration]

TonyVT SkarredGhost

38,264 просмотров • 24 дней назад