🟢 News: GPU compute in the browser is finally... real. WebGPU now ships by default in Chrome, Firefox, Safari, and Edge—not a polyfill, not behind a flag. You can run LLMs client-side. Transformers.js and ONNX Runtime already ship WebGPU backends. Eight years of spec work. No more asterisks. 🔗 #WebGPU #GPU #Chrome #Safari #Firefox #Edgeshow more

WebGL / WebGPU
12,788 Aufrufe • vor 7 Monaten
WOW! 🤯 Language models are becoming smaller and more... capable than ever! Here's SmolLM2 running 100% locally in-browser w/ WebGPU on a 6-year-old GPU. Look at that speed! ⚡️😍 Powered by 🤗 Transformers.js and ONNX Runtime Web! How many tokens/second do you get? Let me know! 👇show more

Xenova
12,557 Aufrufe • vor 1 Jahr
Getting a bunch of questions about this! It's a... full gpu rendering pipeline with built in shader editor for Rive, which can interleave 2D + 3D. Think WebGPU (minus compute, for now) that runs anywhere, not just the browser. You can implement anything at all. Cell shading. Background blurs. Deformation effects. And yes, even Rive vector components as textures. Shaders are precompiled and ship with your .riv. Minimal runtime size impact. Still WIP.show more

Guido Rosso
22,346 Aufrufe • vor 5 Monaten
we just added Apple Pay support in Chrome and... non-Safari browsers on Dodo Payments. you can now use Apple Pay on desktop too. it opens a QR popup, you scan it with your iPhone, and pay. no Safari required. most payment providers still haven't shipped this. they're stuck on older Apple Pay versions that only work in Safari. we updated to the latest spec and built the cross-browser flow from scratch. this is live now for all merchants. no integration changes needed.show more

Ayush Agarwal
31,066 Aufrufe • vor 3 Monaten
Fable 5 built this real-time ocean scene directly in... the browser using Three.js, WebGPU and TSL, with no engine and no baked assets. The interesting part is not just that the water looks good. The whole system is being simulated: wave spectrum, GPU FFT, foam generation and refracted-ray caustics, all running live. This is exactly where frontier models start becoming genuinely useful for technical graphics work. Not by replacing years of rendering research, but by compressing the path from complex math to a working interactive result. A small, focused scene, but technically very serious. Cc: Adem Vessellshow more

Token Gremlin
21,731 Aufrufe • vor 1 Monat
Not sure if this is old news, but I... just discovered that in the Meta Quest browser you can longpress any image on any website and turn it into a 3D photo in a couple of seconds. I love it! Now I wish every image online worked like this by default.show more

Morten Haulik ᯅ
13,017 Aufrufe • vor 6 Monaten
Free NVIDIA GPU with 16 GB VRAM GPU for... Running Local LLMs! If you want to master local LLMs but you're waiting until you can afford a $1,500 GPU, you're honestly not going to make it. The open source AI ecosystem is moving way too fast for you to wait on your budget to catch up. Especially when you can build a bleeding edge inference engine from scratch right now, completely for free. You don't need a heavy local rig to start. Google is literally letting you use an enterprise grade NVIDIA Tesla T4 GPU for $0/hour. At standard cloud computing rates (~$0.20/hr), Google Colab’s 4 hour daily free tier hands you roughly $24 worth of data center tier GPU compute every single month. And most people just waste it. Let’s talk about the hardware you get access to for free. The NVIDIA Tesla T4 is an absolute workhorse: - Architecture: NVIDIA Turing (TU104) - VRAM: 16GB GDDR6 (320 GB/s bandwidth) - Compute: 320 Tensor Cores | 2560 CUDA Cores - Performance: 130 TOPS INT8 | 8.1 TFLOPS FP32 - Power: Sipping energy at a max 70W TDP This is the exact same hardware I used to run DeepMind's Gemma 4 26B A4B QAT MoE at a 250,000 context window without a single Out Of Memory (OOM) crash. If you have a web browser and 10 minutes, you have everything you need. I’ve put together a fully documented, cell by cell Google Colab notebook that teaches you exactly how to do this. Here is what the notebook actually teaches you: - How to provision an Ubuntu Linux environment with CUDA 13.0 and verify your driver stack. - How to pull the source code and compile the latest llama.cpp C++ binaries from scratch, specifically optimizing the build for your exact GPU using the -DCMAKE_CUDA_ARCHITECTURES=native flag. - How to directly download quantized local LLMs (GGUF format) straight from HuggingFace using the CLI. - How to manage 16GB VRAM limits, offload neural network layers to the GPU, and push massive context windows. Compile raw llama.cpp, ollama run a model, or spin up the LM Studio CLI. Pick whatever stack you are comfortable with. just start building. No hardware. No credit card. No excuses. Bookmark this post right now so you don't lose the tutorial. Even if you don't have time to run it today, you are going to want this workflow in your engineering toolkit. The link to the free Colab Notebook is in the comments below. Lemme know if you need more tutorials like this.show more

Alok
178,744 Aufrufe • vor 2 Monaten
YOM 🤝 Onicorn We’ve partnered with Onicorn the exclusive... decentralised platform connecting the best investors, KOLs and companies in Web3. Onicorn is a curated network where serious capital, top KOLs and real projects actually meet no noise, no tourists. They only work with the best, and we’re proud to be in that room. Because YOM is building something the next internet needs: a decentralised GPU edge network powering real-time cloud gaming and AI compute. Idle GPUs turned into a global engine — cheaper, faster, and closer to the player than the hyperscalers can manage. Two teams with the same standard: quality over noise. This is just the start.show more

YOM
15,739 Aufrufe • vor 3 Monaten
BrowseWeb3 is Live: Download the Web3-Powered Internet Browser by... LayerAI 🧬 The native browser for web3 on LayerAI - BrowseWeb3 - is now live for Windows Desktop. Integrating major ecosystem products to create a powerful Internet product. 👉 Download: 🧬 New Era of the Internet: Browse & Earn BrowseWeb3 is the LayerAI's first development catered for the specialized browser market, substitute for Google Chrome or Safari. (1) All essential web3 products in one browser (2) Incentivized browsing via upcoming AI2Earn rewards (3) Enjoy private browsing with Native VPN (4) Less ads, more focus: integrated ad, cross-site tracking and cookie blockersshow more

LayerAI | AI2Earn
125,376 Aufrufe • vor 2 Jahren
The thing no one noticed: There’s a pricing lag... between betting on Bitcoin and trading Bitcoin perps. That gap usually disappears once market makers close it. Until then, you can watch both markets side by side, see the dislocation in real time, and trade it. This is exactly the kind of edge most traders never even look for.show more

Alex Mason 👁△
43,926 Aufrufe • vor 1 Monat
THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T... RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.show more

RetroChainer
11,100 Aufrufe • vor 1 Monat
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 Aufrufe • vor 2 Monaten
This Nvidia GPU farm sits in a spare room... and prints $18,000 a month Ten cards running 24 hours a day, seven days a week. The setup cost $120,000 to build but it paid for itself in seven months. It does not mine crypto. It rents compute to AI companies that need processing power right now. Companies pay by the hour and the demand never stops. At full capacity the farm pulls $18,000 a month after electricity costs. The owner does not touch it. It just runs. Nvidia GPUs are the most in-demand piece of hardware on the planet right now. The companies that figured this out two years ago are already sitting on serious passive income. The barrier to entry is high but the people inside are not leaving. Follow if you want to understand where the real AI money is actually going.show more

winkle.
22,626 Aufrufe • vor 3 Monaten
Over the last 2.5 years, we've secretly assembled a... team of some of the brightest minds across cryptography, GPU programming, and ASIC design to build an internal Jolt Pro team. We surpassed our wildest dreams of what was possible, having already created a 1.61 GHz proving cluster we call a cell. Jolt Pro scales to infinity, the number of cells you can use in parallel is only limited by the size of the datacenter. This work culminated in our demo last night, where we verified a month of Ethereum in just 30 seconds. A special moment and a good start. But so much more to come.show more

raz
636,068 Aufrufe • vor 6 Monaten
A good technical LLM interview question: Your LLM chatbot... takes 12s before it generates the first token, and the users are complaining. So you move the model onto a GPU with 3x the computing power. The time to first token barely improves. Why did this happen? (answer below) Latency in an LLM app is a placement problem disguised as a model problem. If you profile the 12 seconds, the model's prefill itself may only account for around 1.5 seconds of it. So halving the prefill step saves just 750ms out of 12000, which is under 7%. The rest is spread across stages that never touch the GPU. The request first travels to whatever region the app runs in, and a cross-continent round trip could cost over a second before any code executes. Then the request handler starts. On a container-based serverless platform under load, this adds several seconds of cold start, paid before auth, rate limiting, or prompt assembly even begins. Retrieval adds its own hop, and the response streams back across the same distance. Optimizing a stage that was already fast cannot alter the latency that's majorly affected by other stages. Those other stages are slow for a structural reason. An LLM app runs two workloads that want opposite machines. - The request path is short, spiky, and needs to sit close to users - Inference is long-running, GPU-bound, and billed hourly, whether requests arrive or not. So the actual decision is not which model to run, but where each of these two workloads runs. There are three options, each with its own tradeoffs: > A dedicated GPU box removes inference cold starts, but it bills around the clock and lives in one location, so distant users wait out the round trip on every request > Container-based serverless scales to zero, but the request path pays a cold start, and most of these platforms have no GPU behind them. > Edge runtimes start in under a millisecond, because a WebAssembly module carries no OS or container image to boot. They handle the request path well and cannot hold a model. So the answer is not to pick one, but to split the app across two of them. The request path runs close to users, and inference runs on a dedicated GPU it calls into. That also explains the failed upgrade. More compute made a stage that was already fast faster, and left the 10.5 seconds around it untouched. To actually learn how it's done in practice, Akamai's GitHub has a reference implementation for each half. - vllm-on-lke serves Qwen2.5-7B-Instruct behind an OpenAI-compatible endpoint on one RTX 4000 Ada GPU in Linode Kubernetes Engine, with Terraform creating the cluster, both firewalls, and the GPU operator in one apply. - akamai-functions-llm-chatbot covers the front, where a WebAssembly API checks a KV cache and only calls the GPU-backed instance on a miss. Both are available on Akamai’s new Developer Hub, alongside their tutorials and code samples. It also links to Edge Case, their Discord, where four developer advocates architect and deploy a production app live every other Wednesday. If you create a new Akamai Cloud account, you can also get $300 in credits for joining. Join here: That said, this post treats generation as a single 1.5s block, but that block has its own structure, and knowing it well tells you whether a model is slow to start or slow to stream. I wrote a first-principles walkthrough of it, covering the prefill and decode split, KV caching, and where the time actually goes inside each one. Read it below. Thanks to Akamai Cloud for partnering today!show more

Avi Chawla
21,423 Aufrufe • vor 16 Tagen
A guy in Vancouver built an entire operating system... inside a Chrome tab. By himself. Over six years. His personal website is the OS. You open and a Windows-style desktop loads. File explorer. Start menu. Taskbar. You can drag in a zip and extract it. You can play DOOM. You can play Quake III Arena. You can boot Linux from an ISO. You can run Stable Diffusion locally for image generation. You can open a Python terminal. You can edit code in Monaco, the same engine that powers VS Code. All of it runs in your browser tab. Nothing installs. His name is Dustin Brett. Self-taught engineer. Father. Husband. 4,473 commits. All his. He had to swap the Windows icon for the π symbol because of legal pressure. The repo has 12,883 stars. MIT license. His hosting bill is one dollar a month. A single Cloudflare CDN does the rest. This is what the open web was built for. (Link in the comments)show more

Nav Toor
72,629 Aufrufe • vor 2 Monaten
“Wait until we reach port” is not a great... troubleshooting plan. Hyundai Glovis put Starlink on 45 of its 47 ships so crews can get technical help in real time, even when they’re far out at sea. Now they want to run AI predictive maintenance and smarter ship operations over that same connection too. That means catching problems before they become expensive ones and keeping more of the ship connected to the people and systems that can actually help. And yes, calling the family from the middle of nowhere gets easier too. Starlink SpaceX / Writer: Annette, Grok Imagine Designer: Jannéshow more

Mario Nawfal
26,997 Aufrufe • vor 8 Tagen
Behind every OptimAI Node is a mission far greater... than rewards: 🔸To build a decentralized Reinforcement Data Network 🔸To unlock Agentic AI for everyone—not just a privileged few 👉OptimAI Edge Node: Our architecture is now in motion: 🔸OptimAI DePIN to power decentralized infrastructure 🔸OptimAI DeHIN to amplify collective intelligence 🔸Reinforcement Data Layer to train agents smarter, fairer 🔸Compute Layer to enable real-time AI at the edge 🔸OptimAI Chain to govern everything, transparently and efficiently OptimAI Agent Studio, Agent OS, Data Engine, Compute Engine: all part of what we’re building next. This is a network built by people, for people. And you’re not just early, you’re essential. Let’s keep mining, contributing, earning rewards in return, and rewriting the future of Decentralized AI. The future is Agentic. The fuel is your data. The power is decentralized. One Node. One Data. One Agent at a Time. #BUIDL with us!show more

OptimAI Network
60,345 Aufrufe • vor 1 Jahr
Elon Musk gave the entire entertainment industry its expiration... date, and he is the one building the thing that kills it. Musk: “My guess is that we see the first compelling half hour, pure AI show next year.” Next year. A complete show generated entirely by AI. No writers. No actors. No cameras. No sets. No crew. No studio. Just a prompt and enough compute to render a reality that never physically existed. And shows are the easy part. Musk: “I say probably we’re maybe three years away from AI does the whole video game.” A show plays the same way every time. A game has to generate a living world that reacts to every decision in real time across every single frame. That is a fundamentally harder class of problem. And Musk put three years on it. Right now a single AAA title takes seven years and half a billion dollars across thousands of engineers and artists just to ship it. Musk is describing a world where one person types a paragraph and gets something comparable. The entire value proposition of a multi-billion dollar industry lives inside that gap. And it closes in thirty-six months. But the prediction is not the story. The person making it is. This is not an analyst speculating from the sidelines. This is the man building the largest AI compute clusters on the planet. The man who built xAI from zero in under two years. The man stacking hundreds of thousands of GPUs into facilities designed to do exactly what he is describing. When Musk says three years, he is not guessing about what someone else might eventually ship. He is reading you a delivery date off his own roadmap. Every media company on Earth is valued on a single assumption. That quality content is expensive and difficult to produce at scale. That one assumption is the structural foundation underneath every studio, every network, and every publisher in existence. Musk is dismantling it with raw compute. The studios still parading thousand-person production teams are not demonstrating strength. They are advertising the exact cost structure that one person with a prompt and a GPU allocation is about to make irrelevant. And it does not stop at entertainment. If AI can generate an interactive world that responds to human input in real time, it can generate anything. Advertising. Architecture. Training simulations. Product design. Every industry built on humans manually constructing visual experiences frame by frame is sitting on the same countdown Musk just read out loud. Now zoom out. Because this is not just an industry story. For the entire history of human civilization, the distance between imagining a world and actually creating one required thousands of people, millions of hours, and billions of dollars. That distance built Hollywood. That distance built the gaming industry. That distance made content scarce and studios powerful. Musk is collapsing that distance to zero. When the gap between imagining something and it existing disappears, every business model built on the difficulty of creation disappears with it. That is not disruption. That is a full inversion of how human beings create. Musk did not make a casual prediction on that podcast. He told you what he is building. He told you the timeline. And he told you which industries do not survive it. The entertainment industry is still debating whether this future is real. Musk is not part of that debate. He is building. And he just told you the delivery date.show more

Dustin
22,458 Aufrufe • vor 1 Monat
$IREN "we haven't disclosed the specific amount of GPUs"... 1. 🤮 reminds me of $NBIS 2. Setting a terrible precedent here for future deals 3. Making it purposely difficult, to not let analysts properly value your 2027 revenue 4. Increasing the polarized view on IREN by the market However: "approximately 60MW of air-cooled Blackwells" 1. You typically don't talk about gross capacity in a deployment like this 2. If it would be gross capacity, the GPU hour rate at IT level would be crazy high (at PUE 1.2, $680m / 50 = 13.6m/MW) 3. At 60MW IT load, and ~14kW draw at DGX server level, we can get to ~4,286 DGX systems with 8 GPUs per. 4. Based on this we can conclude that 60MW of IT load can run approximately 34k DGX B300. 5. 34k DGX B300 at $680m/yr, would represent a GPU hour price of $2.28 Now this is the problem with not disclosing your GPU quantity. You purposely make your business model look bad, because by approach, you get to a GPU hour price that would imply a payback period of 4 years, where only the last year of the contract is 100% margin. But of course, we can also take "the glass is half full" approach. IREN has ordered 50K B300s from Dell. They have 2 purchase orders for this, 1 between Dell Canada and IE CA Leasing Ltd for 4 phases, and 1 between Dell USA and IE US Hardware 1 Inc (amended from IE US Hardware 4 Inc on April 27, 2026). The order for Canada is divided in 4 phases, and are going to Mackenzie for 80MW of gross capacity, which happens to be 4 buildings of 20MW. The order for Childress is divided in 2 phases, and are going to DC35 and DC36, (as depicted in the earnings presentation) and those are 50MW gross. The purchase price of the order for Childress was $1.2B, and for Canada it was $2.3B If we go with 50,000 B300s for a total of $3.5B then $1.2 would represent 34.285% of the 50,000 GPUs, or 17,140 B300s rounded down. For this calculation I will consider that $IREN will deploy 17,140 GPUs in 50MW gross capacity in DC35 and DC36 of block 3 in Childress.. That would imply at 1.2 PUE, IREN can run 17,140 B300s in 41.67MW IT load. Now by that ratio, they can run 24,680 GPUs in 60MW IT load — a massive difference with 34k units through the Nvidia DGX reference calculation. If common sense is applied, you can still get to 2 completely different outcomes, that show a difference of more than 9k GPUs. The GPU hour rate at 24.68k GPUs would be $3.145 per B300, as MASSIVE difference from the earlier calculated $2.28. Sure, the DGX system may be a factor here. And I'm sure that the reality is somewhere in the middle. But I personally hate this as an investor, to be unable to calculate profitability on unit economic basis. After all, contracts are signed on a $/GPU hour basis. Why hide this from your investors? Not being able to calculate payback periods, unable to calculate ROIC. And most importantly, we cannot properly assess the $NVDA deal on a contract basis. I really hope the payback period of this contract is not 4 years. I want the glass to be half full, but by starting to censor the purchases, IREN is taking a step in the wrong direction. Not a fan of this.show more

Frans Bakker
148,167 Aufrufe • vor 3 Monaten
The Visual Studio Code insiders version that just shipped... and will ship in the next few days will come with an insane amount of new capabilities. A few highlights: - You can now run sub-agents in parallel. Yes, really. I even attached a video. - Major UX improvements for sub agents, especially visible in the chat window - A new search tool wrapped as a sub-agent that iteratively runs multiple search tools: semantic_search, file_search, grep_search Which connects nicely to the point above: multiple searches running in parallel, efficiently and fast - Anthropic’s Message API is now enabled by default - You can choose the model for the cloud agent (three available, all premium) - Extended thinking support when using the Claude cloud agent This is part of the broader multi-vendor cloud support under AgentsHQ I wrote about a few weeks ago - Tasks sent to the background agent (basically the CLI tool) now always run in isolation, each with its own git worktree - In a multi-repo workspace, assigning a task to a cloud agent prompts you to choose the target repo Same behavior when opening an empty workspace with no repo - Support for building an external index for files not supported by GitHub’s default indexing - UI/UX improvements for starting new sessions and switching between local / background / cloud agents - Skills are now first-class citizens, just like prompt files, with better UX indicating when a skill is loaded - Improved API for dynamic contribution of prompt files New V2 includes skills as part of the model. Curious to see the extensions that will leverage this - Finally, initial support for showing context usage percentage per session - Skills are enabled by default - Resizable chat window and session view. Small thing, but it was driving me crazy 😁 - A new integrated browser meant to replace the old simple browser Maybe the beginning of real browser use? - Better UI/UX for token streaming in chat - Ability to index external files not supported by GitHub There’s a lot more. Some of it hasn’t fully landed yet, but everything that has is already in Insiders. The next stable release should drop in early February. As usual, I’m just shocked by the volume of features this team ships every month. After the holiday slowdown, this one is shaping up to be a wild release.show more

Oren Melamed
29,555 Aufrufe • vor 7 Monaten