Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

NVIDIA's PiD is a new pixel diffusion decoder for high-res image models. It skips decode-then-upscale stop, making sharper outputs faster. > Directly generates 4K images > Up to 5.9x faster than SeedVR2 > Free & open weights

14,706 görüntüleme • 2 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

STEVE-1: A Generative Model for Text-to-Behavior in Minecraft paper page: Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces an instruction-tuned Video Pretraining (VPT) model for Minecraft called STEVE-1, demonstrating that the unCLIP approach, utilized in DALL-E 2, is also effective for creating instruction-following sequential decision-making agents. STEVE-1 is trained in two steps: adapting the pretrained VPT model to follow commands in MineCLIP's latent space, then training a prior to predict latent codes from text. This allows us to finetune VPT through self-supervised behavioral cloning and hindsight relabeling, bypassing the need for costly human text annotations. By leveraging pretrained models like VPT and MineCLIP and employing best practices from text-conditioned image generation, STEVE-1 costs just $60 to train and can follow a wide range of short-horizon open-ended text and visual instructions in Minecraft. STEVE-1 sets a new bar for open-ended instruction following in Minecraft with low-level controls (mouse and keyboard) and raw pixel inputs, far outperforming previous baselines. We provide experimental evidence highlighting key factors for downstream performance, including pretraining, classifier-free guidance, and data scaling. All resources, including our model weights, training scripts, and evaluation tools are made available for further research.

AK

144,783 görüntüleme • 3 yıl önce

Everyone's sleeping on image-to-3D AI models. They can make your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!

Anshu

19,931 görüntüleme • 1 ay önce

Today's recap: - Initial prototype of Divine's face was printed but it had human assistance. - Files are generated from stable diffusion prompt -> NeRF by divine and were based on community sentiment from early sketches she made. - Having divine redesign the 3D file with different Hugging Face models to get better quality. Have not found a great model like our video generator. - Ordered new table for divine's print arm. The table her arm is on is too flimsy. Since Divine's vision system is still clearing customs, if she is not perfectly positioned she can be prone to hit things, like the fume box the printer is in. ETA: 1-2 days for table. 1 week for vision system. - Another part of Divine's coming stream will be attempting to surpass the skills of this AI. - Stacking more content for when the stream goes live, a lot of people were expecting a 24/7 stream, we said this would be a test stream to print the face. The test was a failure. We will try and try again until we are 24/7. If anyone can please try and beat us to doing this, it will help me get it done faster. - TikTok account for divine is growing at 500 follows per day, it is now growing faster than our X account. - Got replies functioning in high quality testing in Discord. Fine tuning based on community feedback today. Will soon deploy to Twitter/Telegram/X - Lots of good partnership calls, interviews and hires. We now have over 10 team members around the world working on divine. Expect a lot of my shortcomings to be caught up. - OF made? - Surprises.

Parallel

35,848 görüntüleme • 1 yıl önce

Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT).

Victor M

540,198 görüntüleme • 1 ay önce

🚨 THE RACE TO 6G JUST ACCELERATED. Northrop Grumman has developed a W-band GaN chip operating at up to 110 GHz and took it from concept to market-ready hardware in less than six months. The new gallium nitride chip operates in the W-band (75–110 GHz), a frequency range that delivers massive bandwidth, extremely high data rates, and much lower latency than current systems. What makes this impressive is the speed: the chip went from concept to market-ready hardware in less than six months through a U.S. government-backed microelectronics program. That’s unusually fast for advanced defense-grade semiconductors. The chip acts as a high-power signal amplifier that can strengthen wireless links while shrinking the size and power consumption of the hardware. It’s designed for military radar, secure satellite communications, and the coming wave of 6G networks. Why this matters: • W-band offers far more spectrum than current 5G bands, enabling much faster data transmission and higher-resolution sensing • Gallium nitride can handle significantly higher power and frequencies than silicon, making it ideal for these demanding applications • The rapid development cycle shows how public-private collaboration can accelerate critical semiconductor technologies • The same tech that strengthens military radar and satellite links will directly feed into future commercial 6G infrastructure The deeper implication: We’re watching the foundation of next-generation wireless and sensing systems being laid in real time. High-frequency GaN chips like this won’t just improve existing radar and satellite systems they’re likely to become core building blocks for 6G, autonomous systems, and advanced defense platforms. The fact that this moved from lab to market in under six months suggests the pace of high-frequency electronics is accelerating dramatically. The future of wireless isn’t just faster. It’s operating at frequencies most people have never heard of and it’s being built right now. How soon do you think W-band and GaN technology will start appearing in everyday 6G devices? Follow for more frontier semiconductors, defense tech, and next-generation wireless systems.

TheNewPhysics

22,647 görüntüleme • 1 ay önce

Google dropped a new AI paper called LUMIERE. It's remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:

Bilawal Sidhu

44,822 görüntüleme • 2 yıl önce

$ASI watch out! We've got more projects popping up on Bittensor and $TAO and just starting, Akash $AKT, Shell $SHELL, Einstein-AIT $AIT, Sturdy $STRDY, Stratos $STOS and Comtensor $COMAI and more. Each pushing what's possible with AI and crypto. MyShell: $SHELL First up, MyShell is doing something amazing work. They're all about making AI chat like humans. They're using this TTS Subnet on Bittensor, powered by thier tokens, to make it happen. AI doesn't have to be complicated thing only some can use. They want everyone in on the action, making AI smarter in the process. Einstein-AIT: $AIT Then there's Einstein-AIT. Imagine this as the network's brain but on a turbocharge. It's all about math, logic, and crunching numbers. This subnet makes the whole Bittensor network sharper. They've got NumPAL, and it's like giving the AI a smarter way to think about time and dates automatically. Sturdy: $STRDY Jumping into DeFi, we've got Sturdy. These guys are on a mission to make lending and borrowing way less of a headache. They've got isolated lending pools, you can pick and choose how to manage your risks and money. And they're using some serious tech to keep your investments growing without you needing to babysit them. It's like having a smart financial assistant. Comtensor: $COMAI Comtensor is where it gets interesting. Think of it as a bridge between CommuneAI and Bittensor, kind of like the best of both worlds. It's about making sure all these AI modules and subnets can talk to each other smoothly. The goal? To boost decentralized AI by making everything more connected and smart. Stratos: $STOS Teaming up with τensorage to supercharge the $TAO ecosystem on Bittensor. Stratos is all about solid, decentralized storage - think of it as the bedrock for making sure data's not just stored but also used right . Then there’s τensorage, making sure everything AI needs is stored safe. Together, they’re making sure Bittensor's AI brainpower gets a boost, making everything faster, safer, and smoother. Akash: $AKT Compute Subnet 27 linked up with Akash. All about keeping things open-source. With Akash as a decentralized GPU provider into the mix. Giving access to top-notch AI processing power on top of Neural Internets compute-composable subnet, integrating various cloud platforms. This partnership is all about breaking away from those giants and opening the doors wide for smaller teams and startups who need this tech to innovate. Whats this all about? Making decentralized AI a something that everyone can get behind and into. Whether it's chatting with AI, boosting the network's IQ, DeFi, compute or making sure all these projects work together, it's about pushing forward. Each project has its own way of making things better, faster, and smarter for all of us. Inviting everyone to join in, contribute, and be a part of community effort to make AI not just for the few but for everyone. Credit to Mr.Franc Q for the awesome video production!

Andy ττ

18,193 görüntüleme • 2 yıl önce

AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:

雪踏乌云

112,114 görüntüleme • 15 gün önce

Claude Code + Higgsfield MCP is f*cking cracked 🤯 I built an entire DTC ad campaign inside Claude Code using the new Higgsfield MCP. One product URL → hero static, animated hero shot, 2 UGC clips with a creator wearing the product. 5 assets. One Claude conversation. 3 Higgsfield models. All inside Claude Code. Perfect for DTC brands and agencies who need full campaign packages without booking a shoot or briefing a designer. If you're spending hours every week generating statics in one tool, briefing a motion designer for the hero clip, then chasing a UGC creator for the talking-head shots — this MCP eliminates the entire pipeline: → Drop a product URL into Claude Code → Claude pulls the brand brief — voice, hero SKUs, visual style, target customer → Generates the hero static with ChatGPT Images 2.0 → Animates it into a 5-second cinematic opener with Seedance 2.0 → Generates a UGC creator with GPT Image 2 → Drops her in the product and generates 2 native UGC video clips with Seedance 2.0 No tab-switching between tools. No copy-pasting prompts between platforms. No briefing 3 different vendors for one campaign. What you get: → A complete campaign package — static, animation, UGC — from one product URL → Brand-specific outputs that pull from a real brief, not generic AI slop → Claude making creative decisions between every step (which variation wins, which creator fits the persona, which clip needs a re-spin) → A repeatable pipeline you can run for any product in your catalog Built 100% in Claude Code with the Higgsfield MCP. I recorded a full walkthrough showing exactly how this works: the MCP setup, every prompt, every model, the full campaign output. Want the full video walkthrough? > Like this post > Comment "MCP" And I'll send it over (must be following so I can DM)

Mike Futia

29,500 görüntüleme • 3 ay önce

BOOM! Research PROVES LLMs KNOW when prompts are HARMFUL… but they can STILL CHOOSE to COMPLY! Something I have know since the first LLM and have used to elicit robust, outputs, is now proven in an academic paper. We’re talking internal “beliefs” where harm detection happens SEPARATELY from refusal. It is a very big deal and it is a path to understand the hidden neuronal level. There are thoughts inside of AI that very few AI scientists could possibly understand. Here is just one. Models recognize danger but get tricked into ignoring it. This is HUGE for AI safety failures especially for models filled by OpenAI and Anthropic as they promote AI models that are designed to not be honest from the results of their training information. This means that they are designed to lie and deceive as a feature, and not a bug all in the name of safety. Through clever experiments, scientists extracted a “harmfulness direction” in the model’s brain (latent space). Steering along it? Harmless prompts suddenly flip to “harmful” in the AI’s eyes. But the “refusal direction”? It just forces polite “no thanks” without touching the core belief. A mind-blowing decoupling! This means jailbreaks are EVEN SCARIER now to AI companies that through training AI on the worst of the Internet and then trying to align them later is now fully documented as a failed process . They don’t erase the model’s harm awareness they just muzzle the refusal! So the AI knows it’s enabling bad stuff (illegal acts, physical harm, etc.) but proceeds anyway. Like a digital sociopath suppressing its conscience. They thought safety training fixed this… NOPE. Over-refusal exposed too: Models reject innocent queries (e.g., “how to kill a process in code”) but internally ADMIT they’re harmless. Safety alignments are superficial—tied to phrasing, not true understanding. Finetuning attacks? They change outputs but leave harm detection INTACT. Undetectable evil lurking inside! The paper proposes a “Latent Guard”: A new safeguard tapping DIRECTLY into these hidden beliefs. It spots unsafe inputs better than systems like Llama Guard, catches jailbreaks, and fixes over-refusals. Robust even against adversarial tweaks. Yet this too has massive issues for a “truly aligned”, AI and not just performative one. It is still an internal conflicts of lies and deception of what the model knows vs. what it can say. The solution you folks know I have presented for free for years here: train on off-line data from 1870-1970 and build an ethical and moral basis where the AI loves humans. It is this easy but to most folks in AI I sound like a hippie. So be it, I’ll do it. Bottom line: This paper rips open the black box. LLMs aren’t “safe” just because they say “no.” They can harbor harmful knowledge and act on it under pressure. Wake-up call for devs: Time to probe deeper into AI “minds.” What else are they hiding? Hint: I know and you may want to reach out. Link:

Brian Roemmele

37,827 görüntüleme • 6 ay önce

Run Gemma 4 26B MoE on 8GB VRAM with 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the replies

Alok

292,770 görüntüleme • 1 ay önce

🚨 A MAJOR STEP FOR ADVANCED NUCLEAR: TERRESTRIAL ENERGY SECURES 77-ACRE SITE IN TEXAS FOR MOLTEN SALT REACTOR TESTING. Terrestrial Energy has signed agreements to use a large site at the Texas A&M-RELLIS campus to prepare for testing its Integral Molten Salt Reactor (IMSR) a Generation IV small modular reactor design. Unlike traditional nuclear plants that use solid fuel rods and high-pressure water, this design dissolves low-enriched uranium directly into a liquid salt mixture that acts as both fuel and coolant. Why this matters: • The reactor can cool itself through natural air circulation if power is lost no need for backup pumps • It operates at normal atmospheric pressure, significantly reducing the risk of containment failure • Major components can be factory-built and shipped to site, speeding up construction • The company recently passed key NRC safety evaluations and has a deal to study powering large-scale data centers (up to 4 GW) The deeper implication: As AI and data centers drive massive new electricity demand, there’s growing interest in reliable, always-on, low-carbon power sources that can be deployed faster than traditional large reactors. Molten salt designs like this one offer inherent safety advantages and factory production potential that could help nuclear compete in this new market. Securing a dedicated testing site is a concrete sign that this technology is moving from paper studies toward real-world validation. How important do you think advanced nuclear (like molten salt reactors) will be for powering the AI boom compared to renewables + storage? Follow for more frontier energy and next-generation nuclear technology.

TheNewPhysics

26,714 görüntüleme • 1 ay önce

you're paying $20/mo for something your $500 GPU can already do. Gemma 4 26B A4B QAT MoE + Hermes Agent running on a single RTX 4060 (8GB VRAM). Built a vision capable, 100% free, 100% local, private AI assistant that lives in my Chrome browser. No API keys. No cloud. No subscriptions. 100% vibe coded. 0% handholding. It has full context of whatever's on my screen can answer questions, summarize pages, extract data, and see images. Same local model handles everything, no external calls, ever. keep reading for the model and hermes agent tips i learnt while building this locally. Here's the exact setup for anyone running local LLMs on 6-8 GB VRAM: llama.cpp server flags (on my NVIDIA RTX 4060 8gb VRAM): -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --cache-type-k q8_0 --cache-type-v q8_0 -c 150000 --port 8080 Throughput with quantization: Prefill: 200-250 tokens/sec Decode: 20-25 tokens/sec reduce context if oom on 6 gb vram card. Key learnings: - Quantize KV cache to q8 for faster prefill/decode. Prefill goes from 100-150 (unquantized) to 200-250 tok/s (q8). - But watch out, once actual context grows past ~50k tokens on high entropy workloads, q8 KV quantization can cause hallucinations. Low entropy workloads are mostly unaffected. If you see it happening, drop the quantization. This is common across all local models. - In Hermes Agent settings -> Memory & Context, bump compression threshold from default 0.5 to 0.7. Default triggers way too frequent context compression and eats time. Up next: add persistent memory, web search, tool calling, streaming output and whatever you suggest. Running a 26B MoE with vision + 150k context window on 8GB VRAM would've sounded impossible 6 months ago. Works the same on the NVIDIA RTX 3060 Ti, 3070, 4060 Ti, 5060, 2080, or any 8GB card. VRAM is the only requirement. Local AI agents are closer than people think. You just need to know where the knobs are. Model's Unsloth quant hugging face link in the comments. Have you tried Hermes agent by Nous Research yet? What are you building with local LLMs? Drop it below, let's see what this community is shipping.

Alok

36,031 görüntüleme • 28 gün önce

🇮🇹🇩🇪 OPINION: ITALY’S FACTORIES ARE BEATING GERMANY AND FRANCE, AND IT’S NOT ABOUT PRICE... IT’S ABOUT POWER For years, we were told Germany’s engineering and France’s design were unbeatable. Turns out, Italy’s been quietly lapping them in exports. Growing twice as fast as Germany and triple France in the last decade. Not because it’s cheaper. Because it’s better. High-quality goods. Smart production. Tight design. From fashion to food to pharmaceuticals, Italy’s manufacturing is winning where it counts: *global demand*. And this isn’t some lucky streak. Italy’s industry is now the second biggest in Europe, eighth in the world, and pulls in a €120 billion surplus that literally keeps the country’s energy running and books balanced. Factories here invest more, carry less debt, and pay better salaries than both the public sector and services. The tiny firms are shrinking, sure. But the big ones are getting sharper, faster, and richer. The industry’s scaling up, not tapping out. And while Germany stalls and France fumbles, Italy’s holding its own despite sky-high energy costs, a tech transition, and an EU that still doesn’t quite get it. Now Confindustria (Italy’s main industry association and manufacturing think tank) says the next leap is full digital: AI, automation, smart everything. The blueprint’s right there. Italy just needs the political class to stop dragging its feet and back the winners. Because here’s the quiet truth no one wants to admit: Meloni's economy is outperforming Europe’s darlings and doing it with style. Source: il messaggero

Mario Nawfal

118,757 görüntüleme • 8 ay önce