HOT! MiniMax-H3 Fun Controlnet Union dropped by Alibaba! all-in-one... control for MiniMaxH3 - Canny, Depth, HED, MLSD, Pose - inpainting - single 7GB checkpoint - guidance-distilled for fast 1-pass inferenceshow more

Wildminder
23,080 просмотров • 10 дней назад
Introducing MiniMax H3 by Hailuo AI (MiniMax) Our next-gen... open-weight multimodal video model, built for general intelligence beyond single-task generation. What makes it different: -Native multimodal understanding & generation -Precision editing & control -Commercial-grade output for film, ads, MVs, UI, game CG -Cinematic quality: native stereo audio, up to 2K/24FPS -Multi-asset reference + voice cloning = idea to finished product, in one flow. Try H3 now → #MiniMaxH3show more

COLAW
65,838 просмотров • 1 месяц назад
Got early access to MiniMax H3 and honestly... the... output speaks for itself. 🔥 Tested out a multi-shot cyber-fashion teaser using the new Omni-Reference feature to keep the character and gritty neon vibe completely locked in across the cuts. What blew me away is that this dropped in a single pass with native stereo audio and sharp text rendering built right in—no separate sound design or post-editing mess. What do you guys think? Check it out MiniMax Design (H3) #MiniMaxH3show more

Arcane Ai
44,177 просмотров • 1 месяц назад
Introducing MiniMax H3, our next-gen open-weight multimodal video model,... built for general intelligence beyond single-task generation. H3 understands text, images, video, and audio together, interpreting motion, sound, emotion, and cinematography as one unified creative language, delivering generation with fine-grained precision and cinematic quality. → Native Multimodal Understanding & Generation → Precision Editing & Control → Commercial-Grade Output: film, ads, MVs, UI, game CG → Cinematic Audio-Visual Quality: native stereo audio, up to 2K/24FPS Multi-asset reference for characters, motion, camera movement + voice cloning. Hailuo AI (MiniMax)'s H3 takes ideas from concept to production in one flow. Try it now: #MiniMaxH3show more

SARAH
97,618 просмотров • 1 месяц назад
Introducing MiniMax H3 Our next-gen open-weight multimodal video model,... built for general intelligence, not just single-task generation. Why it matters: 1/ Understands text, image, video & audio together, motion, sound, emotion, cinematography as one language 2/ Native multimodal understanding & generation 3/ Precision editing & control 4/ Commercial-grade output: film, ads, MVs, UI, game CG 5/ Cinematic quality: native stereo audio, up to 2K/24FPS Plus multi-asset reference + voice cloning. Concept to production, one flow. Try it: #MiniMaxH3 Hailuo AI (MiniMax)show more

Ricardo Momo
71,792 просмотров • 1 месяц назад
I've been testing MiniMax H3 in early access, and... it's one of the most capable AI video models I've used recently. What stood out to me most is the level of creative control. Instead of relying on a single prompt, I can combine text, images, videos, and audio to guide the final result much more precisely. A few features I've been enjoying: ✅ Native multimodal understanding and generation across text, images, video, and audio ✅ Omni Reference for more consistent characters, motion, and camera movement ✅ Precise editing that follows detailed instructions surprisingly well ✅ Commercial-grade output built for ads, branding videos, UI demos, game cinematics, and music videos ✅ Cinematic visuals with impressive audio quality I'm still exploring everything H3 can do, but my first impression has been really positive. Looking forward to sharing more examples as I continue testing. What would you create with MiniMax H3? MiniMax Design (H3) #MiniMaxH3 #HailuoAIshow more

Manish Kumar Shah
22,942 просмотров • 1 месяц назад
we sped up distributed inference by up to 5x... with decentralized speculative decoding. many don't realize that AI models normally generate text one single word at a time, waiting for the network after every word. speculative decoding changes this by using a "guess & confirm" system, similar to autocomplete. how it's done: 1. draft locally (the guess) instead of waiting for the network, a tiny, fast model on your device guesses the next few words instantly, without waiting for the network. 2. confirm remotely (the check) the massive remote model doesn't generate from scratch; it just checks the draft. it looks at the guesses in a batch and says "yes, yes, no." you get multiple words in the time it usually takes to get one. 3. adaptive logic dsd is smart. if the topic is creative, it lets the draft flow loose. if the topic is math or code, it checks more strictly. it balances speed and precision automatically so your inference almost feel instant. find out more: paper: blog:show more

Parallax
45,584 просмотров • 7 месяцев назад
The Future of AI Filmmaking Starts Here with MiniMax... H3 MiniMax Design (H3) #MiniMaxH3 Try here : Model: MiniMax H3 Aspect Ratio: 3:4 Resolution: 2K Prompt: Create a 40-second Hollywood-style cinematic film teaser for a sci-fi thriller titled "THE LAST SIGNAL." Scene 1: A lone astronaut walks across an abandoned lunar research station under a dark sky. Dust floats in slow motion. One red emergency light flashes. Ultra-realistic, IMAX quality, volumetric lighting. Scene 2: Deep underground, scientists discover an ancient metallic object buried beneath ice. Strange glowing symbols slowly activate. Cinematic camera push-in. Scene 3: A mysterious signal begins spreading across satellites surrounding Earth. Massive orbital structures light up simultaneously. Scene 4: Cities worldwide lose power. Giant holographic waves move through skyscrapers. Rain, smoke, emergency vehicles, cinematic chaos. Scene 5: A female scientist stares at thousands of floating holographic equations while whispering, "It's communicating." Scene 6: A colossal alien structure slowly rises from the ocean during sunrise. Enormous scale, cinematic drone shot. Scene 7: Fast montage: • fighter jets • exploding satellites • astronauts floating in zero gravity • underground bunker • frightened child looking at the sky • giant alien silhouette inside storm clouds Final Scene: The screen fades to black. Title appears: THE LAST SIGNAL Tagline: "They were never gone." Coming Soon. Style: Hollywood blockbuster, Denis Villeneuve-inspired cinematic language, ultra-realistic, 8K, anamorphic lenses, dramatic contrast, volumetric fog, realistic skin, global illumination, HDR, dynamic camera movement, cinematic color grading, premium VFX, emotionally intense pacing, realistic physics, Dolby-style atmosphere, epic scale.show more

Mr Farman Ai
14,175 просмотров • 1 месяц назад
Here's how to get 100% consistent product ads from... one seedance 2.0 generation. I did it all in a single chat using the Comfy MCP. The real control here comes from calling my existing workflows (shared below) instead of the agent improvising a pipeline. I directed the agent to call my ComfyUI workflow for cinematic product ads. I specified the close-up shot of the sprite animating, the bezel turn flipping the screen, the display changing to the time 10:04. Now for the consistent variations. The driving video does the heavy lifting but you need to get it right → depthanything v3 pass blended with canny edge lines to show the fine detail... it's why the tiny debossed logo is there → the initial sprite outline lived in those edge lines too, and every gen kept inheriting it. claude suggested a sam3 mask over the screen to hide it (s/o the agent) → with the screen masked, the new star sprite is just prompting: one gpt-image-2 still to generate a reference, one extra line in the seedance 2.0 prompt, and it animates oh and the whole process is a claude skill now.show more

rob - comfyui
20,216 просмотров • 1 месяц назад
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
54,264 просмотров • 22 дней назад
This Chinese developer launched Llama 70B locally on a... MacBook on a plane and for a full 11 hours without internet ran client projects. He was sitting by the window on a transatlantic flight with a MacBook Pro M4 with 64 GB of memory. WiFi on board cost $25 for the flight. He declined. No cloud API, no connection to Anthropic or OpenAI servers, no internet at all. Just a local Llama 3.3 70B on bf16 and his own orchestrator script. The model runs through llama.cpp. Generation speed, 71 tokens per second. Context around 60,000 tokens. Memory usage, 48.6 GiB out of 64. Battery at takeoff, 3 hours 21 minutes. And he gave the orchestrator this system prompt before takeoff: "You are an offline orchestrator running on a single MacBook. There is no network. The only resources you have are local files in /Users/dev/work, the Llama 70B inference server at localhost:8080, and a battery budget of 3 hours 21 minutes. Process the queue at /Users/dev/work/queue.jsonl (one client task per line). For each task: draft → run local evals → save artefact to /Users/dev/work/done/. Save context checkpoints every 12 tasks so you can resume after a battery swap. Stop only on empty queue or when battery drops below 5%." So the system knows exactly what resources it is running on. It knows it has no connection to the outside world for the next 11 hours. It knows it has finite memory and a finite battery. It knows the human will not intervene until the plane lands. The system runs in 1 loop. Takes a task from the queue, runs it through inference, saves the artifact, writes a checkpoint. Task after task, just like that. And only when the battery drops below 5% does the orchestrator automatically pause, waits for the laptop to switch to the backup power bank, and continues from the last checkpoint. Here is what the system actually writes in his log during the flight: "saved context checkpoint 8 of 12 (pos_min = 488, pos_max = 50118, size = 62.813 MiB)" "restored context checkpoint (pos_min = 488, pos_max = 50118)" "prompt processing progress: n_tokens = 50 / 60 818" "task 37016 done | tps = 71 s tokens text → /Users/dev/work/done/proposal_westside.md" Outside the window, clouds, blue sky, and no WiFi. On the tray, 1 MacBook, an open terminal on 2 screens, and an inference server on localhost. From what I have observed, this is the cleanest offline AI workflow I have seen in the past year: 11 hours of flight, $0 for WiFi, and the entire client queue closed before landing.show more

Blaze
1,841,161 просмотров • 4 месяцев назад
🚨 Do you understand what Claude just quietly dropped... while everyone was distracted? 1 million tokens. Let me explain what that actually means because the number alone doesn't hit right. > A senior engineer joins a company and spends 3 to 6 months just reading code.. Understanding how things connect. Learning where the bugs hide. Why that one file nobody touches exists. It takes months because a codebase is massive and human memory is small. > Claude just loaded the entire thing in one prompt. 30 seconds. Every file, Every function, Every line. All of it. Sitting in memory like it's been working there for years. And it scored highest among every single frontier model. Not GPT.. Not Gemini, Nobody. > Yesterday Amazon's AI nuked production because it couldn't see the full picture - it made a decision with partial context and deleted everything. Today an AI can hold 1 million tokens of context at once. That's the fix. That's the "before and after" moment for AI coding. > 600 images in one request. Entire PDFs. Full repos. And they dropped it on a Friday on all plans like it was a patch note. The scariest AI updates aren't the ones with press conferences. They're the ones that drop in a tweet at 6pm and change everything by Monday morning.show more

Tuki
206,309 просмотров • 5 месяцев назад
THIS GUY JUST REBUILT A $35,000 ANIMATED SITE FOR... $12. IF YOU RUN A WEB STUDIO, YOU SHOULD PROBABLY KEEP SCROLLING. Every agency billing $100-149/hr is selling you five departments wearing one invoice. Here’s each one - collapsed into a single agentic session. LAYER 1 - THE CONCEPT ROOM (Claude) Reads the brief, pulls references, and scripts the scroll: what the visitor feels at second 3, second 15, second 40. → Used to be a strategist and a wall of mood boards. Now it’s a conversation. LAYER 2 - THE MOTION STUDIO (Higgsfield) Cinematic clips from 30+ generative models - hero shots, transitions, ambient loops - all matched to the story from Layer 1. → Used to be a motion artist on retainer. Now it’s a prompt. LAYER 3 - THE DEV TEAM (Claude Code) Scaffolds the site, writes the GSAP ScrollTrigger timelines and Lenis smooth-scroll, extracts frames, optimizes every asset. → A full scroll-driven build with zero hand-coded keyframes. LAYER 4 - THE DESIGN DEPT (baked-in cinematic layer) Six effects, zero config: film grain, particles, vignette, glass cards, color tints, scroll pacing. → The polish that justified the invoice - now it ships by default. LAYER 5 - THE QA PASS (Claude) Checks load speed, mobile breakpoints, and whether the scroll actually lands - then rewrites whatever doesn’t. → Used to be a client call and a revision cycle. Now it’s one more turn in the same session. Five departments. One operator. One pass. A strategist, a motion artist, a developer, a designer, and a QA lead - weeks of handoffs - now run in a single session. For a Claude subscription and a few dollars of Higgsfield credits. The studio was never selling talent. It was selling overhead. And the overhead just became five layers. Follow me, reply “website” to this post and I will send you the step-by-step Playbook 👇show more

ZEUS⚡️
141,973 просмотров • 2 месяцев назад
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 просмотров • 13 дней назад
I genuinely want our 🇮🇳 economy to reach $5... trillion by 2027. But with GDP growth at 6.4% in Q2 FY26 & FII's pulling out ₹1.75 lakh crore in 2024-25, #Budget2026 on February 1st must deliver bold reforms ! 📈 👉Sharing my 7 critical expectations from our FM Nirmala Sitharaman ji 👇.. 📊 Cut LTCG tax back to 10% & double exemption to ₹2.5 lakh 12.5% tax is killing long term returns for every equity investor. Before July 2024 we paid only 10% above ₹1 lakh & people stayed invested happily. Government hiked it suddenly & broke confidence. Roll back to 10% now. Raise exemption from ₹1.25 lakh to ₹2.5 lakh so small SIP guys with ₹10k-20k monthly build wealth without tax on modest gains. Foreigners pulled ₹1.6 lakh crore in 2024-25 partly due to this. Want them back? Stop squeezing returns & protect investors. 📊 Bring STCG tax down from 20% to 15% Jump from 15% to 20% in July 2024 was brutal. Short term traders pay 33% more tax now. Volumes crashed & retail participation dropped hard in late 2024. Less trading means poor liquidity & bad prices for everyone. Drop to 15% now. Markets will wake up, volumes jump, exchanges compete globally. More action means better valuations & wealth for all. 📊 Abolish STT completely or cut by 50%, end double tax nonsense STT on every buy & sell plus capital gains tax on profit is pure double robbery. No major country does this. Government took ₹78000 crore from STT in FY26. Remove STT fully or slash rates half. Make trading cheap & retail will flood back huge. 📊 Massive infra push, commit ₹15-18 lakh crore capex & finish fast ₹11.2 lakh crore capex is too small for $5 trillion by 2027 or beating China on infra. Announce ₹15-18 lakh crore for roads, ports, airports, metros, defence & digital. But projects stuck years in clearances & land fights. Force single window clearance in 30 days. Set fast track courts for disputes in 6 months max. Give infra status to real estate & affordable housing for cheap loans. Every ₹1 spent creates ₹4-5 in cement steel construction. This pushes infra & capital goods stocks 100-150% in 3 years & millions jobs. 📊 Manufacturing boom, 10 year tax holiday plus 50% first year depreciation Make India better than Vietnam Bangladesh Mexico for factories. Give new units 10 years zero tax, especially Tier 2-3 cities. Allow 50% depreciation on machinery first year for cash flow boost. Extend PLI to 25 sectors like toys footwear auto parts chemicals. Cut customs slabs from 8 to 4 simple. One time amnesty to clear old stuck money. This means export boom, jobs, profits & engineering chemical stocks fly. 📊 Put ₹50k-75k extra cash in middle class pockets yearly Raise standard deduction to ₹1.5 lakh from ₹75000, saves ₹15k-22k tax. Increase 80D to ₹1 lakh from ₹25000 as medical costs exploded post COVID. Boost home loan deduction to ₹3 lakh from ₹2 lakh. These put ₹50k-75k extra in 10+ crore salaried hands yearly. Money goes to cars travel education. Consumption is 55% GDP & drives FMCG auto retail stocks up fast. 📊 GST 2.0, fix ITC delays & make business easy ITC refunds take 3-6 months & block crores cash. Give refunds in 30 days max. One single portal for all. Reduce high GST on EV batteries & parts. Faster refunds let companies invest hire grow. Profits rise, stocks valuations up, dividends better for investors ! Looking forward to a really proactive & well thought budget 2026. 👉 Folks, Are your ready for a blockbuster #Budget2026? @zerodhaonline Nirmala Sitharaman Ministry of Finance PIB India valuepickr Narendra Modi CA Anil Singhvi Zee Business ReserveBankOfIndia Nirmala Sitharamanoffc @varinder_bansal Kaushik Basu Arvind Subramanian Sanjeev Sanyal Prof. Krishnamurthy V Subramanian Jayati Ghosh Karan Bhasin Bibek Debroy @ArvindPanagariya Prof. Shamika Ravi @neelkanthmisra #Budget2026 #TaxReform #InvestInIndia #MakeInIndia #MiddleClassRelief #CapitalGainsTax #InfraPush #ManufacturingBoom #GSTReform #InvestorDemandshow more

Advait Arora
24,580 просмотров • 7 месяцев назад
Introducing Pods Hyperspace Pods lets a small group of... people - a family, a startup, a few friends, to pool their laptops and desktops into one AI cluster. Everyone installs the CLI, someone creates a pod, shares an invite link, and the machines form a mesh. Models like Qwen 3.5 32B or GLM-5 Turbo that need more memory than any single laptop has get automatically sharded across the group's devices - layers split proportionally, inference pipelined through the ring. From the outside it looks like one OpenAI-compatible API endpoint with a pk_* key that drops straight into your AI tools and products. No configuration beyond pasting the key and changing the base URL. A team of five paying for cloud AI burns $500–2,000 a month on API calls. The same team's existing machines can serve Qwen 3.5 (competitive on SWE-bench) and GLM-5 Turbo (#1 on BrowseComp for tool-calling and web research) for free - the hardware is already on their desks. When a query genuinely needs a frontier model nobody has locally, the pod falls back to cloud at wholesale rates from a shared treasury. But for the daily work - code reviews, refactors, research, drafting - local models handle it and nobody gets billed. And when it is idle, you can rent out your pod on the compute marketplace, with fine-grained permissions for access management. There's no central server involved in inference. Prompts go from your machine to your pod members' machines and back: all of this enabled by the fully peer-to-peer Hyperspace network. Pod state - who's a member, which API keys are valid, how much treasury is left - is replicated across members with consensus, so the whole thing works on a local network. Members behind home routers don't need port forwarding either. The practical setup for most pods is three models covering different jobs: Qwen 3.5 32B for code and reasoning, GLM-5 Turbo for browsing and research, Gemma 4 for fast lightweight tasks. All running on hardware you already own. Pods ship today in Hyperspace v5.19. Model sharding, API keys, treasury, and Raft coordinator are all live. What Makes This Different - No middleman. Your prompts travel from your IDE to your pod members' hardware and back. There is no server in between reading your data. - No vendor lock-in. Pod membership, API keys, and treasury are replicated across your own machines using Raft consensus. If the internet goes down, your local network keeps working. There is no database in someone else's cloud that your pod depends on. - Automatic sharding. You don't configure layer ranges or calculate VRAM budgets. Tell the pod which model you want. It figures out how to split it across whatever hardware is online. - Real NAT traversal. Your friend behind a home router with a dynamic IP? Works. No VPN, no Tailscale, no port forwarding. The nodes handle it. - Free when local. This is the part that matters most. Cloud AI bills scale with usage. Pod inference on local hardware scales with nothing. The marginal cost of your 10,000th prompt is the electricity your laptop was already using. Coming soon: - Pod federation: pods form alliances with other pods. - Marketplace: pods with spare capacity can sell inference to other pods.show more

Varun
309,520 просмотров • 4 месяцев назад
Release: LichtFeld Studio v0.5.3 is out! With 316 commits... merged into master, this release is a huge step forward for LichtFeld Studio. What's new in v0.5.3 • Vulkan viewer/rendering migration: New Vulkan viewport pipeline, pass graph, VkSplat renderer, Vulkan point-cloud renderer, 3DGUT/VkSplat support, improved alpha/depth composition, tighter CUDA/Vulkan interoperability, and device matching on multi-GPU systems. • RAD + LOD workflow: Added RAD file export/import, RAD LOD viewer, Spark-style GPU LOD selection, GPU-driven page prefetching, a bounded VRAM pool, out-of-core PLY-to-RAD LOD conversion, and RAD import/export speedups of approximately 3–5×. • HiGS / macro-tile inference: Added a macro-tile inference path for the Vulkan viewer, including macro sorting, batched rasterization, composition, and capacity management. • Asset Manager: Added and significantly enhanced the Asset Manager with thumbnails, SH information, faster synchronization, import-from-URL support, docked mode, data-loading popup integration, and general UI cleanup. • Viewport export: Integrated viewport export directly into the application as a toolbar/overlay tool, added fast render_view_u8-style readback paths, fixed high-resolution clipping issues, improved orthographic export parity, resolved 32K image/video export problems, and added post-export GPU resource cleanup. • Selection and tooling: Added and reworked selection toolbar controls, the Select menu, ring selection, color eyedropper, distance-from-center selection, faster point-cloud and zoomed-out selection paths, Vulkan measurement tool fixes, and drag-and-drop scene graph improvements. • UI/RmlUi platform work: Major RmlUi redesign efforts, hot reloading for RML/RCSS/Python UI files, reactive UI/store integration, viewport toolbar flyouts, improved histogram interactions, input settings enhancements, custom TRS gizmos, and numerous panel, tooltip, and localization fixes. • Windowing and UX: Added borderless window support, title bar drag/maximize/restore behavior, work-area-aware maximize functionality, resize responsiveness and performance improvements, and DPI/UI scaling fixes. • Training and data features: Added adaptive depth loss and depth gradients for the EWA rasterizer, mask loading/application fixes, a new combined Ignore+Segment mask mode, --add-splat, --freeze, improved checkpoint and training state handling, and training speed and VRAM optimizations. • COLMAP/equirectangular support: Added SPHERICAL/equirectangular camera model support and canonical EQUIRECTANGULAR handling, along with fixes for undistortion and camera export. This release will be available to all supporters as a Windows binary via approximately in about an hour. At the same time, LichtFeld Studio remains committed to being free and open source under GPLv3 and can also be built directly from source. Please consider supporting the ongoing development of LichtFeld Studio through a donation via the portal or the supporters page. Thank you to everyone who supports this project financially, contributes code, reports bugs, provides datasets, helps with the website, and contributes in countless other ways. A special thank you to our foundational sponsor Core11 and our Gold Sponsor Volinga, whose support has helped make the current state of the software possible. Thank you as well to every donor and to all of our new Bronze Sponsors. Looking ahead to v0.6 For the next major release, work will focus primarily on stability and user experience. This includes improved cleanup workflows and the ability to modify training parameters while training is in progress. I would also like to introduce a native .licht project format that allows users to save and restore their complete editor state. You can find links to our main sponsors below. Please also visit our website to discover all our Bronze Sponsors. Hint: We do not yet have a Silver Sponsor or Platinum 😉show more

MrNeRF
26,219 просмотров • 2 месяцев назад
🚨 Anthropic committed up to 1M TPU chips for... Claude. Openai is leasing TPUs for chatgpt inference. Here's How kernels work on TPUs (deep dive 2/6 by emi) pallas is Google's answer to kernel writing. a python kernel SDK built on JAX. still very experimental (jax.experimental.pallas). on TPU it compiles through mosaic; on GPU it lowers to triton. if you know CUDA, the syntax will feel familiar but the execution model is completely different. in CUDA, grid=(4,4) launches 16 blocks running simultaneously across SMs. in pallas, those 16 iterations run one after another in lexicographic order. no threads. no warps. no blocks. no occupancy tuning. a TPU is a sequential machine with a very wide vector register — more like a CPU than a GPU. performance comes from width: a 128x128 systolic array doing matmul and an 8x128 SIMD vector unit doing everything else. maximum parallelism on chip: 2, one per TensorCore in megacore mode. three concepts replace CUDA's thread/block/grid hierarchy. Refs are mutable memory references. because execution is sequential, each iteration safely accumulates without atomics. in CUDA you'd need atomics or a separate reduction pass. the memory model is also very different from NVIDIA's. zero hardware caches. VMEM is 32-128 MiB of software-managed scratchpad — 500-1000x larger than GPU shared memory per SM. all data must be explicitly DMA'd from HBM to VMEM before any computation touches it. four levels: HBM → VMEM → VREGs → MXU/VPU, plus SMEM for scalar control data. every byte of data movement is your responsibility. this is like CUDA shared memory except it's 500x bigger and there's no cache fallback. pipelining is mandatory. without double-buffering HBM→VMEM transfers, the MXU just stalls waiting for data. this is the single most important optimization on TPU. and because grid execution is sequential and deterministic, consecutive iterations that need the same input block skip the redundant HBM transfer automatically, impossible on GPU where block execution order is undefined. the compilation pipeline is unlike anything in this series: python → jaxpr → stableHLO → XLA HLO (71+ optimization passes) → LLO (78+ passes) → 322-bit VLIW bundles. the compiler packs instructions for scalar, vector, matrix, and DMA units into a single 322-bit word. everything in that bundle executes in parallel, with no runtime scheduling.show more

wafer
33,134 просмотров • 1 месяц назад
MiniMax H3 - Text to Video, Burger Ad You... can easily adapt this prompt for any food. Prompt: Create a high-end cinematic burger commercial in portrait format, designed like a premium restaurant campaign. Open on an extreme macro shot of a golden toasted bun, revealing tiny bake-specks, soft flour dusting and rich surface texture under warm directional light. The camera glides smoothly across the bun, then pulls back as the burger ingredients separate into a perfectly aligned exploded stack above a matte wooden tabletop. The full burger floats on a single vertical axis: toasted top bun, crisp ruffled lettuce, glossy tomato slices, translucent purple-red onion rings, melted cheddar, a thick char-grilled beef patty and toasted bottom bun. Each layer moves with elegant controlled motion, subtle rotation and realistic weight. Use premium food-commercial camera movement throughout: macro tracking shots, smooth dolly pushes, slow orbital moves, shallow-focus passes between ingredients and precise rack focuses. Add occasional speed ramps into slow-motion beauty moments. Let tiny crumbs, droplets and subtle steam move through the light for additional depth. Integrate bold hand-lettered white typography directly into the composition. Words appear between the floating burger layers, following the camera movement with kinetic typography. Letters can slide behind ingredients, reveal through depth, stretch slightly during transitions and lock cleanly into place. Add minimal white doodle strokes and graphic accent lines that animate around key ingredients. Lighting is luxurious and appetizing: soft directional key from upper-left, controlled fill, warm highlights, deep dimensional shadows and glossy specular detail on tomato, cheese and meat. Background is a refined warm brown gradient with subtle cinematic falloff. Build toward a final satisfying moment where all ingredients rapidly assemble into one perfect burger. The camera performs a fast controlled push-in, then settles into a polished hero shot. Steam rises gently from the patty, cheese settles over the edges and the typography resolves beside the burger. Final frame: centered premium burger hero shot, clean composition, elegant white campaign typography and subtle animated graphic accents. Visual style: luxury food advertising, modern restaurant campaign, cinematic macro photography, rich warm tones, high contrast, shallow depth of field, highly detailed food textures, sophisticated motion design, premium typography, smooth dynamic camera choreography, polished commercial finish.show more

Kōda
24,225 просмотров • 3 дней назад
🚀 Dive into MetaMask Season 1 with Linea: new... on-chain quests are live in daGama! We’re excited to announce that new quests from Linea.eth and MetaMask 🦊 are now live on daGama’s Questboard! Dive into the MetaMask Season 1 campaign with the major Layer 2 network Linea, explore the ecosystem in depth, and earn up to 700 daGama XPs 💫 🌐 What is Linea? Linea is zkEVM Layer2 network bringing Ethereum's security, scalability, and developer tools to millions of users. Built by Consensys, it offers fast, low-cost transactions with full EVM compatibility. Nearly 45% of all MetaMask swaps now happen on Linea, showing how quickly it has become a go-to destination for DeFi, perpetuals, gaming, and everyday on-chain activity. 🦊 What is MetaMask? MetaMask is a trusted Web3 wallet and browser extension that empowers users to explore DeFi, NFTs, and dApps across multiple blockchains while maintaining full control of their private keys. With over 30 million monthly active users and 100 million yearly users worldwide, it remains one of the most secure and widely adopted crypto wallets in 2025. Now you can join MetaMask Season 1 with Linea via daGama Questboard & get XPs in both campaigns: 1️⃣ Explore MetaMask Season 1 Download MetaMask and add your existing address to the Rewards tab in the daGama mobile app: Reward: 100 XP 2️⃣ Linea swaps Switch to the Linea chain, open the “Trade” tab, and swap any available tokens (min. $100). After a successful swap, leave your transaction ID in the answer field. Reward: 300 XP 3️⃣ Explore perps on MetaMask Start perpetual trading on MetaMask using $LINEA (any amount is acceptable), and submit your transaction ID. Reward: 300 XP ⚠️ Perpetual trading involves high risks. Participate responsibly. ⏳ Don’t miss your chance to boost your XP and climb the lead.show more

daGama
72,215 просмотров • 8 месяцев назад
MiniMax H3 MiniMax Design (H3) just dropped, and I... tested it with this 15-second prompt. The results were more interesting than I expected. The character movement, physical interactions, and shot continuity are all impressive. At first glance, it already feels very close to Seedance 2.0. Watch the final video and find the full prompt below 👇 15-second, 16:9 vertical, continuous single-take video that looks like authentic smartphone footage accidentally captured by a passerby in a city park. Overcast natural daylight, subtle handheld shake, limited phone stabilization, occasional autofocus adjustment, and realistic smartphone compression. The absurd event is filmed with a completely serious, unscripted documentary feeling. 0–3s: [Handheld medium shot] A middle-aged man wearing a dark business suit and tie crouches beside the stone edge of a pond. With a completely serious expression, he slowly scatters pieces of bread from a paper bag to several ordinary koi. Small ripples spread across the water as the fish gather in front of him. 3–7s: [Camera instinctively moves closer] An abnormally huge orange-and-white koi suddenly surges out of the murky water, briefly lifting its upper body above the surface and biting down on the entire bread bag in the man’s left hand. He freezes for half a second, then grips the bag with both hands and leans backward. Startled, the person filming steps back. The image briefly loses focus before locking onto the man and the giant fish again. 7–11s: [Close handheld action shot] The giant koi pulls violently toward the deeper part of the pond. The wet paper bag stretches, the man’s arms tense, and his leather shoes slide repeatedly across the wet stone. His knee strikes the edge of the pond. He tries to brace himself with his right foot, but the sole loses traction and his center of gravity moves past the edge. Water, pieces of bread, and fallen leaves scatter from the force as the camera operator hurriedly moves sideways. 11–15s: [Impact and final hold] The paper bag suddenly tears. The man loses all support and pitches forward into the pond, creating one heavy, realistic splash. The camera quickly tilts downward while keeping the center of the pond visible. The man resurfaces with duckweed covering his head and his wet tie stuck across his face. The giant koi calmly swims past him with the remains of the bread bag still in its mouth. The camera holds on the man’s stunned expression while the koi casually swims away beside him. Keep the man’s face, dark suit, tie, and paper bag visually consistent throughout. The koi must retain the same orange-and-white markings and enormous size. The pulling, sliding, loss of balance, and fall must show believable weight, inertia, traction, and water displacement. Natural park ambience and a realistic splash only. No dialogue, no subtitles, no music. Avoid cuts, character teleportation, changes in the fish’s size, extra limbs, and cartoonish acting. #MiniMaxH3 #AIVideoshow more

underwood
12,373 просмотров • 1 месяц назад