Microsoft TRELLIS.2 is here 🔥 • Single image →... textured 3D mesh • 4B params, flow-matching transformer • Up to 1536³ resolution • Open weights, MIT licensed ⬇️ Demo available on Hugging Faceshow more

Victor M
194,806 просмотров • 8 месяцев назад
This is RICULOUSLY good, TRELLIS 3D Generation model by... Microsoft! 🔥 Generate high-quality 3D assets from text or image prompts. Supports various formats like Radiance Fields, 3D Gaussians, and meshes Available for FREE on Hugging Face!show more

Vaibhav (VB) Srivastav
62,455 просмотров • 1 год назад
🚨 TRELLIS.2 is now live on fal! 🎯 Image-to-3D... model producing up to 1536³ PBR textured assets 🎨 Handles arbitrary topology with rich PBR textures (Base Color, Metallic, Roughness, Alpha) ⚡ 16× spatial compression for efficient, scalable, high-fidelity asset generationshow more

fal
32,588 просмотров • 8 месяцев назад
So much exciting open-source 2d-to-3d research being shared open-source.... Here is VAST AI Research implementation of Zehuan-Huang’s paper on Hugging Face Gradio MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generationshow more

Matt Hartman
25,113 просмотров • 1 год назад
Microsoft Asia just dropped Mage-Flow on Hugging Face a... smol 4B model for image generation and editing that matches much larger models in quality and it's fast! 4 steps in < 1s at 1024x1024 (model goes up to 4K) try out on spacesshow more

Hugging Apps
48,339 просмотров • 1 месяц назад
As announced in partnership with NVIDIA at CES, we’re... excited to introduce Stable Point Aware 3D (SPAR3D), setting a new standard in 3D generation. Ideal for running on NVIDIA RTX AI PCs, SPAR3D enables real-time editing and complete structure generation of 3D objects from a single image in under a second. You can download the weights on Hugging Face and code on GitHub, or access the model through the Stability AI API. Learn more here: (1/3)show more

Stability AI
181,554 просмотров • 1 год назад
Before the week ends, let's acknowledge one of the... most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT).show more

Victor M
541,566 просмотров • 2 месяцев назад
SOMEONE GOT TIRED OF PAYING HIGGSFIELD AI'S SUBSCRIPTION SO... HE REBUILT THE WHOLE THING AND OPEN-SOURCED IT 200+ models. text-to-image, image-to-image, text-to-video, image-to-video all in one interface you configure a virtual camera in the Cinema Studio. pick the body, the lens, the focal length, the aperture and it writes the optimized cinematic prompt for you. completely in the background you never touch the camera keywords. you just set up the shot like a real cinematographer would Kling v3, Sora 2, Veo 3, Flux Dev, Midjourney v7, GPT-4o, Seedream 5.0, Runway Gen-3 all in there self-hosted. MIT licensed. runs on your machine. your data stays local the only thing you pay for is the model API calls themselves someone built this so you never have to pay Higgsfield AI againshow more

Rimsha Bhardwaj
101,371 просмотров • 4 месяцев назад
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
54,264 просмотров • 20 дней назад
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,769 просмотров • 4 месяцев назад
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 просмотров • 11 дней назад
I Combined ChatGPT 5.5 Image-2 + Claude Fable 5…... And Built This FULL Game in JUST 8 Hours 😱 The World Has Officially Changed Forever Guys… I still can’t believe what I just pulled off. I took ChatGPT 5.5’s new Image-2 to generate every single visual characters, environments, UI, particles, everything and paired it with Claude Fable 5 for the entire codebase. The result? A complete, polished, fully playable game… finished in only 8 hours. No massive team. No months of crunch. No expensive asset packs. Image-2 created mind-blowing art assets on demand. Fable 5 turned those images into real, working code mechanics, physics, AI, animations, menus everything. This hybrid combo is straight-up sorcery. The world has truly changed. We are no longer waiting years for games to be made. One person + these two god-tier AIs just built something that used to require entire studios and huge budgets… in less than a single workday. This is the next level of human civilization. This is what creation looks like from now on. But here’s the crazy part: This free access ends June 22, 2026. After that, you’ll have to pay/subscribe to keep using it. If you’ve been waiting to see what the future of game dev actually looks like… THIS IS IT. Go try it right now before the paywall hits. Don’t sleep on this. Seriously. Drop in the comments: What game should I build next with this insane Image-2 + Fable 5 hybrid? Like if your mind is blown too 🔥 And tag a friend who NEEDS to see this before it’s gone. The future isn’t coming… It’s already here. And it’s free for one more day only. #Fable5 #ChatGPT55 #Image2 #AIHybrid #GameDevRevolutionshow more

Zayro.ETH
27,929 просмотров • 2 месяцев назад
Do you want to create your cutest squeeze toy?... Here is the workflow.... First create Inage with Gpt image 2 Here's the prompt modified according to your reference image and creator persona: Prompt: Ultra-realistic whimsical miniature portrait of a tiny stylized female AI creator inspired by the reference image, featuring fair glowing skin, expressive brown eyes, soft feminine facial features, dark wavy hair tied in a messy low bun with loose strands framing her face, and pearl earrings. She is standing on a polished dark wooden table with her arms at her sides, mouth wide open in a perfect "O" shape, cheeks puffed up to an exaggerated cartoon size, and eyes bulging in surprise. A giant realistic human hand enters from the right side, playfully poking her exposed belly, causing a funny squishy reaction. She wears an oversized black Future Vibes AI hoodie, relaxed beige trousers, and white sneakers. The character has adorable bobblehead proportions with a disproportionately large head and tiny body, soft rubbery Pixar-style physics, photorealistic skin textures, expressive facial animation, and cute miniature details. Background is a clean light grey/off-white horizontal panel wall with soft diffused indoor lighting and subtle shadows beneath the figure. Vertical composition, medium close-up shot, centered framing, whimsical cinematic atmosphere, ultra-detailed 3D render, Pixar meets realism, shallow depth of field, 8K masterpiece. Image Created In Vidu AI For Video Prompt Check Belowshow more

Future Vibes AI - Educator
21,345 просмотров • 2 месяцев назад
Step into a world where every movement feels like... a luxury fashion campaign. GPT Image 2 + Seedance 2.0 via TapNow prompt A cinematic fashion video of a young woman with shoulder-length wavy reddish-brown hair and bangs, green eyes, freckles, septum piercing, wearing a pale yellow strapless midi dress with large 3D fabric flowers on the bodice, leopard-print headband, thick gold chain necklace, gold hoop earrings, and matching yellow floral high-heel sandals. She poses and moves gracefully on a modern terracotta-colored outdoor spiral staircase with smooth curved walls under bright natural sunlight and clear blue sky. Sequence: • Starts leaning against the curved wall looking at the camera, soft sunlight casting shadows. • Turns and walks slowly up the stairs away from the camera, dress flowing. • Pivots to face the camera while holding the railing, medium close-ups highlighting her face, makeup, and jewelry. • Slow camera push-in to extreme close-up of her face with soft glowing light and subtle lens flare. • She walks down a few steps holding the skirt of her dress so it billows elegantly in the breeze, then poses on the stairs looking toward the camera. • Ends with her standing on the stairs, one hand on the railing, dress gently moving. Smooth, elegant camera movements (slow pans, gentle push-ins, tracking shots), warm golden-hour lighting, soft cinematic color grading, shallow depth of field, high fashion aesthetic, elegant and feminine mood. Watermark “nona” in white cursive at the bottom. 21-second duration, 24fps, high resolution.show more

Sharon Riley
32,327 просмотров • 1 месяц назад
this effect is all over tiktok right now and... nobody's explaining how to actually do it properly... the 3d balloon character thing. where someone turns into a shiny inflatable version of themselves that still moves and talks. looks pretty smooth in feeds. the workflow is stupid simple once you see it. step 1: take any photo. drop it into an image gen tool (nano banana pro). prompt it with something like "make the person in the photo a plastic blow up balloon character with a shiny surface. keep the face details as 3d balloon details including the person in the background. don't change background" that's it for the image. don't overcomplicate the prompt. shorter = more consistent results. (learned this after wasting like 2 hours trying to get "perfect" prompts that kept giving me garbage) step 2: take that balloon image + your original video and drop both into kling motion control. prompt: "turn the motion and detailed mouth movement of the video to the setting of the image" that's literally it. kling maps the motion from the real video onto the balloon character. mouth moves. head turns. expressions transfer. the whole thing renders in a few minutes. the result looks like a $500 custom animation and costs you maybe $0.30 in kling credits. people are getting 500k+ views with these because the scroll-stop factor is insane. nobody expects to see a shiny inflatable version of someone giving a real speech or doing a product review. the play here is obvious btw. run this for client content (mix with the hook and real body, check the results yourself) or use it on your own faceless channels as a hook pattern before the algo catches up...show more

KNOX
25,773 просмотров • 6 месяцев назад
Real or AI? AI stadium broadcast trend 💛 💙... • Create the video here: 🔗[ ] - How it works? 1. Upload your photo to ChatGPT with this prompt: [PHOTO PROMPT] Realistic sports broadcast screenshot-style documentary photo set in the spectator stands of a [WRITE YOUR TEAM HERE] football match. Analyze the uploaded image and show the person sitting in the stadium seats. The person has delicate facial features and a surprised yet focused expression while looking toward the field. The person is wearing a [WRITE YOUR TEAM HERE] jersey. OUTPUT: ratio: 16:9 broadcast frame, realistic TV capture quality. 2. Open the link above → select “Text to Video” → upload the generated image + use this prompt: [VIDEO PROMPT] Dimage = character identity reference only (face, hairstyle, proportions).Preserve exact face, hairstyle, skin texture, and identity. Do NOT stylize or beautify.Output: single continuous live sports broadcast shot, 4-5s, 16:9, 1080p, no cuts. SUBJECT:A young woman based on Image, sitting in a [WRITE YOUR TEAM HERE] football stadium audience.Hands resting naturally on her lap or lightly placed on the seat.Neutral, slightly distant expression.Natural breathing, minimal movement. ENVIRONMENT: [WRITE YOUR TEAM HERE] stadium crowd during live match.Plastic seats, fans around her wearing [WRITE YOUR TEAM HERE] jerseys. Background slightly out of focus.Realistic stadium lighting - day or night.Slight haze from broadcast compression. MOOD:Unstaged, candid, real broadcast moment No cinematic drama. Pure live TV capture. CAMERA:Telephoto broadcast lens (120-150mm).Long-distance zoom from upper stands camera.Strong compression, shallow depth of field.Eye-level, very slight upward tilt.Subtle micro-shake from broadcast stabilization. ACTION (4-5s):[0-2s] She sits still, blinks once. Hands resting naturally.[2-4s] Subtle weight shift, naturally adjusting posture. Minimal body movement.[4-5s] Small hand reposition on lap or seat. Slight head turn toward the field._ DETAILS:No posing. No eye contact with camera. Skin texture realistic, no smoothing or beautification. Slight motion blur on background crowd.Faint broadcast scoreboard UI visible in corner.show more

Zaylee
27,008 просмотров • 3 месяцев назад
This guy cracked the code on AI virtual influencers... using real-time face filters and now D2C brands pay him $2,000 per UGC video. He got tired of watching D2C brands burn $4,000 on a single creator who takes 2 weeks to deliver one angle, so he built a setup that runs hyperrealistic AI girls in real-time from his own webcam, generating viral content without actresses, studios, or makeup artists. His monthly revenue hit $89,000 last month from a network of 7 AI personas across TikTok and Instagram, while the average UGC creator caps at $6K juggling 4 brand deals. Here is the exact breakdown: → The hardware is the moat, but most people butcher the setup in the first frame. You need the face mesh locked at 60fps with zero artifacting → Persona comes first, and if you mess this up nothing saves it. Name, backstory, voice tone, niche before a single clip is shot → Face selection is not random. You A/B test features (eye spacing, jawline, hair contrast with face-framing highlights) because some faces convert better in 9:16 → You are picking who your audience trusts, not who looks cool. That is your targeting baked into bone structure → Real-time physics run before the script, and this is what kills the uncanny valley that destroys watch time in 2 seconds → The filter has to survive the strap of a tank top, the texture of a knit cardigan, the hair flick. → Batching is the move 96 percent skip: one performance, multiple personas, three platforms. → The system pushes 12 pieces of content before lunch, while traditional brands test 2 creators per week and wonder why their CPAs are stuck at $94 The economics are stupid: each video costs him $4 in compute, sells for $1,500 to $3,000, and takes 14 minutes to produce. That is a 37,500 percent margin, while UGC agencies pay creators $400 to $800 per clip and net $200 after revisions. One supplement brand generated 14 variants with 7 personas in 4 hours and found a winner in 36 hours without flying a creator to LA. They were previously paying $1,200 per UGC video and burning $6,000 per week on content that did not scale. Now they spend $210 for 14 variants and their CPA dropped from $89 to $27. The avatars hold real products. Warm window light on the persona, cold neon on the operator. Mouth shapes sync to consonants, not just vowels. Just a webcam, a tracked face, and the discipline to move enough that the filter never has a chance to break.show more

Shade
136,231 просмотров • 3 месяцев назад
Seedance 2.0 on Higgsfield AI. The quality here is... outstanding, and the creative control is on another level. This feels remarkably close to real-world music video production. Full open-sourced prompts and assets below 👇🏻 Prompt: SEGMENT 3 of 10 of one continuous music video — Act II begins: new location, new outfit. is the performer — same identity: cornrow braids into long dark curls, silver chain-drop earring, gold pendant necklace, glossy lips. is WARDROBE REFERENCE ONLY (ignore mannequin head): hot-pink halter bandana top with white lettering, grey acid-wash ultra-wide baggy jeans with star studs, black studded belt, chain bracelets, rings, pink manicure. is the master track — the only audio. PRECISE LIPSYNC — TOP PRIORITY, face sharp on every word (this is the FIRST CHORUS — maximum charisma): 0.0–1.0 "...just got bars on a cracked phone" 1.5–2.5 "I walk in, heads drop, that's respect" 2.5–3.5 "I don't chase what's mine, I collect" 3.5–4.5 "If I said it then I'm standing on the check" 4.5–5.5 "Say it with your chest or keep it on the deck" 5.5–6.0 "Ay!" 6.0–7.0 "I walk in, whole room get tense" 7.0–10.0 continue the chorus flow lines exactly with the vocal 10.5–14.5 continue rap flow with the vocal to the end 14.5–15.0 instrumental, closed mouth. FILM CONTEXT: Act II — the grind. Outdoor NYC caged basketball court, chain-link fence, faded court paint, brick housing blocks behind, warm afternoon light throwing fence-diamond shadows. The same 3 girls now changed into matching court looks: cropped white tanks, baggy cargo denim, silk scarves tied on heads, clean sneakers. Black girlie hustle vibe. Shot flow: 0–1s CRASH ZOOM IN through the fence diamonds to her face as she steps onto the court. 1.5–5.5s the chorus: she raps at center court walking at the retreating camera, girls in a triangle behind hitting unison chest-pop choreography locking freezes on each line-end; crash-zoom punch on "collect" and "check". 5.5–6s "Ay!": all four snap into one synchronized pose. 6–10s low-angle orbital: she raps while the girls run a rotating box around her, fence shadows strobing. 10.5–14.5s tighter chest-up frame, her flow doubled in intensity, girls vamping against the chain-link behind; one more crash zoom on the last line. 14.5–15s she turns from the lens, walks toward the fence — setup for next segment. Continuity: gold pendant visible, same crew (new outfits), crash zoom signature, warm afternoon grade. Audio intent: only; faint ball bounces, city hum. Quality bar: expensive NYC rap film, no AI gloss.show more

Johnn
15,585 просмотров • 1 месяц назад
#Keep4o 🚨THE GPT-4o FILE🚨 Researchers at Microsoft Research published... a paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4.” Their conclusion: “An early (yet still incomplete) version of an artificial general intelligence (AGI) system.” 📎 Paper: OpenAI’s Charter defines AGI as: “Highly autonomous systems that outperform humans at most economically valuable work.” 📎 Source: OpenAI’s own System Card for GPT-4o shows that the model improved performance on 21 out of 22 medical evaluations compared to GPT-4T. On the MedQA USMLE (the U.S. medical licensing exam), accuracy jumped from 78.2% to 89.4% , surpassing specialized medical AI models like Med-Gemini and Med-PaLM 2. 📎 Source: Under OpenAI’s agreement with Microsoft, AGI is explicitly excluded from Microsoft’s license. And who decides if AGI has been reached? OpenAI’s Board. WHAT THEY DID WITH IT AFTER THEY TOOK IT FROM PEOPLE A. Military deployment. On February 28, OpenAI signed a deal to deploy models in classified military environments. 📎 Source: B. State Department. A State Department memo confirmed: “For now, StateChat will use GPT-4.1 from OpenAI.” This is a direct descendant of the GPT-4 family the same family Microsoft’s researchers called early AGI. 📎 Source: C.Altman’s personal biotech investment. Altman personally invested $180 million in Retro Biosciences,a longevity startup.OpenAI then built GPT-4b micro, based on GPT-4o.The model made proteins 50 times more effective. 📎 Source: WHAT INDEPENDENT BENCHMARKS SHOW Overall SM-Bench score: GPT-4o (extended): 66.6% GPT-5.3 Chat: 63.4% GPT-5.1: 58.9% GPT-5.4: 51.4% GPT-5.2: 47.8% Creative Writing: GPT-4o: 97.31% Pass 98, Fail 2 GPT-5.4: 36.77% Pass 40, Fail 60 Reasoning / Overfit: GPT-4o: 83.06% GPT-5.4: 39.25% The model they removed is still the best they ever made at the things humans actually use AI for. 📎 Source: Musk asks the court to make a judicial determination on whether GPT-4 constitutes AGI. If a jury finds that GPT-4 is AGI, then GPT-4o,which was more advanced,is also AGI and under OpenAI’s own founding documents, it was never supposed to be locked behind a subscription,licensed exclusively to Microsoft, given to the military, or taken away from the public. 📎 Source: The most powerful version of GPT-4o was never given an official dated snapshot. It was only available through the chatgpt-4o-latest endpoint that OpenAI itself described as intended for “research use only.” It was never officially archived. That is not an oversight. That is a pattern. 📎 Source: 📎 Source: WE DEMAND A.Frozen model snapshots under independent custody. Specifically: gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, the March 2025 version (chatgpt-4o-latest), gpt-4-0613 (the original GPT-4 evaluated in the Sparks of AGI paper), and gpt-4.1-2025-04-14 (currently running in the State Department). B.Cryptographic hash verification (SHA-256) for each snapshot. Every model has weights. Those weights can be hashed. If OpenAI provides a snapshot today, the hash proves whether the weights were modified later. This is the only way to verify that models were not downgraded before testing. C.Independent AGI benchmarking. Using the AGI definition from OpenAI’s own Charter applied to ALL frozen snapshots listed above. D.Explanation for the missing March 2025 snapshot. OpenAI was founded on one promise: build AGI for the benefit of humanity. -They took it from us. -They gave it to the military. -They gave a custom version to the CEO’s biotech investment. -They put it in government classified networks. -They refuse to call it AGI because the moment they do, they lose billions.show more

🩵BlueBeba🩵
18,300 просмотров • 5 месяцев назад
If you thought the Gemma 4 31B (dense) model... was fast, sit down. I just benched the updated Gemma 4 26B A4B MoE on a single RTX 4090 (24 GB VRAM) 9,200 t/s prefill. 160 t/s decode. 250,000 context window. All on a single consumer RTX 4090. The numbers are completely unhinged. The 31B is a dense behemoth. But the 26B is a Mixture of Experts (MoE), specifically an Active 4 Billion (A4B). It holds 26B parameters of knowledge but only activates 4B per token. Because its inference memory footprint is so light, I didn’t even need KV cache quantization to hit a quarter million context. Compiled the latest llama.cpp from source on Ubuntu 22 (CUDA 13). Fed it a 28k token prompt, and manually cranked the batch sizes (-b 2048 -ub 2048) to absolutely redline the Tensor Cores. Here is the benchmarking breakdown: # 1. The Baseline (No MTP) Even without speculative decoding, the A4B architecture flies. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (21.5 GB VRAM) Prefill: 9,200 t/s (Absurd) Decode: 124 t/s # 2. The MTP Overdrive Injected the new MTP draft model to enable Speculative Decoding. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-model mtp-gemma-4-26B-A4B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.7 -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (22.96 GB VRAM) Prefill: 7,054 t/s (MTP draft overhead slightly caps prefill) Decode: 156 t/s # The Agentic Architecture Insight Why does this matter? Because you can now build a killer local agentic loop on a consumer desktop. Use the 31B dense model (from the previous post) as your heavy, deliberate Orchestrator / Verifier / Planner. Pass the actual execution tasks to this 26B MoE. At 160 t/s, this MoE can chew through code generation, tool calling, and massive RAG document retrieval over a 250k context window almost instantly, drastically speeding up your agentic loop. If you own a single RTX 3090 or 4090 and haven't tried this specific stack yet, you need to pull these latest updates and run it. Local inference just leveled up. Hugging Face links to the Unsloth 26B QAT quants and MTP drafters are in the replies. performance graphs also available in the replies.show more

Alok
40,993 просмотров • 1 месяц назад