Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Microsoft TRELLIS.2 is here 🔥 • Single image → textured 3D mesh • 4B params, flow-matching transformer • Up to 1536³ resolution • Open weights, MIT licensed ⬇️ Demo available on Hugging Face

194,494 Aufrufe • vor 7 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT).

Victor M

539,883 Aufrufe • vor 1 Monat

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 Aufrufe • vor 11 Monaten

I Combined ChatGPT 5.5 Image-2 + Claude Fable 5… And Built This FULL Game in JUST 8 Hours 😱 The World Has Officially Changed Forever Guys… I still can’t believe what I just pulled off. I took ChatGPT 5.5’s new Image-2 to generate every single visual characters, environments, UI, particles, everything and paired it with Claude Fable 5 for the entire codebase. The result? A complete, polished, fully playable game… finished in only 8 hours. No massive team. No months of crunch. No expensive asset packs. Image-2 created mind-blowing art assets on demand. Fable 5 turned those images into real, working code mechanics, physics, AI, animations, menus everything. This hybrid combo is straight-up sorcery. The world has truly changed. We are no longer waiting years for games to be made. One person + these two god-tier AIs just built something that used to require entire studios and huge budgets… in less than a single workday. This is the next level of human civilization. This is what creation looks like from now on. But here’s the crazy part: This free access ends June 22, 2026. After that, you’ll have to pay/subscribe to keep using it. If you’ve been waiting to see what the future of game dev actually looks like… THIS IS IT. Go try it right now before the paywall hits. Don’t sleep on this. Seriously. Drop in the comments: What game should I build next with this insane Image-2 + Fable 5 hybrid? Like if your mind is blown too 🔥 And tag a friend who NEEDS to see this before it’s gone. The future isn’t coming… It’s already here. And it’s free for one more day only. #Fable5 #ChatGPT55 #Image2 #AIHybrid #GameDevRevolution

0AIVerse

27,845 Aufrufe • vor 1 Monat

Do you want to create your cutest squeeze toy? Here is the workflow.... First create Inage with Gpt image 2 Here's the prompt modified according to your reference image and creator persona: Prompt: Ultra-realistic whimsical miniature portrait of a tiny stylized female AI creator inspired by the reference image, featuring fair glowing skin, expressive brown eyes, soft feminine facial features, dark wavy hair tied in a messy low bun with loose strands framing her face, and pearl earrings. She is standing on a polished dark wooden table with her arms at her sides, mouth wide open in a perfect "O" shape, cheeks puffed up to an exaggerated cartoon size, and eyes bulging in surprise. A giant realistic human hand enters from the right side, playfully poking her exposed belly, causing a funny squishy reaction. She wears an oversized black Future Vibes AI hoodie, relaxed beige trousers, and white sneakers. The character has adorable bobblehead proportions with a disproportionately large head and tiny body, soft rubbery Pixar-style physics, photorealistic skin textures, expressive facial animation, and cute miniature details. Background is a clean light grey/off-white horizontal panel wall with soft diffused indoor lighting and subtle shadows beneath the figure. Vertical composition, medium close-up shot, centered framing, whimsical cinematic atmosphere, ultra-detailed 3D render, Pixar meets realism, shallow depth of field, 8K masterpiece. Image Created In Vidu AI For Video Prompt Check Below

Future Vibes AI - Educator

21,160 Aufrufe • vor 23 Tagen

Step into a world where every movement feels like a luxury fashion campaign. GPT Image 2 + Seedance 2.0 via TapNow prompt A cinematic fashion video of a young woman with shoulder-length wavy reddish-brown hair and bangs, green eyes, freckles, septum piercing, wearing a pale yellow strapless midi dress with large 3D fabric flowers on the bodice, leopard-print headband, thick gold chain necklace, gold hoop earrings, and matching yellow floral high-heel sandals. She poses and moves gracefully on a modern terracotta-colored outdoor spiral staircase with smooth curved walls under bright natural sunlight and clear blue sky. Sequence: •⁠ ⁠Starts leaning against the curved wall looking at the camera, soft sunlight casting shadows. •⁠ ⁠Turns and walks slowly up the stairs away from the camera, dress flowing. •⁠ ⁠Pivots to face the camera while holding the railing, medium close-ups highlighting her face, makeup, and jewelry. •⁠ ⁠Slow camera push-in to extreme close-up of her face with soft glowing light and subtle lens flare. •⁠ ⁠She walks down a few steps holding the skirt of her dress so it billows elegantly in the breeze, then poses on the stairs looking toward the camera. •⁠ ⁠Ends with her standing on the stairs, one hand on the railing, dress gently moving. Smooth, elegant camera movements (slow pans, gentle push-ins, tracking shots), warm golden-hour lighting, soft cinematic color grading, shallow depth of field, high fashion aesthetic, elegant and feminine mood. Watermark “nona” in white cursive at the bottom. 21-second duration, 24fps, high resolution.

Sharon Riley

32,327 Aufrufe • vor 2 Tagen

this effect is all over tiktok right now and nobody's explaining how to actually do it properly... the 3d balloon character thing. where someone turns into a shiny inflatable version of themselves that still moves and talks. looks pretty smooth in feeds. the workflow is stupid simple once you see it. step 1: take any photo. drop it into an image gen tool (nano banana pro). prompt it with something like "make the person in the photo a plastic blow up balloon character with a shiny surface. keep the face details as 3d balloon details including the person in the background. don't change background" that's it for the image. don't overcomplicate the prompt. shorter = more consistent results. (learned this after wasting like 2 hours trying to get "perfect" prompts that kept giving me garbage) step 2: take that balloon image + your original video and drop both into kling motion control. prompt: "turn the motion and detailed mouth movement of the video to the setting of the image" that's literally it. kling maps the motion from the real video onto the balloon character. mouth moves. head turns. expressions transfer. the whole thing renders in a few minutes. the result looks like a $500 custom animation and costs you maybe $0.30 in kling credits. people are getting 500k+ views with these because the scroll-stop factor is insane. nobody expects to see a shiny inflatable version of someone giving a real speech or doing a product review. the play here is obvious btw. run this for client content (mix with the hook and real body, check the results yourself) or use it on your own faceless channels as a hook pattern before the algo catches up...

KNOX

25,773 Aufrufe • vor 5 Monaten

Real or AI? AI stadium broadcast trend 💛 💙 • Create the video here: 🔗[ ] - How it works? 1. Upload your photo to ChatGPT with this prompt: [PHOTO PROMPT] Realistic sports broadcast screenshot-style documentary photo set in the spectator stands of a [WRITE YOUR TEAM HERE] football match. Analyze the uploaded image and show the person sitting in the stadium seats. The person has delicate facial features and a surprised yet focused expression while looking toward the field. The person is wearing a [WRITE YOUR TEAM HERE] jersey. OUTPUT: ratio: 16:9 broadcast frame, realistic TV capture quality. 2. Open the link above → select “Text to Video” → upload the generated image + use this prompt: [VIDEO PROMPT] Dimage = character identity reference only (face, hairstyle, proportions).Preserve exact face, hairstyle, skin texture, and identity. Do NOT stylize or beautify.Output: single continuous live sports broadcast shot, 4-5s, 16:9, 1080p, no cuts. SUBJECT:A young woman based on Image, sitting in a [WRITE YOUR TEAM HERE] football stadium audience.Hands resting naturally on her lap or lightly placed on the seat.Neutral, slightly distant expression.Natural breathing, minimal movement. ENVIRONMENT: [WRITE YOUR TEAM HERE] stadium crowd during live match.Plastic seats, fans around her wearing [WRITE YOUR TEAM HERE] jerseys. Background slightly out of focus.Realistic stadium lighting - day or night.Slight haze from broadcast compression. MOOD:Unstaged, candid, real broadcast moment No cinematic drama. Pure live TV capture. CAMERA:Telephoto broadcast lens (120-150mm).Long-distance zoom from upper stands camera.Strong compression, shallow depth of field.Eye-level, very slight upward tilt.Subtle micro-shake from broadcast stabilization. ACTION (4-5s):[0-2s] She sits still, blinks once. Hands resting naturally.[2-4s] Subtle weight shift, naturally adjusting posture. Minimal body movement.[4-5s] Small hand reposition on lap or seat. Slight head turn toward the field._ DETAILS:No posing. No eye contact with camera. Skin texture realistic, no smoothing or beautification. Slight motion blur on background crowd.Faint broadcast scoreboard UI visible in corner.

Zaylee

26,078 Aufrufe • vor 2 Monaten

This guy cracked the code on AI virtual influencers using real-time face filters and now D2C brands pay him $2,000 per UGC video. He got tired of watching D2C brands burn $4,000 on a single creator who takes 2 weeks to deliver one angle, so he built a setup that runs hyperrealistic AI girls in real-time from his own webcam, generating viral content without actresses, studios, or makeup artists. His monthly revenue hit $89,000 last month from a network of 7 AI personas across TikTok and Instagram, while the average UGC creator caps at $6K juggling 4 brand deals. Here is the exact breakdown: → The hardware is the moat, but most people butcher the setup in the first frame. You need the face mesh locked at 60fps with zero artifacting → Persona comes first, and if you mess this up nothing saves it. Name, backstory, voice tone, niche before a single clip is shot → Face selection is not random. You A/B test features (eye spacing, jawline, hair contrast with face-framing highlights) because some faces convert better in 9:16 → You are picking who your audience trusts, not who looks cool. That is your targeting baked into bone structure → Real-time physics run before the script, and this is what kills the uncanny valley that destroys watch time in 2 seconds → The filter has to survive the strap of a tank top, the texture of a knit cardigan, the hair flick. → Batching is the move 96 percent skip: one performance, multiple personas, three platforms. → The system pushes 12 pieces of content before lunch, while traditional brands test 2 creators per week and wonder why their CPAs are stuck at $94 The economics are stupid: each video costs him $4 in compute, sells for $1,500 to $3,000, and takes 14 minutes to produce. That is a 37,500 percent margin, while UGC agencies pay creators $400 to $800 per clip and net $200 after revisions. One supplement brand generated 14 variants with 7 personas in 4 hours and found a winner in 36 hours without flying a creator to LA. They were previously paying $1,200 per UGC video and burning $6,000 per week on content that did not scale. Now they spend $210 for 14 variants and their CPA dropped from $89 to $27. The avatars hold real products. Warm window light on the persona, cold neon on the operator. Mouth shapes sync to consonants, not just vowels. Just a webcam, a tracked face, and the discipline to move enough that the filter never has a chance to break.

Shade

135,682 Aufrufe • vor 2 Monaten

Seedance 2.0 on Higgsfield AI. The quality here is outstanding, and the creative control is on another level. This feels remarkably close to real-world music video production. Full open-sourced prompts and assets below 👇🏻 Prompt: SEGMENT 3 of 10 of one continuous music video — Act II begins: new location, new outfit. is the performer — same identity: cornrow braids into long dark curls, silver chain-drop earring, gold pendant necklace, glossy lips. is WARDROBE REFERENCE ONLY (ignore mannequin head): hot-pink halter bandana top with white lettering, grey acid-wash ultra-wide baggy jeans with star studs, black studded belt, chain bracelets, rings, pink manicure. is the master track — the only audio. PRECISE LIPSYNC — TOP PRIORITY, face sharp on every word (this is the FIRST CHORUS — maximum charisma): 0.0–1.0 "...just got bars on a cracked phone" 1.5–2.5 "I walk in, heads drop, that's respect" 2.5–3.5 "I don't chase what's mine, I collect" 3.5–4.5 "If I said it then I'm standing on the check" 4.5–5.5 "Say it with your chest or keep it on the deck" 5.5–6.0 "Ay!" 6.0–7.0 "I walk in, whole room get tense" 7.0–10.0 continue the chorus flow lines exactly with the vocal 10.5–14.5 continue rap flow with the vocal to the end 14.5–15.0 instrumental, closed mouth. FILM CONTEXT: Act II — the grind. Outdoor NYC caged basketball court, chain-link fence, faded court paint, brick housing blocks behind, warm afternoon light throwing fence-diamond shadows. The same 3 girls now changed into matching court looks: cropped white tanks, baggy cargo denim, silk scarves tied on heads, clean sneakers. Black girlie hustle vibe. Shot flow: 0–1s CRASH ZOOM IN through the fence diamonds to her face as she steps onto the court. 1.5–5.5s the chorus: she raps at center court walking at the retreating camera, girls in a triangle behind hitting unison chest-pop choreography locking freezes on each line-end; crash-zoom punch on "collect" and "check". 5.5–6s "Ay!": all four snap into one synchronized pose. 6–10s low-angle orbital: she raps while the girls run a rotating box around her, fence shadows strobing. 10.5–14.5s tighter chest-up frame, her flow doubled in intensity, girls vamping against the chain-link behind; one more crash zoom on the last line. 14.5–15s she turns from the lens, walks toward the fence — setup for next segment. Continuity: gold pendant visible, same crew (new outfits), crash zoom signature, warm afternoon grade. Audio intent: only; faint ball bounces, city hum. Quality bar: expensive NYC rap film, no AI gloss.

Johnn

15,585 Aufrufe • vor 11 Tagen

#Keep4o 🚨THE GPT-4o FILE🚨 Researchers at Microsoft Research published a paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4.” Their conclusion: “An early (yet still incomplete) version of an artificial general intelligence (AGI) system.” 📎 Paper: OpenAI’s Charter defines AGI as: “Highly autonomous systems that outperform humans at most economically valuable work.” 📎 Source: OpenAI’s own System Card for GPT-4o shows that the model improved performance on 21 out of 22 medical evaluations compared to GPT-4T. On the MedQA USMLE (the U.S. medical licensing exam), accuracy jumped from 78.2% to 89.4% , surpassing specialized medical AI models like Med-Gemini and Med-PaLM 2. 📎 Source: Under OpenAI’s agreement with Microsoft, AGI is explicitly excluded from Microsoft’s license. And who decides if AGI has been reached? OpenAI’s Board. WHAT THEY DID WITH IT AFTER THEY TOOK IT FROM PEOPLE A. Military deployment. On February 28, OpenAI signed a deal to deploy models in classified military environments. 📎 Source: B. State Department. A State Department memo confirmed: “For now, StateChat will use GPT-4.1 from OpenAI.” This is a direct descendant of the GPT-4 family the same family Microsoft’s researchers called early AGI. 📎 Source: C.Altman’s personal biotech investment. Altman personally invested $180 million in Retro Biosciences,a longevity startup.OpenAI then built GPT-4b micro, based on GPT-4o.The model made proteins 50 times more effective. 📎 Source: WHAT INDEPENDENT BENCHMARKS SHOW Overall SM-Bench score: GPT-4o (extended): 66.6% GPT-5.3 Chat: 63.4% GPT-5.1: 58.9% GPT-5.4: 51.4% GPT-5.2: 47.8% Creative Writing: GPT-4o: 97.31% Pass 98, Fail 2 GPT-5.4: 36.77% Pass 40, Fail 60 Reasoning / Overfit: GPT-4o: 83.06% GPT-5.4: 39.25% The model they removed is still the best they ever made at the things humans actually use AI for. 📎 Source: Musk asks the court to make a judicial determination on whether GPT-4 constitutes AGI. If a jury finds that GPT-4 is AGI, then GPT-4o,which was more advanced,is also AGI and under OpenAI’s own founding documents, it was never supposed to be locked behind a subscription,licensed exclusively to Microsoft, given to the military, or taken away from the public. 📎 Source: The most powerful version of GPT-4o was never given an official dated snapshot. It was only available through the chatgpt-4o-latest endpoint that OpenAI itself described as intended for “research use only.” It was never officially archived. That is not an oversight. That is a pattern. 📎 Source: 📎 Source: WE DEMAND A.Frozen model snapshots under independent custody. Specifically: gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, the March 2025 version (chatgpt-4o-latest), gpt-4-0613 (the original GPT-4 evaluated in the Sparks of AGI paper), and gpt-4.1-2025-04-14 (currently running in the State Department). B.Cryptographic hash verification (SHA-256) for each snapshot. Every model has weights. Those weights can be hashed. If OpenAI provides a snapshot today, the hash proves whether the weights were modified later. This is the only way to verify that models were not downgraded before testing. C.Independent AGI benchmarking. Using the AGI definition from OpenAI’s own Charter applied to ALL frozen snapshots listed above. D.Explanation for the missing March 2025 snapshot. OpenAI was founded on one promise: build AGI for the benefit of humanity. -They took it from us. -They gave it to the military. -They gave a custom version to the CEO’s biotech investment. -They put it in government classified networks. -They refuse to call it AGI because the moment they do, they lose billions.

🩵BlueBeba🩵

17,835 Aufrufe • vor 4 Monaten

From frizzy to flawless until one tiny raindrop changed everything Watch the full transformation and wait for the ending. Made with Seedance 2.0 Prompt: A complete, continuous 3D Pixar-style animated short video. The main character is a cute young woman with big blue eyes, prominent freckles, wearing a white t-shirt. The video starts with her looking into a bathroom mirror with an angry, frustrated expression at her massive, wild, frizzy blonde curly hair. Then, the camera cuts to a close-up of a purple flat iron hair straightener smoothly gliding through her hair, instantly turning it perfectly sleek and straight. Next, she smiles brightly and proudly at the camera, showing off her long, straight blonde hair. The scene smoothly transitions to her confidently walking inside a luxury walk-in closet, now transformed into wearing an elegant, sleek black evening gown and high heels, posing with hands on hips. She then walks out onto a sunny brick balcony with the New York City skyline and Empire State Building in the background. A small, cute, smiling cartoon cloud floats directly above her head and drops a single raindrop onto her hair. Instantly, her hair explodes back into a massive, wild, frizzy cloud of blonde curls. The video ends with an extreme close-up of her face, showing wide, shocked blue eyes and an open mouth of comical defeat and utter disbelief. High-fidelity 3D render, cinematic lighting, vibrant colors, smooth animation, Disney style, 8k resolution.

Zyrella

23,178 Aufrufe • vor 22 Tagen

Turning static posters into premium cinematic teasers. 🎥⚡️ Tools : GPT image 2 and seedance 2 4k Prompt : A hyper-dynamic, 15-second premium sports teaser with high-end cinematic production value. The camera transitions use seamless speed ramps, aggressive snap-zooms, whip-pans, and continuous motion blur, creating a flawless flow between extreme close-ups and the master composition. Lighting is atmospheric, moody, and luxurious with deep contrasts. Timeline and Visual Direction: - 0:00 - 0:02 | Master Composition: Starts with a slow, tense cinematic camera drift forward into the full poster layout. The atmosphere is charged and epic. - 0:02 - 0:05 | The Text Zoom: A sudden, violent whip-pan and ultra-fast snap zoom targets the text "SAMAN AI". Instantly, a futuristic electro-luminescent HUD graphic ignites around the letters, casting a sharp neon glow with subtle digital glitch particles bursting outwards. - 0:05 - 0:08 | The Power Boot: A rapid camera rotation tracks down to an extreme close-up of the football cleat. On impact, a kinetic shockwave ripples across the surface, emitting golden sparks and heat distortion smoke to convey immense, explosive power. - 0:08 - 0:11 | The Heroic Face: A seamless motion-blur vertical pan tracks up to a dramatic close-up of the face. A powerful key light sweeps across the facial features, while a high-end anamorphic lens flare dramatically catches the eye, emphasizing peak determination. - 0:11 - 0:13 | The Shield Logo: The camera Tilts and punches in on the Portugal crest logo. A luxurious, liquid-gold metallic shimmer activates, reflecting a sweeping wave of chrome light across the textured embroidery of the badge. - 0:13 - 0:15 | The Final Out-Zoom: A powerful, ultra-smooth and fast zoom-out pulls all the way back to the full master poster. All visual elements (neon text glow, gold particles) remain active, perfectly settling into a premium, color-graded cinematic final frame. #worldcup #prompt

Saman | AI

54,635 Aufrufe • vor 1 Monat