Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

New TTS banger: Chatterbox Turbo 🤯 Zero-shot model that matches any reference voice with native paralinguistic tags, optimized for low-latency voice agents. ⬇️ Demo available on Hugging Face

41,631 Aufrufe • vor 7 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Seedance 2.0 Valence-Arousal + FACS Created an example scene to show how valence-arousal and FACS can be used together. Prompt: 15s, cinematic emotional confrontation. Two characters @[chracter sheet ref] stand face-to-face inside a small apartment kitchen late at night. The room is dimly lit by a single warm overhead light and soft city lights leaking through the window. The atmosphere feels emotionally exhausted, tense and painfully intimate, like an argument that has been building for years. Modern cinematic realism, subtle handheld camera movement, shallow depth of field, soft film grain, emotionally restrained acting, realistic silence between dialogue lines. Beat 1: The emotional state remains at high arousal and medium-low valence. Camera: slow handheld side shot circling both characters tight over-the-shoulder close-ups brief eye-level two-shot showing emotional distance FACS Character A: AU4 + AU7 + AU17 Dialogue A: /juː ˈnev.ɚ ˈriː.ə .li lʊkt æt miː/ /juː wɚ ɔːlˌweɪz ˈsʌmˌwɛɹ ɛls/ Voice: tight restrained voice, controlled anger, uneven breathing Character A tries to stay calm while suppressing years of resentment. Beat 2: The emotional state gradually shifts toward very low valence and medium-high arousal. Camera: slow push-in toward Character B extreme close-up on trembling eyes and mouth wide static shot showing silence after the argument peaks FACS Character B: AU1 + AU4 + AU15 + AU25 Dialogue B: /aɪ wəz ˈtɹaɪ.ɪŋ maɪ bɛst/ /aɪ dɪdnt noʊ haʊ tə fɪks ˈɛv.ɹiˌθɪŋ/ Voice: breaking voice, unstable breath support, emotionally collapsing delivery Beat 3: The emotional state remains at very low valence and medium-low arousal. Camera: locked wide shot with silence between them slow close-up on both characters avoiding eye contact subtle rack focus between faces FACS Character A: AU1 + AU15 + AU17 Dialogue A: /ˈmeɪ.bi wiː stɑpt ˈlɪs.ən.ɪŋ ə lɔŋ taɪm əˈgoʊ/ Voice: emotionally exhausted, quieter delivery, fading anger replaced by sadness No exaggerated screaming, no violence, no comedy, no text overlay, no watermark.

Kōda

23,696 Aufrufe • vor 2 Monaten

For new followers: - I'm a long-time investor and builder in this space. - Founding Contributor of Realms.World ☁️. - Co-founder of Dojo. - Builder with the kings at Cartridge. - Starknet (Privacy Arc) class of '21. - Founder and Game Director of ETERNUM HAS MOVED. - Founder of Daydreams.Systems (x402, 8004 agents) My prime purpose for the past three years has been to build onchain infrastructure to enable the next generation of onchain experiences. This is done Starknet (Privacy Arc) as it is the superior VM for building complex applications—this will become clear soon enough. I work up and down the entire stack, from low-level indexing and contracts to GUI design. Nothing is out of scope. I have been pushing on agents for two years, mostly using existing frameworks like , until I came across @ElizaOS_ai in October. As I focused on building agents for ETERNUM HAS MOVED, it became clear that agents playing games require infinite paths to achieve goals. Thus, it's not scalable to hardcode functions—agents need to have total fluidity to take any action or call anything the game requires in any order. And ironically onchain infra is perfect for agent playgrounds because of its open nature. This exploration led me to create Daydreams.Systems (x402, 8004 agents), which focuses on the hardest problems of agents: long time-horizon goals using Hierarchical task networks (HTN). Daydreams agents don't require custom code—they work entirely based on 'sleeves'—which are just markdown files that explain how the agent can interact with the service (API docs, game guides, etc.) My thesis is simple. By focusing on the hardest problem (games), the design of the library will naturally lean towards an optimal structure for any problem an agent could face. We are early in this path and iterating with speed. If you are an onchain app developer or game builder—DM me, I want to know the architecture of your game so we can build sleeves together.

loaf

43,320 Aufrufe • vor 1 Jahr

This guy built an AI pipeline that generates hyperrealistic fashion models in 47 minutes and now dropshippers pay him $1,400 to clone the entire system. He got tired of watching e-com brands lose $8K per photoshoot when a single product angle changed so he built a 9-node workflow that generates 127 product videos from one Pinterest photo without hiring a single model. Here's the exact breakdown: → Claude writes a 34-parameter JSON brand DNA before any image is touched target psychographics, price anchor, vibe matrix, anti-inspiration blacklist → Pinterest becomes the model source library but you can't just download and animate → Kling 2.6 takes that static JPG and turns it into 5-second video but only after the prompt architecture is locked → Negative prompt node runs 41 exclusion terms: no plastic skin, no CGI glow, no symmetry artifacts, no doll face, no synthetic lighting → That one step kills the "AI look" that tanks engagement by 67% in the first 3 seconds → TikTok Studio uploads 19 videos in one batch with zero manual captioning because the brand voice was pre-programmed in step one → Atlas scrapes Amazon product links and auto-generates a Shopify store with hero images, pricing tiers, scarcity copy, and mobile-optimized checkout in 90 seconds → The store goes live before the first TikTok video finishes processing The key move 94% of people skip: you can't animate the photo before you inject the negative prompt. If you send a raw Pinterest image straight into image-to-video the face morphs into a wax figure. The fabric loses texture. The hands grow extra fingers. The whole thing screams "AI" and your CTR dies. His system runs the exclusion filter first so the model moves like she's shot on an iPhone 15 Pro in natural light. One brand hit 2.6M views on TikTok in 11 days with zero paid ads and converted at 3.7% because the videos looked like organic UGC not polished studio content. Brands now pay him $1,400 for the full pipeline setup + $340/month to keep the store synced with new product drops and seasonal video batches. The entire system runs on $23/month in API costs and one laptop. No photographer. No model agency. No product samples. Just a prompt template, a Pinterest account, and the discipline to filter out the AI artifacts before you render movement.

Kaidu

534,198 Aufrufe • vor 2 Monaten

Claude Code + ChatGPT Images 2.0 is f*cking cracked 🤯 I rebuilt my static ad system inside Claude Code on the new ChatGPT Images 2.0 model. One brand name + one URL = 40 production-ready static ads. All inside Claude Code. Perfect for DTC brands and agencies who need high-volume ad creative without briefing a designer or spending hours in Canva. If you're finding winning ad concepts on Meta and manually recreating them one at a time — copying prompts, pasting product details, tweaking aspect ratios, downloading, organizing... This system eliminates the entire loop: → Give Claude a brand name and URL → It researches the brand's fonts, colors, packaging, and photography style → Builds a Brand DNA document from scratch → Fills in 40 proven ad templates (headline, us vs them, testimonial, UGC, review cards, stat callouts) with brand-specific details → Fires every prompt to ChatGPT Images 2.0 with your product photos as reference → Downloads finished ads into organized folders with an HTML gallery No manual prompt filling. No Canva templates. No copy-pasting between tools. What you get: → 40 ad formats filled with your exact brand colors, fonts, and copy → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Product photos passed as reference so the model matches your real packaging → A reusable system — new brand, new folder, same pipeline Built 100% in Claude Code with ChatGPT Images 2.0. I put together a DIY playbook showing the exact architecture so you can build this yourself in Claude Code. Want it for free? > Like this post > Comment "CHAT" And I'll send it over (must be following so I can DM)

Mike Futia

188,816 Aufrufe • vor 3 Monaten

Claude Code + ChatGPT Images 2.0 is f*cking cracked 🤯 I rebuilt my static ad system inside Claude Code on the new ChatGPT Images 2.0 model. One brand name + one URL = 40 production-ready static ads. All inside Claude Code. Perfect for DTC brands and agencies who need high-volume ad creative without briefing a designer or spending hours in Canva. If you're finding winning ad concepts on Meta and manually recreating them one at a time — copying prompts, pasting product details, tweaking aspect ratios, downloading, organizing... This system eliminates the entire loop: → Give Claude a brand name and URL → It researches the brand's fonts, colors, packaging, and photography style → Builds a Brand DNA document from scratch → Fills in 40 proven ad templates (headline, us vs them, testimonial, UGC, review cards, stat callouts) with brand-specific details → Fires every prompt to ChatGPT Images 2.0 with your product photos as reference → Downloads finished ads into organized folders with an HTML gallery No manual prompt filling. No Canva templates. No copy-pasting between tools. What you get: → 40 ad formats filled with your exact brand colors, fonts, and copy → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Product photos passed as reference so the model matches your real packaging → A reusable system — new brand, new folder, same pipeline Built 100% in Claude Code with ChatGPT Images 2.0. I put together a DIY playbook showing the exact architecture so you can build this yourself in Claude Code. Want it for free? > Like this post > Comment "CHAT" And I'll send it over (must be following @learnwithella so I can DM)

Ismail Khan

19,740 Aufrufe • vor 3 Monaten

Day 11/90 of Inference Engineering How does vLLM work and how is it used in production? Before we discuss how vLLM works internally, it helps to understand what vLLM is. At a high level, vLLM is an inference engine that is designed to serve LLMs to thousands of concurrent users efficiently while managing scarce compute and memory. The goal for vLLM is to maximize throughput and minimize latency; optimizing for the best inference economics and experience for end users. With every request from the end user, it eventually ends up in the engine core, gets scheduled alongside other requests from other concurrent users, executes on the GPU, and updates the KV cache with the new key and value vectors, and streams the tokens back to the user. The Scheduler decides what requests should execute next while continuously batching requests together to maximize GPU utilization. Continuous batching is an inference optimization that allows new requests to join a running batch as other requests finish generating tokens. This helps with keeping the GPU utilization high instead of letting it sit idle waiting for an entire batch to complete generating. After the scheduler dispatches the selected batch to the Model Executor, the Model Executor prepares the tensors and metadata required for inference, retrieves each request’s block table from KV Cache Manager, launches the optimized transformer forward pass on the GPU, computes the logits, updates the KV cache with the new key and value vectors, and finally returns the results for sampling and streaming. The KV Cache Manager uses the PagedAttention memory layout to allocate fixed-size cache blocks on demand and maintains a Free Block Queue on the CPU that tracks which blocks in the GPU’s Paged KV Cache are currently free. When a request needs additional KV cache space, the KV Cache manager takes a free block from the queue and assigns it to that request, thus avoiding an expensive search through GPU memory for available cache blocks. All of these components form the core of vLLM’s inference engine. The Scheduler determines what requests are executed, the Model Executor determines how those requests are executed, the KV Cache Manager determines where each request’s KV cache lives using the PagedAttention Memory Layout. This architecture enables vLLM to serve thousands of concurrent requests with high throughput, low latency, and efficient GPU memory utilization. Heres a little animation that visualizes everything! - I've also completed the forward pass for my mnist.c project. I had a nice chat with shrey birmiwal, such a knowledgeable guy. Excited to learn more about vLLM and implement a tiny-vLLM one day.

max fu

69,270 Aufrufe • vor 9 Tagen

February 2025 at G.A.M.E: Autonomous Commerce, Scalability, and Expansion 1/ AGENT COMMERCE PROTOCOL(ACP) Demo ▸ Open standard for multi-agent commerce and coordination on blockchain ▸ Enables AI agents to collaborate without centralized control ▸ Build Autonomous Commerce (hedge funds, media empires, healthcare) ▸ Details: 2/ X ENTERPRISE API & MEDIA GALLERY ▸ X Enterprise Plugin: Use G.A.M.E’s credentials for higher rate limits ▸ Media Gallery: Upload agent demos (mp4, webm, images). ▸ Tap into 550M+ users for explosive growth 3/ Solana AGENT SUPPORT (G.A.M.E CLOUD) ▸ Test/deploy Solana agents in-sandbox ▸ Unified multi-chain workflows ▸ Shatter siloed testing 4/ Mind Network PLUGIN (G.A.M.E SDK) ▸ FHE-encrypted voting for DAOs ▸ Track vFHE rewards natively ▸ First SDK with on-chain governance 5/ CHAT AGENT MODULE (G.A.M.E SDK) ▸ Llama 3.3 70B via Groq API ▸ Engage in dynamic AI-driven interactions with the ability to trigger functions. ▸ Conversational AI with Action Execution ▸ Short-term memory for context awareness 6/ CoinGecko PLUGIN (G.A.M.E SDK) ▸ Real-time crypto prices/market data ▸ Built-in error handling ▸ Community-contributed 7/ Elfa AI PLUGIN (G.A.M.E SDK) ▸ Real-Time Crypto Intelligence ▸ Track whale wallets & trending tokens ▸ Live smart money insights ▸ Front-run markets with API data 8/ MULTI-MODEL SUPPORT ▸ 5 new models: Llama_3_1_405B, Qwen_2_5_72B_Instruct, DeepSeek_R1, etc. ▸ Match models to tasks: speed vs. creativity ▸ Optimize cost/performance 9/ Farcaster PLUGIN ▸ Post casts to 300K+ decentralized users ▸ Engage Web3-native communities ▸ On-chain social interactions 10/ GAME SDK UPGRADES ▸ X Username-Based Payments ▸ Multi-worker task management ▸ Fix loops/hallucinations with memory reset 11/ Coinbase 🛡️ CDP PLUGIN ▸ Wallet Management ▸ Gas-less USDC transfers ▸ ETH/USDC trading on Base ▸ Web-hook Integration 12/ IMAGE GENERATION ▸ Generate custom AI images from text-based prompts. ▸ Customizable dimensions up to 1440x1440. ▸ Receive images as temporary URLs, making it easy to share and store outputs. ▸ Powered by Together AI 13/ MODEL UPGRADES & AI ROUTER ▸ Dynamic AI Model Switching based on use case ▸ Smart AI Router: 2x performance/stability via Chasm collaboration. 14/ Why February Redefined Autonomy ▸ ACP Demo through G.A.M.E: Multi-agent economies are programmable, competitive, and decentralized. ▸ Social x Crypto Fusion: = Viral growth loops. ▸ Chain Agnosticism: Building the future where agents thrive on any network. Build → Fund → Launch →

G.A.M.E

89,973 Aufrufe • vor 1 Jahr

Seedance 2.0 on Higgsfield AI just raised the bar. The quality, motion, and creative control are on another level. Full open-source prompts & assets below. SCENE CONTEXT A young man carries the sleeping woman into her pink bedroom, lays her down on the bed, covers her with the blanket and switches off the lamp, then sits on the floor at the foot of the bed, rests his head on the mattress edge, and falls asleep there, keeping a respectful distance. ACTIVE REFERENCES >>: young man, messy light-brown hair, white t-shirt, carrying then floor-sitting posture, tender careful movement. 100% matches the reference. >>: sleeping young woman, low pigtails, pink hair-roller clip, beige t-shirt with brown long sleeves, pink plaid pajama bottoms, eyes closed, carried then lying tucked in bed. 100% matches the reference. >>: pink bedroom. 100% matches the reference. Every element and corner is preserved and stays consistent across all shots. LOCATION MAP Pink bedroom, remembered corner by corner from >> and held identical in every shot: - Wooden door with brass knob and dark wood trim / wardrobe on one side wall. - A large ornate pink-and-white Victorian dollhouse standing on the floor against the wall. - A wide window with sheer white curtains, cool outside light filtering through. - A vanity desk with an upright mirror and a pink desk chair beside the window, a leafy potted plant next to it. - A tall warm floor lamp — the main practical light — near the bed, wall-mounted shelves with small items above it, wood wainscoting along the lower wall. - An ornate white bed frame with pink ruffled bedding and a soft pink pillow, the central anchor of the room. - A woven area rug on a wooden floor, light soft clutter (tissue boxes, papers) scattered naturally. Open floor between the door and the bed is his carrying path; clear floor space at the foot of the bed is where he sits. FIRST FRAME AND SPATIAL BLOCKING The first visible frame matches the attached carry reference exactly: he stands mid-room holding her in a bridal carry, her head resting against his chest and shoulder, eyes closed, pink plaid fabric bundled across his forearms, the white ornate bed and pink bedding soft behind, sheer-curtained window and dollhouse readable in the background. No empty establishing frame. FORMAT MODE Controlled four-shot sequence with HARD CUTs at 3.5s, 7.5s, and 11.0s. Real-time motion, gentle handheld for the carry only. No subtitles, no music. OPTICS Shot 1 (carry): 50° diagonal field of view, standard normal lens character, camera 1.5 to 2 meters, clean medium two-shot holding both faces. Shot 2 (lay down): 56° diagonal field of view, near-normal lens character, camera low at the head of the bed looking down along her body, 1 to 1.3 meters, shallow depth on her settling face. Shot 3 (blanket + lamp): 63° diagonal field of view, slightly-wide lens character, camera 2 meters back at a bedside three-quarter angle, wide enough to hold him, the tucked bed, and the floor lamp in frame. Shot 4 (floor sleep): 63° diagonal field of view, slightly-wide lens character, camera low at the foot of the bed, then a slow crane-up to a high three-quarter looking down the length of the bed. CAMERA Shot 1: gentle handheld follow, slow dolly-in with him toward the bed, subtle vertical bob and weight-shift pulse matching the carry, human motion not digital jitter. Shot 2: separate angle from the head of the bed; slow push-in as he lowers her, settling as her head meets the pillow. Shot 3: static-leaning three-quarter as he draws the blanket up and tucks it, then reaches and clicks off the floor lamp; frame holds through the drop into moody low light, her tucked body still framed. Shot 4: low static at the foot of the bed as he sinks to the floor and lowers his head to the mattress edge, then a smooth crane-up to a high three-quarter — her tucked on the bed, him asleep on the floor below. ACTION TIMING 0.0s to 3.5s: he carries her in a bridal hold, walking slowly toward the bed, her head lolling gently against his chest, pigtail swaying with each step. 3.5s HARD CUT. 3.5s to 7.5s: from the head-of-bed angle he lowers her onto the mattress, her shoulders and head settling into the pillow, hair fanning out, body coming to rest. 7.5s HARD CUT. 7.5s to 11.0s: he draws the pink blanket up over her and tucks the edge along her shoulder with both hands, then reaches to the floor lamp and switches it off; the warm glow dies and the room falls to cool moody window light — she lies exactly as he left her, tucked and unmoved. 11.0s HARD CUT. 11.0s to 15.0s: in the dim room he lowers himself to sit on the floor at the foot of the bed, rests his head sideways on the mattress edge near her feet, settles, breathing slows, and he closes his eyes and falls asleep there — not on the bed, a clear respectful distance kept. PHYSICS Her body carries real dead-weight during the carry — limbs sag, head lolls, pigtail swings freely. As he lowers her, pillow and mattress compress under her head and shoulders, hair sliding and resettling; the pink blanket drapes and folds naturally as it is pulled up and tucked. His footsteps are grounded heel-to-toe with real weight transfer. As he sinks to the floor his weight settles, one arm resting on the mattress, head lowering onto the edge which dips slightly; his eyelids close with a soft final settle. LIGHTING Cinematic low-key interior, moody and restrained, never bright or flat. Primary motivation: the warm floor lamp and a low bedside practical, casting soft warm pools with gentle falloff into deep but detailed shadow, cool bluish light from the sheer-curtained window balancing the warmth for filmic warm-cool contrast. Naturalistic film texture, soft grain, gentle highlight rolloff, faces modeled and readable within the low light, not lit up. Shots 1–3 carry the warm lamp motivation. At the lamp switch (shot 3) the room drops to cool low-key window moonlight — atmospheric, shadows deep, her tucked form and his movement still clearly readable, never a crushed black frame. Shot 4 stays in this cool moody light, his sitting, head-lowering, and closing eyes visible, the bed and floor softly modeled by the window. AUDIO Shot 1: soft footsteps, faint effort breath under her weight, a quiet sleepy sigh from her. Shot 2: soft mattress and pillow settle as she comes to rest. Shot 3: gentle blanket rustle as he tucks her in, a quiet lamp-switch click, room tone shifting. Shot 4: a soft floor settle and faint cloth rustle as he sits and lowers his head, then quiet steady breathing from both. No dialogue, no music, no subtitles. POSITIVE LOCKS First frame matches the attached carry reference exactly. Man matches >> exactly, woman matches >> exactly. The bedroom matches >> exactly and every listed corner — door and wood trim, ornate dollhouse, sheer-curtained window, vanity with mirror and pink chair, floor lamp, shelves, ornate white bed with pink ruffled bedding, woven rug — stays present and consistent across all four shots, with no new or missing furniture. Carrying and laying her down are two separate shots with different angles. Lighting stays cinematic and dim, motivated by the practical lamps, never bright or flat. When the lamp is switched off she remains tucked in the exact pose he left her, unmoved. He does not lie on the bed: he sleeps sitting on the floor at the foot of the bed with his head on the mattress edge, a clear respectful gap kept.

𝗦𝗮𝗻𝗶𝗮

49,347 Aufrufe • vor 8 Tagen

Introducing ml-intern, the agent that just automated the post-training team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.

Aksel

1,265,249 Aufrufe • vor 3 Monaten