正在加载视频...

视频加载失败

What a crazy week in AI! 🚀 MiMo V2 Flash Gemini 3 Flash Qwen Image Layered Trellis 2 TurboDiffusion SCAIL HY World 1.5 StereoSpace LongVie 2 RealVideo V-RGBX Ray 3 Modify Kling Motion Control SVG-T2I Wan 2.6 Seedance 1.5 Pro LongCat Video Avatar Flux 2 Max GPT Image 1.5...

19,186 次观看 • 7 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Gemini-1.5 Pro has its spotlight stolen today, and people are poking fun at Sora vs Google memes. Well, I think it's the biggest boost in LLM capability so far in 2024. v1.5's 10M token context (1) excels at retrieval; (2) generalizes zero-shot to extremely long instructions like full tutorials and codebases; and (3) works across modalities such as text, audio, and video. Here's a stunning example: v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning. I talked about the Myth of Context Length many times before: don't get too excited by claims of 1M or even 1B context tokens. LSTMs already achieved literally infinite context length 25 yrs ago! What truly matters is how well the model actually uses the context to solve real-world problems, and Gemini-1.5 has surpassed the SOTA with flying colors. The paper is also well-written with lots of solid quantitative analysis on in-context memorization and generalization. Paper: “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context” Congrats to Jeff Dean Oriol Vinyals Sundar Pichai and team!

Jim Fan

278,483 次观看 • 2 年前

🎬 Kendi kısa reklamını, filmini veya ürünü AI ile tanıtmak artık hayal değil! Hemde Çok Kolay!! Kaydet Kesin Lazım Olur !!! Bu videoda gördüğünüz HER ŞEY (Görseller, geçişler, müzik, 3D animasyonlar) tamamen yapay zeka ile yapıldı. Peki mutfakta neler var? İşte Sümela Manastırı videosunun arkasındaki "Workflow" ve kullanılan araçlar: 🧵👇 🛠️ Araç Çantası:🧠 Planlama & Akış: Gemini 🎨 Görsel Üretimi: Nano Banana Pro (Nano Banana 2 Lite) & each::labs 🎥 Görselden Video (Text-to-Video): Kling 2.6 🎞️ Start/End Frame Video: Kling O1 (Kling AI) 🎵 Müzik: Suno AI (Suno) ✂️ Edit: CapCut (CapCut) 📝 Adım Adım Tarif: 1. Adım: Görselleri Planla ve ÜretGemini'ye stratejiyi verdim: "Bana Kız Kulesi için FPV drone ile çekilmiş gibi, biri havadan iniş (başlangıç), diğeri içinden bir görüntü (bitiş) olacak şekilde iki görsel prompt hazırla." Prompt 1 (Dış Çekim): A breathtaking high-altitude FPV drone photograph diving steeply towards the Sümela Monastery... Prompt 2 (İç Çekim): An immersive interior photograph capturing the final destination... inside the Cave Church... 2. Adım: Videoya Dönüştür (Sihir Burada!) ✨Kling O1 Image-to-Video modeline Başlangıç ve Bitiş görsellerini yükleyip şu prompt'u girdim: 🚀 Video Prompt: An FPV drone shot starting from the input image's perspective, accelerating rapidly downwards... 3. Adım: 3D Metin EkleKling 2.6 ile havada süzülen metinler oluşturduk: 🔠 Text Prompt: Slowly rotating 3D text, floating in the air, cinematic lighting... 4. Adım: Atmosferi Müzikle TamamlaSuno AI ile epik bir başlangıç ve mistik bir final: 🎵 Music Prompt: Cinematic, Epic Orchestral, intense buildup, sudden transition to Ethereal, Gregorian Chants Diğer Posta giderek detaylara bakabilirsiniz. Kaydet Lazım OLUR....

Tugser Okur

62,302 次观看 • 7 个月前

Claymotion ads are crushing it on Meta right now. Built a free claude skill to make them 👇 If you've been scrolling Meta lately, you've seen them — stop-motion clay characters, tactile textures, weirdly satisfying to watch. CTRs are 2-3x the feed average. Almost nobody is running them. The problem: they look impossible to make unless you have a studio. They're not. You just need the right prompts. So I packaged the prompt system as a Claude Code skill. It's free. Here's what it does: Paste your product URL. Out comes a full claymotion ad plan: 1/ Shot-by-shot storyboard 5-7 shots with the narrative arc. Setup → product reveal → payoff → CTA. 2/ Image prompt per shot Exact prompt you paste into Midjourney, Nano Banana, or any image gen. Camera angle, lighting, clay texture specs, character details — dialed in for consistency across shots. 3/ Video prompt per shot The animation prompt you paste into Kling, Veo, Seedance, or Sora. Motion direction, pacing, transitions — so the shots actually flow. 4/ VO script per shot Voiceover copy written for rhythm. Timed to the shot length. Hook, body, CTA — all on brand. 5/ Music + sfx direction Tone notes for the track. Specific sfx cues per shot (squish, pop, whoosh) You take the outputs. Paste them into your image + video generators. Stitch the shots. Record the VO. A full claymotion ad in under an hour, at the cost of a few API credits. Instead of $3,000 and 3 weeks with an animation studio. Why claymotion works right now: → Pattern break — nothing else in the feed looks like it → Tactile feel — clay reads as "real" even when AI-generated → High dwell time — people watch the whole thing → Cheap to test — 5-10 variations per product is now feasible Comment "Clay" and I'll send you: → The Claude Code skill (free) → A starter prompt pack → 3 example storyboards so you can see the output (must be connected)

Ahad Shams

16,779 次观看 • 3 个月前

🔥HOLY SMOKES! $TAO holders! 🚀 SUBNET 19 (VISION) ON BITTENSOR IS ABSOLUTELY CRUSHING IT! In my 5+ years covering crypto and AI, this is one of the most impressive implementations I've seen. The combination of scale, performance, and decentralization is absolutely next level! 🚀 @namoray_dev @Corcel_X 💨 INSANE Speed Performance: - Llama 3.1 8B: 196.18 tokens/s with +107.23% advantage - Llama 3.1 70B: 124.96 tokens/s with +154.96% advantage - Llama 3.2 3B: 166.69 tokens/s with +21.66% advantage 🔥 Top Tier Model Integration: - Meta-Llama-3-70B & 8B Instruct - FLUX.1-schnell for Text-to-Image - ProteusV0.4-Lightning (Text & Image) - Multiple model variations for redundancy 🔥 What Makes This INSANE: - Complete decentralization - No single point of failure - Multiple model choices for redundancy - Real-time performance tracking - Transparent incentive structure The incentive distribution curve shows a healthy network with: - Strong rewards for top performers - Fair distribution across all participants - Clear path for growth and improvement - Sustainable economic model What's truly MIND-BLOWING is how they've managed to: 1. Scale to millions of operations 2. Maintain high quality across multiple tasks 3. Create a fair, competitive marketplace 4. Build in redundancy and reliability 5. Achieve true decentralization This isn't just another subnet - this is the future of decentralized AI inference happening RIGHT NOW! 🔥 1. MASSIVE Scale & Adoption: - We're seeing 7M+ tokens being processed - 14K+ processing steps being executed - Multiple AI models running simultaneously - Incredible miner participation across the network 2. Revolutionary Task Distribution: - Llama 3.1 70B leading with 20% weighting - Avatar Generation at 15% - Perfectly balanced task distribution for optimal network performance - Multiple specialized tasks including Text-to-Image and Image-to-Image processing 3. Elite Performance Metrics: - Top miners hitting 0.00775 incentive rates - Consistent performance across the network - Impressive scaling from top to bottom performers - Strong incentive curve maintaining network quality 📈 Network Performance: - Consistent upward trend in tokens/s - Quality scores maintaining high levels (>0.9) - Steady improvement in miner performance - Rock-solid network reliability ⚡ Platform Highlights: - Permissionless, serverless architecture - Global network of Always-On GPUs - Instant API access - Full decentralization - Multi-model support with seamless switching What makes this TRULY SPECIAL is the consistent upward trajectory in both speed and quality, while maintaining a decentralized architecture. The performance advantages over industry standards (+154.96% for 70B!) are absolutely mind-blowing! 🚀 This isn't just another AI subnet - it's a glimpse into the future of decentralized AI inference! The combination of speed, reliability, and model variety makes this one of the most impressive implementations in the space! 🔥 📽 Watch Now on YouTube and TikTok: Source 🔗

Andy ττ

11,616 次观看 • 1 年前

RIP Arcads 🤯 I built a Claude skill that vibe-codes UGC ads on demand. One product + one prompt = the AI creator, the script, the scene-by-scene shot list, and finished video. All inside Claude. Perfect for DTC brands and agencies who can't afford to keep paying $500-$1,500 per UGC video and waiting 2 weeks for revisions. If you're briefing creators, mailing PR boxes, waiting on first cuts, then asking for re-shoots because the hook didn't land... This skill eliminates the entire loop: → Tell Claude the product, ad angle, and length → Skill writes the GPT Image 2.0 prompt to generate the AI creator from scratch → Skill writes every scene prompt, dialogue line, and delivery direction → Pipe it into Seedance 2.0 with character + product + voice locked → Speed up + caption in CapCut → Ship the ad in 20 minutes No more paying $11 per video on Arcads. What you get: → Perfect character consistency across every scene → Voice consistency that holds clip-to-clip (small CapCut hack inside the tutorial) → Real product fidelity using your actual product photo as a reference → Multi-scene day-in-life, testimonial, and action-shot formats out of the box Built 100% with a Claude skill + Seedance 2.0. I recorded a full step-by-step tutorial showing the exact workflow + the 3 finished AI UGC ads I made for Rhode, AG1, and Barebells. Want to see the full breakdown? > Like this post > Comment "UGC" And I'll send it over (must be following so I can DM)

Mike Futia

34,117 次观看 • 3 个月前

if you want to create an AI channel with guaranteed views do this: this is the 2nd channel i see in less than a week that has done the same thing🚨 1. identify a channel that is already successful 2. take screenshots and send them to Gemini, 3. ask it “for a prompt that imitates that visual style” 4. use it every time you generate an image 5. Veo 3.1 or Kling 2.1/2.5 turbo for image to video 6. basic editing for example, this is the visual style i would use at Calliopelabs: { "channel_visual_style": { "aesthetic": "High-End 3D CGI Documentary", "tone": "Dark Tech, Investigative, Cinematic, Mystery", "render_style": "Octane Render, Ultra-realistic textures, 8k, Ray-tracing", "core_elements": { "characters": { "type": "Featureless Mannequins", "material": "Polished gray-white metal, smooth, reflective, seamless", "features": "No face, no eyes, no mouth, abstract representation of humans" }, "environment": { "style": "Minimalist 3D stages, clean abstract voids", "background": "Slightly blurred (Bokeh), uncluttered, focus on subject" }, "lighting": { "mood": "Cinematic & Dramatic", "palette": "Cool tones (Deep Blues, Cyans, Emerald Greens), Neon accents, High Contrast", "technique": "Volumetric lighting, Rim lights, Studio setup" }, "camera": { "style": "Macro photography feel, shallow depth of field, cinematic angles", "movement": "Slow, deliberate, floating, dolly-in" } } } } paste it into the Visual Style, create a template (voice over, music, effects etc..) and automate video the generation download the video and review it, create a thumbnail with Nano Banana Pro imitating that channel mercilessly - upload the video - get views - repeat until one channel explodes if you want me to better explain how to do this with my tool ask for it in a comment and i’ll write you ⬇️

Sergio Gil

50,733 次观看 • 6 个月前

$VET, #VeFam. In this video, I demonstrate in less than 4:30 minutes how to create an AI agent on veworld(.)ai. Watch me build a Mr. Robot Monologue Writer agent. If you haven't seen Mr. Robot, I suggest you watch it! This is just early bird access. The options for tools and integrations and such are limited, but what exists is already working quite well. The process is easy peasy. The UI is simple, but effective. It asks you for... 1. Role & Purpose 2. Voice & Style 3. Behavior 4. Rules 5. Tags 6. Avatar image 7. Welcome text. 8. Test drive before publication. ... and that's about it. This free version lets you have at most 3 agents, I am told. This implies that there is also a paid version. I'm all for it, because it sounds to me like VeChain is ready to do real business! I am providing feedback to Jérôme Grillères in order to help improve VeChain's AI agent marketplace. I didn't have to set up anything. The web UI is all I needed! The agent is running on Claude Sonnet 3.7. I did not have to provide a Claude API key. We seem to be riding along on VeChain's. I hope there'll be a choice for more models, including ChatGPT, in the future. This is so user friendly, that I can easily imagine that this would take off in a big, big way. I'm definitely building on this, when it goes into production with full features. Even if my own AI agents aren't successful, then I'm sure others' will be. And that means the $VET / $VTHO / $B3TR flywheel is going to take off in a big, big way. I, for one, am here for it. (See the reply below for the listing of the AI agent I just created.)

₿lackthorne AI

16,562 次观看 • 1 个月前

Lucas Crespo (Lucas Crespo 📧 ) is the mastermind behind Every 📧 's visual vibe—and he does it one prompt at a time. As our creative lead, Lucas uses tools like native image gen in ChatGPT and Midjourney to generate the cover images you see every day. He also designs the interfaces for our products—Cora, Spiral, and Sparkle—and makes everything on our site feel as thoughtful and delightful as possible. We get into: - Why Every’s aesthetic feels familiar and new at the same time. Every’s aesthetic plays with the tension between the old (like Greek statues and Baroque symbols) and the new (like saturated colors and modern motifs) to make the glamor of the past feel fresh. - Art direction matters more than ever today. As AI makes it easier to generate images, Lucas says the real work of design is shifting toward art direction, specifically, curating an aesthetic that feels “organic;” on his X timeline that’s showing up as clouds, earthy landscapes, and textures. - Reimagining what a website can be with AI. Lucas compares most websites to identical buildings—predictable, efficient, and forgettable—and wonders how AI can help us break that mold by designing experiences that prioritize serendipity over speed, and curiosity over control. - Behind the scenes of Cora’s visual aesthetic. How Lucas designed the landing page and launch video for Cora by rooting it in the product’s philosophy: turning the inbox from a source of chaos into something that feels calm, thoughtful—like stepping into spring. - The future of internet interfaces. Lucas believes the future of digital interfaces will be curated with the same care as a film set or ad campaign, where every detail is chosen with intention. Lucas also walks us through how he created the headline image for Every’s consulting page—a human and robotic hand fist-bumping—using Midjourney to iterate from rough prompt to polished visual. This is a must watch for anyone interested in the future of design and making the internet a little more beautiful every day. Watch below! Timestamps: 1. Introduction: 00:01:41 2. How AI changed the course of Lucas’s career: 00:04:02 3. Why Every’s aesthetic feels both familiar and fresh: 00:11:38 4. Why Lucas thinks minimalism is overrated: 00:16:19 5. Art direction matters more than ever in the age of AI: 00:22:01 6. How to reimagine what a website can be with AI: 00:25:06 7. Lucas’s process in Midjourney to generate cover images: 00:33:41 8. Midjourney v. image generation in ChatGPT: 00:43:01 9. Behind the scenes of Cora’s design language: 00:49:43 10. How AI is rewriting the role of a designer: 00:59:57

Dan Shipper 📧

16,038 次观看 • 1 年前

🚨12 HOUR NEWS RECAP 1.⁠ Trump slammed the media for asking about the Epstein coverup after AG Pam Bondi was asked why a minute was missing from the CCTV tape of his cell and if Epstein had worked for any intelligence agencies: “Are people still talking about this guy, this creep? That is unbelievable.” 2.⁠ Tucker warned that the Epstein coverup could spark a revolution: “How can you say thousands of children were raped, but I'm not gonna find out who raped them? I don't want a revolution, but if you wanted a revolution, this is how you would act.” 3.⁠ Flash flooding hit New Mexico as the Rio Ruidoso surged from 1.5 feet to over 20 feet in under an hour, sweeping homes away. At least 3 people have died in the flooding so far. 4.⁠ Ukraine faced the largest drone assault of the war - and possibly in history. According to Ukraine’s Air Force, Russia hit it with 728 drones and 13 missiles overnight. Of those, 296 were shot down and 415 jammed or lost to electronic warfare. 5.⁠ At least 3 people are dead after the Houthis sunk a Greek-operated cargo ship in the Red Sea. 5 crew members have been rescued so far, and a search is underway for the rest of the crew. 6.⁠ A wildfire sparked by a burning car has turned into a full-scale disaster in southern France - torching 720+ hectares, damaging homes, and choking Marseille under a blanket of smoke. 7.⁠ Tom Barrack, Trump's trusted envoy and ambassador to Turkey, warned that while Trump backs Lebanon with strength and clarity, that support won’t last forever. The option of protracted discussions, until after the country's elections in May 2026, isn't on the table. 8.⁠ A bridge collapsed near Vadodara, India, sending multiple vehicles plunging into the river below. The death toll has risen to 10, with rescue efforts still underway. 9.⁠ In a rare video appearance, jailed PKK founder Abdullah Ocalan has called for a full end to the group’s armed insurgency against Turkey. “The phase of armed struggle has ended,” he said, urging a shift to democratic politics and disarmament. 10.⁠ Russia is testing a new AI drone armed with thermal vision and onboard target selection. It flies without GPS, shrugs off jamming, and adapts to the battlefield in real time, selecting targets autonomously.

Mario Nawfal

104,421 次观看 • 1 年前

📖SEEDANCE 2.0 JUST MADE EVERY FILM SCHOOL IRRELEVANT FOR SOLO CREATORS Solo creators with the right workflow are closing clients that used to require a full production studio. Seedance 2.0 inside Dreamina holds character consistency across scenes in a way no other tool at this price point comes close to. Same face, same costume, same lighting logic — frame after frame after frame. That's the feature that turns a single prompt session into a short film. The lava demon materializing inside a gothic cathedral. The girl in black holding her ground while everything burns around her. Two characters with completely different visual languages sharing the same atmospheric world — and Seedance holds both of them consistent across every cut. That's not a generation. That's a production. Here's the 7-step workflow that produced this: • Step 1 — Define the character before you define the scene. Write a complete physical description — face structure, hair, clothing, skin, posture. This becomes the anchor every future generation references. • Step 2 — Build the world separately from the character. Gothic cathedral, candlelight, fog, cracked stone, scattered bodies. Define the atmosphere as its own entity before you place anyone inside it. • Step 3 — Generate the reference frame. One image that establishes the visual language, the color grade, the lighting temperature. Lock this as your style reference before generating any video. • Step 4 — Feed the reference into Seedance's image-to-video pipeline with a motion prompt. Camera behavior only — slow push, hold, circle. The image handles the subject. The prompt handles the direction. • Step 5 — Generate four variations per scene. Delete the two that look generated. Keep the one where the character's face holds and the atmosphere feels physical rather than rendered. • Step 6 — Edit in CapCut or Premiere Pro. Add music that matches the emotional temperature of the grade — the visual already tells you what the sound should feel like. Dark orchestral, slow tempo, single instrument carrying the melody. • Step 7 — Save the character description and reference frame as a template. The next episode starts from the same character in the same world. Series content becomes a system, not a restart. How a freelancer sells this: Dark fantasy content for game studios, music artists, and fantasy brands is a real market with real budgets. A musician dropping an album needs a visual world. A game studio needs promotional cinematics. A fantasy brand needs a story. Where to find clients: • Music artists on SoundCloud and Spotify releasing dark, gothic, or cinematic albums — search by genre, find artists with 1k–50k listeners who have no visual content. They have the audience and the need but no production budget for traditional video. • Indie game studios on Itch and Steam launching fantasy or horror titles — they need promotional cinematics and trailers but can't afford a production company. A single free scene built from their game's character art opens every conversation. • Dark fantasy and gothic brands on Instagram and TikTok with strong photo content but zero video presence — jewelry brands, clothing labels, occult lifestyle brands. They have the aesthetic already built. You just add motion to it. • Fantasy and horror fiction authors on Instagram and Substack launching new books — they need visual teasers, trailers, and world-building content to build pre-launch audiences. Most have no idea this kind of production is accessible at this price point. • Tabletop RPG creators and Dungeon Masters on Patreon and Kickstarter — they build entire fantasy universes and need cinematic content for campaigns, promotional videos, and subscriber rewards. The niche is underserved and the creators inside it spend consistently on content tools and services. 📥Tomorrow I'll show you what's sitting right next to this opportunity. 🔖Save this if you are looking for practical AI methods that actually pay.

Zentrix⌚️

50,098 次观看 • 1 个月前

HERMES AGENT SUPPORTS 300+ MODELS. PICKING THE RIGHT ONE PER TASK IS THE DIFFERENCE BETWEEN $5/MONTH AND $50. STARTING OUT: Claude Sonnet 4.6. official recommendation from Nous Research. "the model this project was built and tested with." strong reasoning. reliable tool calling. mid-range pricing. PREMIUM TIER: Claude Opus 4.8. best coding benchmarks available. self-correcting reasoning. catches its own mistakes. 1M context. use for demanding tasks where quality matters. GPT-5.5. #1 Chatbot Arena. #1 GPQA Diamond reasoning (94.1%). #1 creative writing. 2M context. handles entire codebases in one pass. Grok 4.30. the only frontier model with live X firehose access. real-time social data, breaking news, market sentiment. connects via Grok OAuth. no separate API key. Grok-Composer-2.5-Fast (v0.17.0). Cursor's coding model. 200K context. available through your Grok subscription via OAuth. no extra cost if you already pay for Grok. MID-RANGE TIER: Claude Sonnet 4.6. best balance of quality and cost for daily use. strongest prose and tool calling in this tier. Gemini 2.5 Pro. Google Search grounding built in. cites sources. verifies claims. pulls current data. 2M context. best for research-heavy workflows. GPT-4.1. reliable tool calling. solid general reasoning. good middle ground when you need OpenAI compatibility. BUDGET TIER: Claude Haiku 4.5. fastest Anthropic model. cheapest paid Claude option. strong at classification, routing, simple queries. use for auxiliary tasks: compression, vision, web extraction, approval scoring. DeepSeek V4. best cost-to-quality ratio in the market. 90% cache discount on repeated context. use for sub-agents and bulk parallel work. DeepSeek V4 Flash. cheapest paid model worth using. 1M context. MIT license. self-hostable. use for cron jobs, monitoring, routine searches. MiniMax M3. Nous Research and MiniMax collaborating on optimization. 1M context via lightning attention. 59% SWE-Bench Pro. beats several premium models on coding. one of the most-used models inside Hermes. FREE / LOCAL: Qwen 3.5 27B via Ollama. 16GB VRAM. reliable tool calling. best free local model for Hermes as of mid-2026. Qwen 3 8B. 8GB VRAM. fits a $7 VPS. handles routine tasks at zero API cost. Llama 4 Maverick. best open-weight tool calling. 1M context. needs more VRAM but strongest local option. HOW TO ASSIGN MODELS: main model: Desktop app / Dashboard → Models → switch sub-agent model: set in Desktop app, Dashboard, or config.yaml: delegation: model: "deepseek/deepseek-v4" auxiliary models (compression, vision, web extract): Desktop app / Dashboard → Models → Auxiliary Haiku 4.5 or Gemini Flash work well here. saves significantly when your main model is premium. per-profile: each Hermes profile gets its own model. Scout on DeepSeek. Analyst on Sonnet. Briefer on budget model. Coder on Opus. per-cron-job: pin a specific model to any cron job. morning brief on Haiku. deep research on Sonnet. monitoring on DeepSeek Flash. each job uses only the model it needs. per-session: /model deepseek/deepseek-v4-flash hot-swap mid-conversation. no restart needed. FALLBACK CHAINS: if your primary model is unavailable, Hermes automatically switches to the next provider. rate limit or server error = next model in the chain. no failed runs. no manual intervention. set in Desktop app, Dashboard, or config.yaml: fallback_providers: - openrouter - nous - codex PROVIDER PATHS: OPENROUTER: 300+ models under one API key. pay per token. most flexible. NOUS PORTAL: 300+ models + Tool Gateway (web search, image gen, TTS, browser). one OAuth. one subscription. 10% off token-billed providers. CHATGPT SUB: GPT-5.5 + Grok via OAuth. included tokens with $20 subscription. OLLAMA: free. local. private. zero API cost. your hardware only. mix providers across profiles and tasks. Scout on OpenRouter. Analyst on Nous Portal. Coder on ChatGPT sub. Monitor on Ollama. THE RULE: premium for work that needs deep reasoning. mid-range for daily driver tasks. budget for volume and background work. free for monitoring and routine jobs. pricing changes fast. check openrouter ai for current rates before committing. Which is your favourite model and for what task? full 15 levels breakdown in the article 👇

YanXbt

17,138 次观看 • 1 个月前

Voice AI turn taking is a solved problem. The single most common complaint about voice AI, today, is that agents interrupt too often. But the voice agents I build for myself now respond quickly and interrupt me less often than the people I talk to every day. (I actually measured this.) Mark Backman made a Pipecat AI PR two weeks ago that was the last piece of the puzzle for turn taking so good that I no longer ever think about it. The approach combines three layers of processing: 1. Voice activity detection, with a short (200ms) trigger. 2. A native audio turn detection model that's small, fast, and runs on CPU. This model captures audio nuances like inflection and filler sounds that don't get transcribed. 3. A prompt mixin for the conversation LLM that decides turn completion based on conversation context. None of these are new. We've been using VAD for a long time. We trained the first version of the Pipecat Smart Turn native audio model in December 2024. And we've been experimenting with prompt-based large model turn detection (sometimes called "selective refusal") for more than a year. Now, the Smart Turn model and the SOTA LLMs we're using in voice agents have both gotten so good that using them together feels like we've finally "solved" turn detection. Mark also figured out how to elegantly apply a "single-token tagging" technique to this problem. We sometimes use single-token tagging in place of tool calling, when we need a near-zero latency programmatic trigger. Mark's Pipecat mixin defines three single-token characters and prompts the LLM to output exactly one of them at the beginning of every response. - ✓ means the agent should respond normally (immediately) - ○ is a "short incomplete" - the agent should wait 5 seconds - ◐ is a "long incomplete" - the agent should wait 10 seconds The wait times, and the details of the prompt, are configurable, of course. Watch the video to see me talk to an agent that handles all my various pauses and inflections, plus phrases like "let me think," pretty much the way a person would handle them, in terms of response latency. Also, in the second half of the video, I ask the agent to adjust its response pattern because I'm going to tell it a phone number. This kind of "in-context" adjustment of response wait times is really useful. The LLM in the video is GTP-4.1. We've tested the prompt and single-token adherance with GPT-4.1, Gemini 2.5 Flash, Anthropic Claude Sonnet 4.5, and AWS Nova 2 Pro. Note that older models in all these families (and, in general, smaller open weights models) aren't able to reliably output these single-token tags. But the new models we're using these days are pretty amazing.

kwindla

26,935 次观看 • 5 个月前