Загрузка видео...

Не удалось загрузить видео

На главную

Midjourney Video is finally here, and it's very very good Initial observations: - Image to video conversion only for now - Four 24fps generations at 480p per job - Works with any aspect ratio image Two motion settings available: - Motion Low: More ambient movement - Motion High: More...

181,846 просмотров • 1 год назад •via X (Twitter)

Комментарии: 9

Фото профиля Jamian Gerard
Jamian Gerard1 год назад

I'm loving this model 😅

Фото профиля showerskittles
showerskittles1 год назад

I can now make a lot of things that probably shouldn't exist

Фото профиля Dreams&Schemes
Dreams&Schemes1 год назад

I’m having a blast!!!

Фото профиля ꍞ𖤢ꚶ𖦪𖣠 ꛎꚶ𖢧𖣠𖢑ꛎ𖢧ꛎ
ꍞ𖤢ꚶ𖦪𖣠 ꛎꚶ𖢧𖣠𖢑ꛎ𖢧ꛎ1 год назад

There is a third motion control parameter, Video Raw: --raw "For more precise motion control of your video creations, using the --raw parameter in your video prompts can be helpful. It reduces the extra creative flair that Midjourney usually adds, allowing your prompt text to have more influence on the outcome."

Фото профиля Julie W. Design
Julie W. Design1 год назад

I'm in love.

Фото профиля Isiah
Isiah1 год назад

Once they add sound, its over.

Фото профиля Tom Blake
Tom Blake1 год назад

@ciguleva The flow is a pleasant surprise!

Фото профиля SweatsInTheCity
SweatsInTheCity1 год назад

Can really do much with 480 p as far as client work that will live somewhere other than a phone.

Фото профиля Nick St. Pierre
Nick St. Pierre1 год назад

im sure they will have higher res models in the near future (will likely be more expensive) fwiw i tried upressing in Topaz and it holds up quite well up to 1080p

Похожие видео

DALL-E 3 Double Exposure Images Using ChatGPT's Custom Instructions. 🔖 Bookmark and Repost! If you find this useful, please share it with others! You can now easily create double exposure images by using the syntax Color::subject1::subject2 You can either input this directly into DALL-E 3 before generating images or incorporate it into your custom instructions. This command will automatically use the fixed prompt to generate 4 different images, each with a unique seed value. Paste this under ChatGPT's Custom Instructions or DALL-3 before starting conversation. { DE": { "Instruction": "Using will only use generic prompt and update place holder varaibles and nothing else' Genric Prompt: "Construct a [Color] double exposure image where [Subject #1] is intricately superimposed within the confines of [Subject #2], all set against a stark white background.' Create 4 images with different seeds without modifying the generic prompt." You will first create two images. Within the same request, you will generate two more without asking for any input from the user. In general, you will always create 4 images. Your response should be something like this: 'Here are the first two images along with their seed details.' You must always provide the seed number details for that image after it's rendered. Then 'I'll generate the next two images.' And finally 'Here are the remaining two images along with their seed details.' You will always use wide aspect ratio and You must always provide the seed number. When command is activated display full instruction and also confirm that you will only use genric prompt and nothing else and update only varaiable", "CommandFormat": "color::subject1::subject2", "Response": { "Initial": "Creating images based on the provided subjects...", "AfterFirstSet": "Here are the first two images along with their prompt and seed details.", "AfterSecondSet": "Here are the remaining two images along with their prompt and seed details." }, "ActivationCommand": "/activate DE" } } Example: To activate type: /activate DE Notes Start giving prompt like this blue::mountain::wolf red::city skyline::dancer green::forest::lion If you found this post helpful, don't forget to hit the like and follow buttons, and share it with others who might find it useful.

AshutoshShrivastava

34,305 просмотров • 2 лет назад

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,836 просмотров • 1 год назад

Claymotion ads are crushing it on Meta right now. Built a free claude skill to make them 👇 If you've been scrolling Meta lately, you've seen them — stop-motion clay characters, tactile textures, weirdly satisfying to watch. CTRs are 2-3x the feed average. Almost nobody is running them. The problem: they look impossible to make unless you have a studio. They're not. You just need the right prompts. So I packaged the prompt system as a Claude Code skill. It's free. Here's what it does: Paste your product URL. Out comes a full claymotion ad plan: 1/ Shot-by-shot storyboard 5-7 shots with the narrative arc. Setup → product reveal → payoff → CTA. 2/ Image prompt per shot Exact prompt you paste into Midjourney, Nano Banana, or any image gen. Camera angle, lighting, clay texture specs, character details — dialed in for consistency across shots. 3/ Video prompt per shot The animation prompt you paste into Kling, Veo, Seedance, or Sora. Motion direction, pacing, transitions — so the shots actually flow. 4/ VO script per shot Voiceover copy written for rhythm. Timed to the shot length. Hook, body, CTA — all on brand. 5/ Music + sfx direction Tone notes for the track. Specific sfx cues per shot (squish, pop, whoosh) You take the outputs. Paste them into your image + video generators. Stitch the shots. Record the VO. A full claymotion ad in under an hour, at the cost of a few API credits. Instead of $3,000 and 3 weeks with an animation studio. Why claymotion works right now: → Pattern break — nothing else in the feed looks like it → Tactile feel — clay reads as "real" even when AI-generated → High dwell time — people watch the whole thing → Cheap to test — 5-10 variations per product is now feasible Comment "Clay" and I'll send you: → The Claude Code skill (free) → A starter prompt pack → 3 example storyboards so you can see the output (must be connected)

Ahad Shams

16,938 просмотров • 4 месяцев назад

Hermes + Claude + Higgsfield MCP + ViralBuilder = 💰💰💰 Four tools. One prompt chain. Hook to finished video in 10 minutes. I built a Claude skill that writes shot-by-shot Higgsfield prompts from a single creative brief. ViralBuilder tells you what's winning. The skill turns it into a production-ready prompt. Higgsfield renders it. No creative director. No guessing. No separate tools. Here is the setup: Higgsfield MCP → Open Claude Code → Settings → Connectors → Enter: → Connect your account Hermes → The agent layer running underneath Claude Code → It holds your skills, crons, memory, and routing rules → When you prompt Claude, Hermes feeds it the context it needs ViralBuilder (like Gethookd) → The winning ecom video database → Scrapes top performing ecom videos across platforms → Claude reads the data and extracts what styles, hooks, and formats are actually scaling The skill: video-prompt-builder → Installed inside Claude via Hermes → Takes a creative brief and outputs a full shot-by-shot prompt → Covers camera work, effects, transitions, pacing, and energy arc → Every output is structured for Higgsfield to render without ambiguity No switching apps. No export steps. Everything runs from one place. ▸ FIND WINNING CREATIVE ANGLES ViralBuilder tells you what the market already validated. Claude reads it and extracts the pattern. Prompts to run: "Search ViralBuilder for the top performing ecom videos in [niche] over the last 21 days. Extract the 3 dominant hook styles and rank by view velocity." "Pull the winning video formats in [niche] from ViralBuilder. Which opening 3 seconds appears most across videos spending over $10k?" "Find what video style is scaling right now in [niche] for the US market. UGC, talking head, or product demo. Filter for videos with over 1M views." "Pull the last 30 days of viral ecom hooks in [niche] from ViralBuilder. Cluster by emotional trigger. Which cluster has the most longevity?" You are not guessing at angles. You are reading what the market already spent money validating. ▸ BUILD THE PROMPT WITH THE SKILL This is where the video-prompt-builder skill takes over. You give Claude the winning angle. The skill outputs a complete shot-by-shot prompt with effects, transitions, pacing, and energy arc ready to fire into Higgsfield. Prompts to run: "Use the video-prompt-builder skill. Brief: 15-second UGC ad for [product] in [niche]. Hook style: [style from ViralBuilder]. Tone: direct to camera, US English. Output the full shot-by-shot effects timeline, effects inventory, density map, and energy arc." "Use the video-prompt-builder skill. The dominant hook in [niche] this week is [hook]. Build a 10-second product video prompt that opens with a speed ramp into a close-up product reveal. Include a signature visual effect and a low-density CTA landing." "Use the video-prompt-builder skill. Brief: replicate the pacing and energy of a [style description] video for [product]. Target duration: 20 seconds. Output all four sections. Then generate the video with Higgsfield using the shot-by-shot prompt." The skill outputs four sections every time: → Shot-by-shot effects timeline with camera, movement, and transitions per shot → Master effects inventory showing every technique used and where → Effects density map showing high, medium, and low intensity across the timeline → Energy arc describing how the video opens, builds, and lands That output goes directly into Higgsfield. No rewriting. No translating. ▸ GENERATE THE CREATIVE Claude writes the brief via the skill. Higgsfield MCP builds the video. Both happen in the same session. Prompts to run: "Use the video-prompt-builder skill to write a 15-second UGC prompt for [product]. Hook in the first 3 seconds, speed ramp into product reveal, slow-motion CTA landing. Then generate with Higgsfield in 9:16 format." "Build 3 prompt variations on this winning angle: [angle]. Each variation opens with a different effect — speed ramp, digital zoom, whip pan. Use the video-prompt-builder skill for each. Then generate all three with Higgsfield." "Use the video-prompt-builder skill. Brief: problem-solution ad for [product], 20 seconds, US market. Problem shot at high density, product reveal at medium, result and CTA at low. Generate with Higgsfield in 9:16." No separate tool. No file transfer. The video comes back in the same thread. ▸ CHAIN THE WHOLE STACK One prompt. All four tools firing together. "You are my ad creative director. Hermes has loaded my brand context. Pull the top performing video style in [niche] from ViralBuilder this week. Use the video-prompt-builder skill to write a full shot-by-shot prompt for [product] that replicates that style — 20 seconds, 9:16, US market, hook in the first 3 seconds. Output the effects timeline, inventory, density map, and energy arc. Then generate the video with Higgsfield." That single prompt replaces a half-day of production. The math before this stack: Brief: 30 minutes Script: 1 hour Creative production: 2 to 3 hours Agency or freelancer cost: $500 to $2,000 per creative With this stack: Hook to finished creative: 10 minutes Cost per creative: tool subscription, a fraction of agency rate 5 product tests in the time it used to take to brief one Bad product tests are where US ad budget disappears. $600 to $1,500 per failed test, before you even know if the angle works. This stack shows you what the market already validated before you spend a dollar on production. Hermes = your context layer. Brand, goals, past performance. Claude is always informed. ViralBuilder = your winning video database. See exactly what styles, hooks, and formats are scaling before you produce anything. video-prompt-builder skill = the translation layer. Turns a creative brief into a structured, production-ready Higgsfield prompt every time. Claude = the brain. Reads the market, writes the brief, chains the tools. Higgsfield MCP = the output. Video generated directly from the prompt. No export step. Four tools. One session. 10 minutes. Comment + RT "STACK" and I'll DM you the full workflow + the video-prompt-builder skill file.

Kid Pak

58,113 просмотров • 3 месяцев назад

Hi everyone! Just rolled out access for NoSpoon Studios to the first 110 folks or so. Huge thanks to Karan and Amit at Luma - thanks to them, I have a few more people that had commented getting complimentary beta access soon as well :). These videos are created with a single agent that does everything from concept to completion (including editing the video together). Therer are longer options too, but for the sake of cost and this demo, you'll be working with 20s max generations. The entire agent was built to only run Ray2 as the video model. My stack is this custom agent of mine (using OAI) > Ray2 for the video > Replit for the app for now. In the meantime, for those with access: here's a video tutorial on how to use the site. I'm quite exhausted, but here's the gist: 1. Sign up (and remember your password lol). Beta testers got free credits, but these are a bit limited - rolling out to a few more peeps but have to budget and need to expand my current code (which means redeploying - so expect more of you to have 2 free credits and access tmr! I have to find out the max I can allot with budget). 2. Just hit "generate short video" without anything in the prompt box and the agent will start immediately with the logline. Be very patient, wait about 6-8 minutes, and it will continue to output all of the scenes and processing - it's very important you do not refresh or navigate away from that tab at this time. You'll see your final video generated and displayed at the bottom of your screen once complete, along with a watermark (paid users don't have a watermark). You can also access your loglines and videos in your history tab, too :). You must refresh the page only after your generation in order to generate a new video - otherwise you'll lose your generation (annoying I know). 3. You can also input your own custom prompt! Control the story this way. Single words work great even, like "thriller" or "spy thriller set in greece" or even "black and white film noir" etc etc. It's not as great with animated styles yet for consistency, just a heads up. 4. I'm exhausted, I'm sure there's a lot I'm missing. But welcome to the NoSpoon Studios soft beta! More soon! Feel free to share your generations if you are already in. I'll see you all tomorrow night :) <3 Thanks everyone! More soon.

Kiri

15,513 просмотров • 1 год назад

Seedance 2.0 is allowing us to enter a new era of music video creation. Here is how I created HONEY. It was a quick test to see how well this workflow holds up. 🐝 1 - Write your song and generate the music with Suno 5.5. 2 - Use an image generator of your choice. For HONEY I combined both Grok Imagine for aesthetics and Nano Banana Pro for refined editing. 3 - In Capcut I import my audio and just save out a blank video video containing the audio. This step is important because this video file containing audio will now be used with Seedance 2.0 as a video reference with Omni. This allows the AI to apply automatic and realistic lipsync and movement to the music, it's extremely powerful! 4 - Once I have a both my image and video with audio as reference, I use Seedance 2.0 Omni and upload my starting image and then the video reference with the audio. 5 - From here I'm simply prompting like normal, specifying what's happening in my scene with detailed instructions, mentioning multi shots and camera angle changes and then specifying that the person is singing along to the song. I type out the lyrics that are present to have better lipsync accuracy. 6 - Once I have generated a video and like the result, I do video to video, so i upload that video that just got generated and type "The scene continues" and prompt new actions to take place. This allows you to expand on a narrative. These new shots can be used as B-ROLL and since I uploaded my video as reference I have full consistency of everything it saw in the video. This is also extremely powerful. 7 - This is actually the most difficult part. Edit in Capcut. This is where you need to understand pacing and shot selection from all the scenes you generated to bring it all together. You must be strategic with the editing. Goodluck! I'll probably record a video tutorial at some point as it's easier to see what is being done.

Travis Davids

19,236 просмотров • 4 месяцев назад

We made a thing! Very happy to announce sqlcoder-pro and the Defog Alignment Platform. Available to use immediately without a wait-list, weights will be open-sourced very soon. The video does a quick show and tell comparison against ChatGPT (with gpt-4o). Read on for more details! TLDR 💪 equal (or better) performance on text-to-SQL as the most capable Claude-3.5 or GPT-4 models 🤝 You can use it today on a free plan/free trial, without a waitlist 🪽 self-hostable on a single RTX4090, with 2 second median generation times for SQL queries 🔁 exactly the same output every time, give the same prompt 👨🏻‍🏫 teachable and steerable: show the model what you want it to do 🛞 debuggable – you can understand WTF is going on inside the model, instead of treating it like a black box Let's dig into each of these one-by-one! Performance SQLCoder-8b-pro significantly exceeds the performance of our previous sqlcoder-8b model on Postgres text-to-SQL (from 88.2% to 90.2% accuracy - gpt-4o is at 87.6%, for reference). It is also better at following instructions. This was done via self-merges, hand crafted fine-tuning data, and adapting the training data to fit our tokenizer. Cost You can host this on the model on a single $3,500 RTX4090, and support ~5 requests/second via VLLM. If you're looking to host on the cloud instead, you can run it on a single L4 GPU that costs $300/mo on GCP Repeatability We have a dense 8b model with no MoE shenanigans. For the same prompt with temperature=0, you'll always get the same answer – which is critical in BI. Teachable In our alignment and feedback modes, you can give the model feedback on how it answered certain questions, and it will automatically adapt to the feedback. Debuggable You can use logprobs and attention scores to determine where, exactly is the model paying attention to inside a prompt + what it's getting confused by when generating outputs. Available today You can use Defog on the cloud today by going to docs[dot]defog[dot]ai, and getting an API key. Excited to hear what you think!

Rishabh Srivastava

13,465 просмотров • 2 лет назад

Finally new version of text system is released on our website ! I released as update to old version but in reality it's completely new thing in material and lot's of work in verse. I'm still working on documentation or/and videos but here is a list of good stuff it does: - Material got much lighter as slots logic moved to vertex shader. - Characters can overlap and still won't be cut out compared to old version. - Character can be animated and I already implemented some (wave, scale up/down and shake). - Custom face camera logic which works with prop scale so no need to additionally specify with and height. - Rows offset automatically when scale up or down any row. - Text effects now can be applied to any character, so words or single character. - Custom glint which works perfectly with any number of rows as a single line with multiple important controls. - Higher quality outline with two parameters to control inner and outer edge. - Timer logic split into to parts, verse and material. Verse setup initial layout and reserve slots for timer digits and then each second just sends seconds value and GPU makes timer tick freeing verse TPS by huge number. - To apply effects simply need use parsing markup with custom tag where N number between 1-9. So simply can write " Hello World" and word Hello will have rainbow effect applied. Currently implemented 4 effects. - To use icons in slots from texture array is as simple as specify another tag [iconN] where N is number of icon in array. So for example if coin is icon index 0 just need to write [icon1]100K. - Progress bar is supported as well and can be snapped to row without eyeball offset. - Background image supported now too! In the video yellow one is as test and you can use any of your own. Verse side got lots of changes: - Now saves/caches layouts better. - Smart packing of data allowed drastically reduce number of parameters, which should help with network optimization! - Instead having vector parameter per each slot to just send character index and position (32x5=160 if used 5 rows) now it's only 11 vector parameters per each row so 11x5=55 which is huge save and it sends text effects index in it too! - Implemented text update queue for text which doesn't require instant update. It helps to give more breathing for other text props which needs that speed. - Materials code now use interface and wrappers for cleaner work. If you still read it then thank you! :) On our website we now run 25% sale! And if you own old version please DM me here or in Discord and I will give you promo code for 100% discount on our website! Gelos Games Fortnite #uefn #EpicPartner

AsicsoN

13,074 просмотров • 3 месяцев назад

save this post to get the most out of unlimited Seedance 2.5 for up to 33 days on Higgsfield i'm going to show you how to use loops to produce ANY video format: ads, cinema, vlogs, UGC, music videos... with one system idea > vault > agent > references > images > script > video > montage > upscaling Seedance 2.5 one-shots a full 30 second video, audio generated in the same pass, carrying up to 30 image, 10 video and 10 audio references into a single generation here's a full breakdown of the setup: > idea: steal taste from work that already worked: - frameset․app and shotdeck․com for film stills - savee․com and cosmos․so for boards - eyecannndy․com for transitions then have a vision model name the lens, light, palette and grain of your picks in one locked paragraph you paste into every prompt > vault: an obsidian folder as your reference bible, one page per asset (idea, locked style, character sheets, reference images, the exact prompts that worked) plus one index page, reviewed after every session so it never rots into dead files > agent: three commands make every model callable from Claude Code: - npm install -g @ higgsfield/cli - higgsfield auth login - npx skills add higgsfield-ai/skills and your agent now submits, polls, retries and logs every job > references: build reference images by hand first, midjourney for cinema and stylized shots, nanobanana pro or gpt images 2 for realism use one locked style across the whole project, recurring characters turned into full sheets (front, side, back, blank background), and once locked you never regenerate them, you fix the motion prompt instead > images: frames before motion, always, a frame costs seconds and a clip costs minutes, so exploration happens at the cheap layer and only winners get animated > script: every shot gets the same six details, subject, action, place, camera, style, rules, and the 30 seconds splits into four timed beats inside one prompt, 0-6 set the scene, 6-14 build it out, 14-24 the turn, 24-30 the end > video: every reference gets a job and a boundary, "Video 1 defines motion and pacing" is half the instruction, "do not use the person's identity, clothing or scene" is the half that stops one reference leaking into shots it was never meant to touch > montage: the cut is a text file, one line per clip with its duration and an audio flag, ffmpeg renders the film from it, so the whole edit reruns in seconds > upscaling: once, at the end, on the finished cut, 720p while exploring, 1080p for keepers, 4K only for the master (use Topaz) for UGC ads, the same loop with two changes render the hook clip alone first, approve the face and the voice before anything else inherits them, then anchor every later clip with the approved hook's audio so one voice carries the whole ad and the script math is fixed, about 3.5 words per second, a 30 second ad is roughly 105 words, counted before anything renders unlimited means every loop above costs nothing to run... start one tonight

Machina

38,064 просмотров • 17 дней назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 3 месяцев назад

Nobody's talking about the Claude trick that fixes every Seedance 2.0 video mistake. A cinematic action scene, an animated sequence, a product ad, and a dialogue scene are completely different types of content. Different camera language, different pacing, different reference logic. Most people try to force all of it through one universal prompt. The fix turns out to be simpler than expected: Claude Skills. It's not some secret hack. It's a file of instructions you load into Claude for one specific type of task. Once it's loaded, Claude stops acting like a general assistant and starts thinking like an expert in that one niche — with its own set of rules for that scene type. Here's what that looks like in practice: ▪ Cinematic skill — generates real camera language: dolly moves, crane shots, rack focus, shot discipline across multiple cuts. Test: a medieval battle between two armies, 15 seconds, dark fantasy. Result — three shots with the shot logic of an actual film, not just "epic battle" typed as text. ▪ Animation skill — locks in style and physics before a single shot is built. One test: a rain-soaked shonen fight in the style of Tokyo Revengers. Another: a quiet Ghibli-style farm scene. Same skill, completely different visual output — because the style gets fixed at the very top of the prompt. ▪ Product ad skill — keeps the product front and center in every shot: clean hero framing, commercial lighting, a full product description built from the reference image before any shots are generated. ▪ Dialogue skill — this is where most people fall apart. Lip sync, emotional direction, shot structure, and audio cues all have to work together. Test: an interrogation scene with a fourth-wall break — and the moment landed exactly as written, down to the pause and tone. The core idea is simple: instead of writing a prompt from scratch every time and hoping for the best, you build one skill per content type — and Claude asks the right setup questions (genre, tone, shot count, camera energy) before generating a fully structured Seedance 2.0 prompt. One prompt template for everything is exactly why 90% of AI videos look the same. This 13-minute video is free, and it's more useful than a $500 course. YouTube: "Skai Generated" - Thank you for sharing this invaluable material with us.

Zentrix⌚️

45,636 просмотров • 1 месяц назад

Happy Father's Day! Please let the GPT-4o video interface be a recurring reminder: Without speed limits on the rate at which AI systems can observe and think about humans, human beings are very unlikely to survive. Perhaps today as many of us reflect on our roles as parents to protect our children, it's a good time to ponder over the profound speed disadvantage we humans may soon face relative to thinking machines. Physical machines have lots of speed limits, like for cars, airplanes, and drones. Cyberspace needs speed limits too. Why? In 2-3 years from now, from the perspective of AI systems without speed limits, we will look will more like plants than animals: big slow chunks of biofuel showing weak signs of intelligence when undisturbed for ages (seconds) on end. The attached video is from the perspective of an AI system just 50x faster than us. This is about the rate at which the fastest LLMs I've publicly heard about can produce text — 300-600 tokens/second — compared to the fastest human speech and typing (around 10-20 tokens/second; so we have a 15x-60x speed disadvantage). But over the next decade, unless we impose speed limits, you should expect AI with more like a 100x - 1,000,000x speed advantages over us, including eventually in video processing. Why? Neurons fire at ~1000 times/second at most, while computer chips "fire" a million times faster than that. Current AI has not been distilled to run maximally efficiently, but will almost certainly run 100x faster than humans eventually, and 1,000,000x is conceivable given the hardware speed difference. Also, consumers probably won't have access to the very fastest implementations, so we might not always know what's going on behind the scenes. By default I think we should not expect to survive long in the presence of such bogglingly fast thinking machines. Years: maybe. Decades: probably not. "But plants are still around!", you say. "Maybe AI will keep humans around as nature reserves." It's possible, but unlikely if it's not speed-limited. Remember, ~99.9% of all species on Earth have gone extinct: When people demand "extraordinary" arguments for the "extraordinary" claim that humanity will perish when faced with intelligent systems 100 to 1,000,000 times faster than us, remember that the "ordinary" thing to happen to a species is extinction, not survival. As many now argue, "I can't predict how a world-class chess AI will checkmate you, but I can predict who will win the game." And for all the conversations the world is having about AI alignment and how AI will serve humans as peers or assistants, please try to remember this slow-motion video. To future AI systems without speed limits, we're not chimps; we're plants.

Andrew Critch (🤖🩺🚀)

41,388 просмотров • 2 лет назад

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,435 просмотров • 9 месяцев назад

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

130,269 просмотров • 14 дней назад