ben's banner
ben's profile picture

ben

@contraben15,214 subscribers

building @contra and @contralabs_ai

Shorts

For Anthropic’s Claude Design, your opening prompt sets the ceiling of the project. If your prompt is a comprehensive brief, it will land somewhere near production-ready Seems obvious, but here is something interesting we observed: one designer wrote: >16 prompts >climbed in specificity over the session >and still finished worse than they started A single "reconsider" prompt at iteration 2 tanked output quality by -0.7 and the session never recovered. 2/3 of every prompt typed across all 5 sessions was refinement or correction. TLDR: Your first move is the whole game. and prompting is an art. What a time to be alive.

For Anthropic’s Claude Design, your opening prompt sets the ceiling of the project. If your prompt is a comprehensive brief, it will land somewhere near production-ready Seems obvious, but here is something interesting we observed: one designer wrote: >16 prompts >climbed in specificity over the session >and still finished worse than they started A single "reconsider" prompt at iteration 2 tanked output quality by -0.7 and the session never recovered. 2/3 of every prompt typed across all 5 sessions was refinement or correction. TLDR: Your first move is the whole game. and prompting is an art. What a time to be alive.

229,536 Aufrufe

Fable is back (enjoy the final hours of in-sub usage) so we made it fight Anthropic's Opus 4.8: >5 real landing page and portfolio briefs, built in Claude Code, >judged blind by 9 working designers across 90 matchups. Fable's best brief won 88.9% of its head-to-heads. Its worst output only won 11.1%. The difference was the brief 🧵

Fable is back (enjoy the final hours of in-sub usage) so we made it fight Anthropic's Opus 4.8: >5 real landing page and portfolio briefs, built in Claude Code, >judged blind by 9 working designers across 90 matchups. Fable's best brief won 88.9% of its head-to-heads. Its worst output only won 11.1%. The difference was the brief 🧵

104,724 Aufrufe

In our Anthropic Claude Design study, 5 designers approved a design system before they typed their first prompt. >Brand palette >type system >components the whole thing all set up. Only 1 of them named any of it in their opening prompt. That designer was the only one to finish production-ready. The other 4 assumed Claude would carry the system over. It didn't. TLDR: Claude doesn't reliably carry the design system you just approved. If you don't name it in the prompt, it doesn’t exist. It's never been a better time to be a designer, but you must learn the art of the prompt.

In our Anthropic Claude Design study, 5 designers approved a design system before they typed their first prompt. >Brand palette >type system >components the whole thing all set up. Only 1 of them named any of it in their opening prompt. That designer was the only one to finish production-ready. The other 4 assumed Claude would carry the system over. It didn't. TLDR: Claude doesn't reliably carry the design system you just approved. If you don't name it in the prompt, it doesn’t exist. It's never been a better time to be a designer, but you must learn the art of the prompt.

163,778 Aufrufe

hey x, Contra Labs is hiring Researchers in SF/NY that are interested in making AI better for creativity. You will collaborate directly with the frontier labs on cutting edge projects that directly impact the tools creatives use every single day tag someone who could be interested!

hey x, Contra Labs is hiring Researchers in SF/NY that are interested in making AI better for creativity. You will collaborate directly with the frontier labs on cutting edge projects that directly impact the tools creatives use every single day tag someone who could be interested!

20,188 Aufrufe

What is the best video editing agent for short form social? Does it actually work? We watched professional video editors, step by step, as they built short-form social reels in Adobe Premiere Pro. Today we're open-sourcing this preview dataset on Hugging Face, to make AI agents better at editing videos. The data set is 234 annotated steps across 4 computer-use trajectories. Editors narrated their reasoning aloud as they worked, so every step pairs a screenshot with the expert's own thought, a structured action, and executable grounding: >a Premiere MCP tool call, keyboard shortcut, menu path, or coordinate click. >The format follows the AgentNet trajectory schema, extended with a Premiere action taxonomy and multi-path execution. ***That makes it directly usable for computer-use agent SFT, reasoning mid-training, tool-use and function calling, and benchmarking agents against a human expert baseline. Enjoy!

What is the best video editing agent for short form social? Does it actually work? We watched professional video editors, step by step, as they built short-form social reels in Adobe Premiere Pro. Today we're open-sourcing this preview dataset on Hugging Face, to make AI agents better at editing videos. The data set is 234 annotated steps across 4 computer-use trajectories. Editors narrated their reasoning aloud as they worked, so every step pairs a screenshot with the expert's own thought, a structured action, and executable grounding: >a Premiere MCP tool call, keyboard shortcut, menu path, or coordinate click. >The format follows the AgentNet trajectory schema, extended with a Premiere action taxonomy and multi-path execution. ***That makes it directly usable for computer-use agent SFT, reasoning mid-training, tool-use and function calling, and benchmarking agents against a human expert baseline. Enjoy!

39,867 Aufrufe

It's never been a better time to be creative. (ever) The two best frontier models (ever) have been released within 30 days of each other. Here’s what we learned from running 4 frontier models head to head. >Sol has taste and Fable takes direction. 10 identical landing page briefs, judged blind by working creatives. GPT 5.6 Sol won 82% on loose briefs. But when handed a real design spec it finished last. 🧵

It's never been a better time to be creative. (ever) The two best frontier models (ever) have been released within 30 days of each other. Here’s what we learned from running 4 frontier models head to head. >Sol has taste and Fable takes direction. 10 identical landing page briefs, judged blind by working creatives. GPT 5.6 Sol won 82% on loose briefs. But when handed a real design spec it finished last. 🧵

35,609 Aufrufe

It's safe to say Meta is back in the game. We ran a blind "style-transfer" tournament with AI at Meta's new Muse Image vs >OpenAI's GPT Image 2 >Google Gemini's Nano Banana Pro >BlackForestLabsAI - Unofficial's FLUX.2. on 10 real world briefs, 55 tournaments, with every output ranked blind by professional working creatives. GPT won the most, but Muse placed top-two in 59%, more than any other model, already beating Nano Banana Pro 👀 (Image reference created with Muse Image btw)

It's safe to say Meta is back in the game. We ran a blind "style-transfer" tournament with AI at Meta's new Muse Image vs >OpenAI's GPT Image 2 >Google Gemini's Nano Banana Pro >BlackForestLabsAI - Unofficial's FLUX.2. on 10 real world briefs, 55 tournaments, with every output ranked blind by professional working creatives. GPT won the most, but Muse placed top-two in 59%, more than any other model, already beating Nano Banana Pro 👀 (Image reference created with Muse Image btw)

36,620 Aufrufe

hey X fam Contra Labs is hiring multiple Strategic Project Leads in NY/SF to work on some very, very... cool projects tag someone who could be interested

hey X fam Contra Labs is hiring multiple Strategic Project Leads in NY/SF to work on some very, very... cool projects tag someone who could be interested

38,387 Aufrufe

who's at Config next week ? Contra Labs is hosting our annual "Config Kick Off Party" on Monday the 22nd for >creatives >researchers >designers to talk about the latest AI powered workflows, the role of human taste in creative work + all things creative intelligence. + a live panel with product leaders from Open AI, Qiver, Krea, and of course Figma See you there! (link in comments)

who's at Config next week ? Contra Labs is hosting our annual "Config Kick Off Party" on Monday the 22nd for >creatives >researchers >designers to talk about the latest AI powered workflows, the role of human taste in creative work + all things creative intelligence. + a live panel with product leaders from Open AI, Qiver, Krea, and of course Figma See you there! (link in comments)

29,461 Aufrufe

Anthropic just shipped Claude Opus 5. It's state of the art for coding, close to Claude Fable 5. So we re-ran our landing page study: >10 client briefs, >4 models, >10 expert designers, >600 blind comparisons, >400 write ups. TLDR: designers liked its words more than its pages. more in the 🧵

Anthropic just shipped Claude Opus 5. It's state of the art for coding, close to Claude Fable 5. So we re-ran our landing page study: >10 client briefs, >4 models, >10 expert designers, >600 blind comparisons, >400 write ups. TLDR: designers liked its words more than its pages. more in the 🧵

14,210 Aufrufe

What are the biggest failure points for AI generated landing pages ? >8 working designers. >40 AI landing pages. >754 failure points found. OpenAI’s Sol fumbles layout Anthropic’s Fable fumbles the finish xAI’s Grok fumbles interaction AI at Meta’s Muse Spark has room to improve across the board Every AI model breaks a landing page in its own signature way. We mapped all 4. 🧵

What are the biggest failure points for AI generated landing pages ? >8 working designers. >40 AI landing pages. >754 failure points found. OpenAI’s Sol fumbles layout Anthropic’s Fable fumbles the finish xAI’s Grok fumbles interaction AI at Meta’s Muse Spark has room to improve across the board Every AI model breaks a landing page in its own signature way. We mapped all 4. 🧵

17,255 Aufrufe

We benchmarked @openai’s ChatGPT Images 2.0 against the top image models. It won everything. Then we asked the harder question: can brand designers really ship with it? Here's where the model breaks.

We benchmarked @openai’s ChatGPT Images 2.0 against the top image models. It won everything. Then we asked the harder question: can brand designers really ship with it? Here's where the model breaks.

30,778 Aufrufe

Can AI really do logos ? Seems to be a hot topic. So, we ran 155 blind tournaments with 10 professional brand designers instead to dig in. >OpenAI's GPT Image 2 crushed the field. >Google Gemini's Nano Banana Pro, Microsoft AI's MAI Image 2.5, and AI at Meta's Meta Muse split the scraps. But still, winning a tournament and being production-ready are very different realities. 🧵

Can AI really do logos ? Seems to be a hot topic. So, we ran 155 blind tournaments with 10 professional brand designers instead to dig in. >OpenAI's GPT Image 2 crushed the field. >Google Gemini's Nano Banana Pro, Microsoft AI's MAI Image 2.5, and AI at Meta's Meta Muse split the scraps. But still, winning a tournament and being production-ready are very different realities. 🧵

15,816 Aufrufe

Kimi K3 just jumped 17 places to #1 on Arena.ai's Frontend Code Arena. So we gave it real client work to see where leaderboards meet reality: >10 landing page briefs, >judged blind by 8 working designers >against GPT 5.6 Sol, Gemini 3.5 Flash, and Claude Fable 5. It tied the best model in the study. And on detailed briefs, it won. Details in 🧵:

Kimi K3 just jumped 17 places to #1 on Arena.ai's Frontend Code Arena. So we gave it real client work to see where leaderboards meet reality: >10 landing page briefs, >judged blind by 8 working designers >against GPT 5.6 Sol, Gemini 3.5 Flash, and Claude Fable 5. It tied the best model in the study. And on detailed briefs, it won. Details in 🧵:

13,798 Aufrufe

We ran a blind head-to-head on the leading image models for one specific job: >product detail shots. Seedream 5.0 Lite (BytePlus, ByteDance) beat the flagship models from Google, OpenAI, and Black Forest Labs. It won 2 out of 3 times.

We ran a blind head-to-head on the leading image models for one specific job: >product detail shots. Seedream 5.0 Lite (BytePlus, ByteDance) beat the flagship models from Google, OpenAI, and Black Forest Labs. It won 2 out of 3 times.

26,202 Aufrufe

For Google Gemini, your prompt is the difference between Client V1 and Production Ready. We observed 10 designers going through a real world campaign workflow: >Hero stills >Social cuts >Secondary assets total: 29 deliverables. Only 24% were "Production Ready" And they all wrote prompts the same way. Below are the prompts and outputs 👇⬇️

For Google Gemini, your prompt is the difference between Client V1 and Production Ready. We observed 10 designers going through a real world campaign workflow: >Hero stills >Social cuts >Secondary assets total: 29 deliverables. Only 24% were "Production Ready" And they all wrote prompts the same way. Below are the prompts and outputs 👇⬇️

21,383 Aufrufe

Wow, killer work from the Krea team 👀 Krea 2 Large is the #2 style-transfer model, already closing in on GPT Image 2. Average Style Fidelity gap from Krea's latest model to OpenAI's GPT Image 2: 0.14 points. Gap from Krea to the rest of the field: more than 4x larger. Full study below.

Wow, killer work from the Krea team 👀 Krea 2 Large is the #2 style-transfer model, already closing in on GPT Image 2. Average Style Fidelity gap from Krea's latest model to OpenAI's GPT Image 2: 0.14 points. Gap from Krea to the rest of the field: more than 4x larger. Full study below.

20,158 Aufrufe

"Spotify Wrapped" was built with Rive. LinkedIn's "Year End Review" was built in Rive. The UI inside of every BMW is built with Rive. And now, Rive just launched Scripting! Want to try it out, and win $5k? Today we are launching the "Scripting with Rive" challenge on Contra. What are you going to build ?

"Spotify Wrapped" was built with Rive. LinkedIn's "Year End Review" was built in Rive. The UI inside of every BMW is built with Rive. And now, Rive just launched Scripting! Want to try it out, and win $5k? Today we are launching the "Scripting with Rive" challenge on Contra. What are you going to build ?

32,925 Aufrufe

Convergence on quality, divergence on taste. >Across our Contra Labs benchmark >12 frontier models from OpenAI Google DeepMind Black Forest Labs ByteDance >across 5 creative domains ~15,000 evaluator judgments this theme dominated.

Convergence on quality, divergence on taste. >Across our Contra Labs benchmark >12 frontier models from OpenAI Google DeepMind Black Forest Labs ByteDance >across 5 creative domains ~15,000 evaluator judgments this theme dominated.

16,690 Aufrufe

Can Adobe Firefly's "Edit" feature compete with Photoshop? Select a region, describe the fix, keep the rest of the frame intact. That's the pitch. Here’s what we observed across 4 sessions and 8 targeted edits with real working creatives: >1 edit landed cleanly 🥇 >5 landed partial 🥈 >2 missed entirely ❌ Firefly understood the request almost every time, it just couldn't hold the rest of the image still while executing it. Avoiding drift is hard. In the partials, the target moved but something else broke. >A product lost prominence / focus >A shadow stayed broken >A new artifact appeared on the wall while the format crept closer to the brief Social was the hardest category. 0 clean wins across 4 attempts. With social content, crop, pose, and product placement ARE the deliverable, so any drift fails the job. TLDR: Firefly can edit like Photoshop, but avoiding drifts is the real challenge. I personally use this feature a lot, how have you solved for this ?

Can Adobe Firefly's "Edit" feature compete with Photoshop? Select a region, describe the fix, keep the rest of the frame intact. That's the pitch. Here’s what we observed across 4 sessions and 8 targeted edits with real working creatives: >1 edit landed cleanly 🥇 >5 landed partial 🥈 >2 missed entirely ❌ Firefly understood the request almost every time, it just couldn't hold the rest of the image still while executing it. Avoiding drift is hard. In the partials, the target moved but something else broke. >A product lost prominence / focus >A shadow stayed broken >A new artifact appeared on the wall while the format crept closer to the brief Social was the hardest category. 0 clean wins across 4 attempts. With social content, crop, pose, and product placement ARE the deliverable, so any drift fails the job. TLDR: Firefly can edit like Photoshop, but avoiding drifts is the real challenge. I personally use this feature a lot, how have you solved for this ?

13,452 Aufrufe

Videos

Keine weiteren Inhalte verfügbar