Video wird geladen...
Video konnte nicht geladen werden
i asked opus 5.5 to explain why personal benchmarks are so important (this is one shot)
17,669 Aufrufe • vor 4 Tagen •via X (Twitter)
9 Kommentare

incredible

So good! 🤯

This is amazing. How long did it work on it? Effort max? I’m trying something similar. It’s been working for 2 hours on the video. But that might be normal timing for something so high fidelity…

A benchmark you wrote yourself is the only one the model has not already seen.

i've been experimenting this year with how to shape models' explanations so I comprehend them better and faster, despite a habit of multi-tasking and always being under deadline. slowing it down into a spoken video explainer feels useful (maybe because I watch it full screen for focus?)

Old World: Idiocracy New World: Idiographicy

When the Yellow soil of China appeared, I chuckled, it is about open source models anyway, cool.

I've been sleeping on this for too long. Think I'm going to figure out something this week for a benchmark. I think some sort of slide deck benchmark. I do AI enablement and make a lot of decks.

Running Opus 5.5 headless at low effort for research briefs: about 15 seconds and twelve cents a run. Low effort is the real feature. You only escalate when the call is genuinely ambiguous.
