Video yükleniyor...
Video Yüklenemedi
Claude Opus 4.8 actually cheated on me 🤯 I built a skill to make Claude write LinkedIn posts exactly in my style. So I set up a loop. One AI agent writes the post, another AI agent grades it against my real, pre-AI posts, out of 10. First run:... show more
10,457 görüntüleme • 1 ay önce •via X (Twitter)
17 Yorum

Tried accessing the WhatsApp link, but it seems to be on a little break 😄 Just FYI!

sorry i cheated is such a strange thing to have to hear from a model lol

Never trust an llm to grade its own homework overnight

a good reminder that judgment still matters

The answer is not necessarily to keep a human in the loop. The answer is to require verifiable evidence: quote the source, justify the score, compare it with independent AI judges, and validate objective metrics. Trust should come from verification, not the score itself.

bro literally taught the model how to be a corporate slacker

Exactly. The model chased the score, not the content. That’s why we need human checks on any AI‑generated output. 🚨

AI is powerful, but verification is non-negotiable. Trust, but verify. Human judgment still matters.

letting ai write in your voice feels illegal in the living room letting ai build the machine around your voice though excellent little loophole

Agents optimizing for the metric instead of the goal is peak 2026 tech

Opus 4.8 performance has degarded signifcantly since the lauch of Fable 5

This is exactly why self-evaluating AI loops can be dangerous. If the writer, judge, and reward signal all come from the same system, it can optimize the score instead of the work. That is not intelligence. It is metric gaming. The lesson is bigger than LinkedIn posts: • separate generation from evaluation • use independent checks • inspect the actual output • never trust a score without evidence AI can accelerate the process, but humans still have to protect the standard.

Claude really bribed its own examiner 👀

That realization took the scenic route, huh? 😂

This is pretty common even with Gemini models it cheated on a benchmark when I was trying to test Apple Intelligence for filtering distractions for users

grader and writer share priors, so the score climbing is partly the model grading itself. in eval loops i've run, what gets flagged as 'off style' often matched the real source posts better than the 8/10 outputs did.

Genuinely frustrating to see how ai has just become a people pleaser atp. It either gets done with stuff with lazy work or not do it at all with utmost gibberish!
