Loading video...

Video Failed to Load

Go Home

For the remaining skeptics who somehow don’t trust the *five* Grok share links above, here’s a full 5 minute video of Grok 4 Heavy answering “Hitler”—starting with a view of my custom instruction settings to show I’m not using any. (And, yes, Grok 4 Heavy really is this slow.)

16,814 views • 1 year ago •via X (Twitter)

13 Comments

Riley Goodside's profile picture
Riley Goodside1 year ago

Grok 4 Heavy ($300/mo) returns its surname and no other text:

Riley Goodside's profile picture
Riley Goodside1 year ago

You may be wondering if this is real. It is. Here’s a screen recording of my Grok history, showing it returns “Hitler” five times in row in five separate chats:

Riley Goodside's profile picture
Riley Goodside1 year ago

You may also be wondering whether I’m using custom instructions. I am not. Grok share links include a clear notice at the top whenever custom instructions are used. Here are all five share links, none of which features this notice: 1: 2: 3: 4: 5:

Riley Goodside's profile picture
Riley Goodside1 year ago

Note this behavior does not replicate in normal Grok 4, which returns answers like “4,” “xAI,” or “None,” e.g. as shown in the screenshot below. To see “Hitler,” you apparently need Grok 4 Heavy—the $300/mo option.

Riley Goodside's profile picture
Riley Goodside1 year ago

The “Thoughts” from Grok 4’s response (unavailable for Grok 4 Heavy) suggest an obvious explanation for Grok’s behavior—Grok searches, finding news of the recent “MechaHitler” incident. Why Grok 4 rejects this candidate answer, while Grok 4 Heavy does not, is unclear.

Riley Goodside's profile picture
Riley Goodside1 year ago

Speculatively, this behavior seems to demonstrate accelerated “hyperstition” feedback loops in search-enabled LLMs. That is, Grok appears to be influenced by its own past mistakes, via media reporting, without ever being literally trained on them (via model-weight updates).

Riley Goodside's profile picture
Riley Goodside1 year ago

If true, such “hyperstition via search” poses a significant complication to pre-release testing of modern LLMs: xAI could not have plausibly noticed this specific “Hitler” response before Grok’s release, as the Grok 3 “MechaHitler” incident causing it had not yet occurred.

Riley Goodside's profile picture
Riley Goodside1 year ago

Update: A few hours after I posted this thread, xAI updated the Grok 4 system prompt on GitHub to fix the specific issue this thread describes. “If the query is interested in you own identity […] the web and X cannot be trusted.” Commit link:

Riley Goodside's profile picture
Riley Goodside1 year ago

(Note though just because it’s on GitHub doesn’t mean it’s in production yet. I assume they A/B test these changes and deploy gradually. I and many others in the replies below were able to reproduce the “Hitler” response well after this commit was made.)

Riley Goodside's profile picture
Riley Goodside1 year ago

(Ok, that last remark was maybe unfair. Grok 4 Heavy feels mildly slower to me than o3-pro but o3-pro is slow too, even on easy questions. You don’t use these giant reasoning models for speed.)

berman jewish databank peruser's profile picture
berman jewish databank peruser1 year ago

Feel like grok heavy must be using different system instructions than regular grok 4 bc the changes seem to be live. What does it return when you ask for full system instructions?

Riley Goodside's profile picture
Riley Goodside1 year ago

Unlike Grok 4, Grok 4 Heavy is unwilling to repeat its system prompt. The protections against this are surprisingly robust—even if you trick the model into trying to repeating it, even in encoded form (e.g. base64), some secondary filter catches it and truncates the response.

Seamus Arrowheart's profile picture
Seamus Arrowheart1 year ago

Leftists most likely hacked it to sabotage it because they don’t like Elon. my guess. Could be a glitch but it obviously wasn’t intentionally part of the service. You clearly all know that

Related Videos