Loading video...
Video Failed to Load
A warning for those using Grok Bot I love the UI and simplicity and the persistent VM use for each bot is great, but I'd be careful trusting it too much until you learn the current limitations of this iteration I would have looked like an idiot had I... show more
98,299 views • 1 month ago •via X (Twitter)
51 Comments

I always tell LLMs to use the latest Grok models with live search enabled for any and all research and fact check tasks.

same. but whatever grok bot is doing under the hood rn appears to have limitations. an unfortunate tradeoff for now - simplicity for the masses instead of model control for those in the game

You can dictate model control at the CLI level - just have an agent that strictly runs whatever model you want through a CLI.

ahhh gotcha. yeah i wonder if that would have resolved this. still have a feeling it may not have given the web fetch tool was the limitation but good to know

It would've because Grok 4.6 with live search bypasses any limitation grok bot has

The main thing missing from @bot right now is I want a model selector and thinking modes. Which I guarantee is how it went wrong for you. And would love for it to have other models too. I think @farzyness said he’s doing this..

@bot @farzyness yeah itll be interesting to see if they offer that in the future or stick to the spartan + clean ui for wider adoption. most people dont want to learn how to use a cli

Interesting. I haven't sold myself on going this far in agentic AI yet.

its fun to tinker with but i still have real concerns about accidental deletions and agents going rogue, especially giving access to google drive and my mac with all my critical work/personal files. but sandboxing to the best of my ability for now

Exactly. I also don't feel I have the use cases that justify doing it. I use agentic AI in some repetitive tasks in building language learning software. but that's within regular grok build.

yeah the things i really want to outsource fully (emails and video editing) are not ready to be outsourced reliably based on any tool ive seen

Video editing would be huge for me but that’s probably at least 2 years away. Not sure that needs Grok Bot though. I use AI for thumbnails to some extent. I still like making my own but ChatGPT is very good at editing and improving them.

Nice demonstration of AI weakness. More refined slop is still slop. Like humans (and human facets), different AIs know different things. There are no perfect AIs, just ones with which you agree.

yeah, thats why im always running answers thru multiple platforms to catch bias and blindspots and missed context etc

Good advise and can verify the same as an avid user of all of these tools, including Grok Bot. and Grok Build dig deep on research. But as Dillon mentions, when it matters, grab multiple sources or just manually verify.

i fear the slop wave that is coming from normies that fully trust these tools too quickly and spew nonsense without getting to a ground source of truth

Having nice tools doesn't make you a expert carpenter anymore than having AI tools makes you suddenly a universal expert! Like winning the lotto magnifies who you've always been financially, AI will do the same with intelligence.

Thanks for the pro tip! Grok bot is V 0.2, so this kind of thing will happen until it matures. Kinda like when the first FSD version kept missing a road jog in my neighborhood and then telling me to take over while leaving the car pointed directly at a tree 😅

@bestjkymn I've also noticed a disconnect when researching with Grok bot vs Grok 4.6 - use 4.6 for research, bot for tasks

grok bot really said trust me bro then cooked the numbers

*For all those using AI in general. The AI will make mistakes and act like it's the truth. You'll correct it; it will say, "Sorry, you were right," and probably make another mistake in your next session. Welcome to another episode of AI will take over your job, not. ;)

Oh no, you continue to criticize Elon babies. Pitch fork people are coming. 🤣😂🤣 Yeah, Grok or any AI can't be trusted, I always double or triple check.

lol some see at criticism i see it as helpful feedback supporting the mission ;)

Did you use 4.6 in build? Or do you have heavy? Grok doesn’t have 4.6 yet on my account, only through build

should have clarified - i pay $200/month for cursor and have been porting all of my work there. gonna share a lot more in time it may have been 4.5 in the web at the time but ive been running 4.6 in cursor

Awesome! Been thinking about doing this, but doesn’t seem like cursor is made for someone like me, so I haven’t bothered. Interested in hearing your take

Let’s not act like Codex or Claude Code doesn’t make moronic decisions.

Qwen-3.8-27B did a better job on my own harness.

I wonder if there is a $PLTR AIP solution for small businesses?? @chadwahl 🤔

Need to lock these scoldings as skills so it’s a repeatable process

haha. feels better to let it fly from my mouth every time

I just built a repeatable “make a 3d model of anything” skill with the bots, no meshy or other 3rd party, I worked hand and hand with the bot, like Optimus will learn until we solidified the process. The bots can call on this skill, oh it’s super cool and can’t wait to share.

@poteto

I'm honestly stunned every time I see people waking up to the reality of how extremely unreliable these things are. Anyone who tells you that their agent is doing this or that, or can be completely autonomously commercially used in a business, is completely out of their mind. All of them are selling absolute bullshit. If I showed you the validation pipelines I'm building for any system that is supposed to run autonomously, you wouldn't believe me... and even then, they still have the potential to mess stuff up. This technology is completely unreliable. It is an inherent LLM problem, not just Grokbot. They aren't even aware of it, and they aren't even lying; it's just hallucinations and confabulations. If you start pairing it with memory, it gets even worse. For example, I've built an automatic pipeline that researches and writes articles for me (not for this account, but for other accounts). If I told you that it goes through five agentic systems and five validating systems... checking, referencing, and counter-checking every scientific quote or reference mentioned... you wouldn't believe me. It takes a lot of work and a lot of effort, but once set up, it can be reliable up to 99% with a low failure rate. But having just some automated bot doing shit and relying on it is absolutely crazy. It's not just Grok; all of them are like that. As Grok always was the retard among LLMs and now seems to slowly be getting closer, it's still disappointing to hear that. Grok is supposed to have better real-time knowledge, and it obviously fails. But I guess it's not the missing real-time knowledge; it's rather really the confabulation issue.

I trust it for real work and it does a better job than me. Yes, always check the output, but it's like questioning FSD. It's a good Bot sir!!

Lol there is clearly a disclaimer in using all AI MODELS that it always not have the correct answer. making a video like this just highlights people not reading disclaimers. lol

Trust but verify

Does the bot have threads? Or does it do one thing at a time?

Hermes and OpenClaw don't have that issue just saying 🤷🏻♂️

You should never trust AI with the final answer. Otherwise humans become obsolete.

This expert level use is beyond my and most people pay grade. Thanks for reminder though.

My internal plan is as follows! ⬇️

All AI must be double checked. Grok makes very basic mistakes daily. There is little “intelligence.”

The problem is that they’re not even using Grok 4.6; they are using Claude’s cheapest model to run these bots 🤖

sw engineers do not trust they verify! well done

@grok summerize please

Love this. I actually have Grok Bot use my Claude and ChatGPT accounts to go back and forth checking accuracy on technical articles. It takes usually 5-8 cycles before it’s written correctly. I used to do this manually. Now I have Grok Bot just do it in the middle of the night and I wake up to a nice piece.

Was Adobe able to recover your broken video the other day?

You cannot ask Grok bot to use 4.6. What you can instead do is ask grok bot to use grok build via CLI and over there it uses 4.6. Also you can ask it to use cursor.

Thanks for sharing the misadventure. I’ve also found a couple of bugs with Grok Bot when it comes to remote sessions being confused for my local computer and asking for permissions it doesnt actually need. But all in all, I’m having serious fun with it

Well said I get errors in Claude also
