Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Generally, free AI comes with weird limits. However, with this stealth AI model, Ox Alpha + OpenCode Zen gateway, I can personally verify that it is: ✅ Almost unlimited running our internal cybersec bench (I burned 4 billion tokens in 4 hours) ✅ Decent speed ✅ Intelligent model (solves...

39,953 görüntüleme • 26 gün önce •via X (Twitter)

38 Yorum

Kartik profil fotoğrafı
Kartik26 gün önce

It's mostly likely either of these: -GLM model -Hy4 -MiMo V3 -MiniMax

Mehul Mohan profil fotoğrafı
Mehul Mohan26 gün önce

I’ll be surprised if this is not an American lab model

Shaun Seah profil fotoğrafı
Shaun Seah26 gün önce

@1kartikkabadi1 how would an american frontier lab be able to run cybersec with no guardrails, look what happened with fable. why would gemini be above them?

Mehul Mohan profil fotoğrafı
Mehul Mohan26 gün önce

@1kartikkabadi1 Grok 4.6 also has no guardrails as such. No guardrails don’t mean that model is SOTA in cyber security tasks (5.6 sol crushes this model in our bench)

Quantum Voyager profil fotoğrafı
Quantum Voyager26 gün önce

It's a chinese model

CryptoKnight23 profil fotoğrafı
CryptoKnight2326 gün önce

Running your internal cybersec bench on a free stealth model is ballsy. Hope that zero retention claim comes with audit rights.

Mehul Mohan profil fotoğrafı
Mehul Mohan26 gün önce

Like our tasks we sell labs, this bench includes offensive multi service web security RL environment tasks. AI can’t see/exfil environment code

CryptoKnight23 profil fotoğrafı
CryptoKnight2326 gün önce

Ok fair. Exfil angle is covered. Still wonder where those 4B tokens end up.

Zeid Kazi profil fotoğrafı
Zeid Kazi26 gün önce

Leaning more towards it being a Chinese model than Gemini 3.5 pro, Testing it out right now

Pluggi profil fotoğrafı
Pluggi26 gün önce

gemini doesn't have max thinking level, probably minimax new model, if not, xiaomi

Martin Ronfort profil fotoğrafı
Martin Ronfort26 gün önce

Four billion tokens in four hours confirms these are real numbers. We actually went deeper on what makes it genuinely different here:

seastart profil fotoğrafı
seastart26 gün önce

4B tokens in 4 hours with zero guardrails sounds less like a standard gateway leak and more like an unthrottled red-teaming sandbox in the wild. The throughput is genuinely insane though.

علي عسيري profil fotoğrafı
علي عسيري25 gün önce

It’s new GLM model.

Benz Affleck™ profil fotoğrafı
Benz Affleck™26 gün önce

Would be funny if it is Gemma 5

Lalit Dhalia profil fotoğrafı
Lalit Dhalia26 gün önce

Oh shit. Lol. Nobody even bothers to think that this could be a google model. And you are right. Who has this much capacity to serve their model? Ofcourse that's google. I think u are the only one I have seen on my timeline mention this being a google model. Let's see next week.

Amandeep profil fotoğrafı
Amandeep26 gün önce

Comparable to deepseek v4 flash in outputs?

Luke ➨ rm2.online profil fotoğrafı
Luke ➨ rm2.online26 gün önce

É chinês amigo!

rohan profil fotoğrafı
rohan26 gün önce

Deepseek V5 Flash for sure

alvinyap.base.eth profil fotoğrafı
alvinyap.base.eth25 gün önce

It's already proven the model has the exact same tokenizer and video encoder footprint as GLM models... why would American labs has the same tokenizer as GLM?

Quantum Voyager profil fotoğrafı
Quantum Voyager26 gün önce

Pretty humble model NGL

Avinash profil fotoğrafı
Avinash25 gün önce

It's GLM

Pulkit Lather profil fotoğrafı
Pulkit Lather26 gün önce

Idk man, they have so much compute 100 trillion is a fucking huge number

Krishna Sharma profil fotoğrafı
Krishna Sharma26 gün önce

Were you always this dumb it's glm 5.3

Haseeb Mir profil fotoğrafı
Haseeb Mir26 gün önce

Gonna Try 0x Alpha for my project from GPT Luna

Haris Sulaiman profil fotoğrafı
Haris Sulaiman26 gün önce

Not Gemini. I think I learned while working with Gemini models is that they have the best support for regional languages like Malayalam that other models don’t. Prompt it with obscure Malayalam words, and you will know the difference.

hit profil fotoğrafı
hit26 gün önce

what's the actual per-request token limit or context window you hit

waishnav profil fotoğrafı
waishnav26 gün önce

It’s astra btw

Srinath Nagarajan profil fotoğrafı
Srinath Nagarajan25 gün önce

My guess -

Fūrin profil fotoğrafı
Fūrin26 gün önce

lmao🤣

Aayush Khare profil fotoğrafı
Aayush Khare26 gün önce

Why would Gemini 3.5 Pro not have Cybersecurity Guardrails?

Pulkit Lather profil fotoğrafı
Pulkit Lather26 gün önce

It's so crazyyy I'm gonna try it today

Jacob Rothfield profil fotoğrafı
Jacob Rothfield26 gün önce

No one else thinks it’s Google

Dylan | LedgerPe profil fotoğrafı
Dylan | LedgerPe25 gün önce

Same here burned through 1B tokens and worked as well/better than sol for most the tasks

Shishir Srivastav profil fotoğrafı
Shishir Srivastav25 gün önce

I mean google did launch nano-banana like this, although the difference is, everybody was impressed with it, unlike this one.

LofiGroove profil fotoğrafı
LofiGroove24 gün önce

Can I use in claude?

RG profil fotoğrafı
RG26 gün önce

But we can see the thinking does gemini models shows thinking texts?

Aditya profil fotoğrafı
Aditya26 gün önce

Interesting

AgenticAussie profil fotoğrafı
AgenticAussie26 gün önce

Its GLM 5.3 Vision

Benzer Videolar

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 görüntüleme • 2 ay önce