
J A Z I I
@notjazii • 9,976 subscribers
• Exploring AI & Agents • Testing models & covering AI/tech news • Ambassador @Cognition • Building https://t.co/HQWNYNHAUc
Shorts
Videos

Grok 4.7 vs SWE 2 tested both models with same prompt at highest reasoning available and results came out really different > grok 4.7 took 32 minutes and cost $8.14 to make this > swe 2 took 15 minutes and cost $0 to make this really surprised by how good a free model is and no idea what grok was doing here which one did better?
J A Z I I50,932 görüntüleme • 1 gün önce

opus 5 is getting routed to opus 5.2 tested opus 5 on xhigh with the same prompt in devin and claude code but results came out really different > claude code one is clean and more detailed > while devin one is almost there but not the same opus 5 in cc used to be lazy and hard to communicate with, but something changed in the last 2 days now it works for longer and is way easier to communicate with anthropic did same with fable release too where some users were being routed to fable 5.1 so it’s very likely opus 5 is getting routed to 5.2 in claude code can you guys tests and lemme know if it's getting routed for you or not?
J A Z I I140,113 görüntüleme • 8 gün önce

sonnet 5 is routing to sonnet 5.2 looks like all 3 claude models are now routing for some users and from testing: > opus 5.2 is the biggest jump > sonnet 5.2 looks way better too > fable 5.2 is only a small jump at this point anthropic got whole new claude 5.2 family getting ready to release when do you think they will release them?
J A Z I I41,760 görüntüleme • 2 gün önce

fable 5.2 > gpt 6 astra tested both models with same prompt at high reasoning and results came out really different fable next model entered stealth testing today, so i don’t think it's coming out anytime soon from early testing: > huge step up from current fable > output looks better than astra > slower and expensive still running more tests and will drop the outputs soon
J A Z I I39,450 görüntüleme • 4 gün önce

anthropic cooked hard with opus 5.2 after testing it for a whole day, i can say this model is so good > huge step up from opus 5 > on par with astra > faster than opus 5 > way better at frontend and 3d here’s one frontend test i ran and it’s the best output any model has given me even better than astra you can check fable 5.1 output with same prompt in QT and tbh opus 5.2 just mogs everyone here really hoping that anthropic will release this model officially this week
J A Z I I40,402 görüntüleme • 7 gün önce

meta muse just mogged fable 5.1 tested fable 5.1 and muse spark 1.3 with same prompt at highest reasoning available and results came out really different > muse spark 1.3 completed task in one minute and costed almost nothing > fable 5.1 completed task in 70 minutes and costed $13 it's clear that meta is stepping up their game in ai, they don't wanna left behind which one did better here?
J A Z I I75,566 görüntüleme • 20 gün önce

Opus 5 vs DeepSeek v4.1 Flash tested both models with same frontend prompt at highest reasoning available but results came out really different > opus took 80 minutes to complete and costed $20 > v4.1 flash took 60 minutes and costed $0.2 which one did better here?
J A Z I I39,045 görüntüleme • 12 gün önce

Breaking: Anthropic is preparing a double drop opus 5.1 and fable 5.1 are already live in claude for some users from early tests it looks like models are really improved from their current versions it's highly likely that anthropic will drop both models either today or tomorrow here are some outputs for both models, looks good or nah?
J A Z I I77,278 görüntüleme • 26 gün önce

GPT 6 Astra vs SWE 2.0 tested both models with same prompt at highest reasoning available > astra took 80 minutes to complete this task and costed $64 > swe 2.0 took 25 minutes and costed zero dollars really surprised by swe results here and astra looks nerfed right now, no idea why which one did better here?
J A Z I I31,322 görüntüleme • 12 gün önce

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time overthinking and running random tests, which slowed it down thanks AIHubmix for providing api access so i could test v4.1 flash hoping that team will cut down the overthinking before the full release, but for now here are the results which model did better here?
J A Z I I29,132 görüntüleme • 14 gün önce

GPT-5.6 Sol vs Fable 5 both models were tested on the same reasoning with the same prompt > 5.6 Sol one shotted everything in 70 minutes in devin > Fable 5 one-shotted everything in 90 minutes in claude code really surprised by Sol results here which one do you think won?
J A Z I I79,418 görüntüleme • 2 ay önce

Fable 5 vs GPT 5.5 tested both on High settings with the same prompt > Fable one shotted everything in 15 minutes and used 3% of weekly limits > GPT 5.5 took 20 minutes and used 3% of weekly limits 5.5 is not looking good here at all which one do you think won?
J A Z I I107,777 görüntüleme • 3 ay önce

GLM 5.3 Flash vs Kimi K3 tested both models on same prompt at highest reasoning available but results came out really different > 5.3 flash took 20 minutes to complete this and costed $0.7 via AIhubMix API > kimi k3 took 30 minutes to complete this and costed $4.32 via Kimi Plan which one do you think did better here?
J A Z I I23,105 görüntüleme • 27 gün önce

GLM 5.2 vs Kimi K2.7 both were tested on same settings with same prompt > GLM one shotted everything in 35 mins > Kimi needed extra prompts to fix movements and took 30 mins super surpised by how good GLM 5.2 is, while being so cheap which one do you think won?
J A Z I I63,741 görüntüleme • 3 ay önce

Opus 4.8 vs MiniMax M3 tested both on default settings with the same prompt > Opus one shotted everything in 7 minutes > M3 needed an extra prompt to fix the "break block" feature and took 20+ minutes both got super close, judge both and lemme know which one looks better?
J A Z I I67,862 görüntüleme • 3 ay önce

Opus 4.8 vs GLM 5.2 everyone is praising Glm on my timeline, so i had to test it myself > both models were given the same prompt > GLM was super fast and used like 2% of my weekly limit > Opus took over 20 mins and used like 3% of weekly limits which one do you think won?
J A Z I I48,130 görüntüleme • 3 ay önce