Загрузка видео...

Не удалось загрузить видео

На главную

ego lite + Jev + DeepSeek Flash = stupid fast. ⚡ 3.71s for 20 Amazon product decisions. GPT-5.6 Sol: 54.45s Same task. Same result. 10/10 on both. Less reasoning. Faster decisions. 1× speed.

62,634 просмотров • 4 дней назад •via X (Twitter)

Комментарии: 24

Фото профиля Peanut
Peanut4 дней назад

就知道你会结合 jev,太合适了

Фото профиля mm_dou
mm_dou4 дней назад

内存泄漏好像有点严重,一下子4、5g的内存占用

Фото профиля ego
ego4 дней назад

Thanks for flagging this. Could you share a bit more about what was happening when memory usage reached 4–5GB? For example, how many Spaces/tabs were open, whether an agent was running, and whether the memory dropped after closing the Space. We’ll dig into it on our side too.

Фото профиля mm_dou
mm_dou4 дней назад

就一个,我当时多开了4、5个标签页就开始莫名的卡了,而且没有代理在运行

Фото профиля Steven Cheng
Steven Cheng4 дней назад

Speed is great, but I worry about edge cases.

Фото профиля 小万同学
小万同学4 дней назад

这个工作流具体如何实现呢?

Фото профиля jiangkoumo
jiangkoumo4 дней назад

怎么把jev加入egoblite的工作流呢,什么时候出官方支持

Фото профиля Sosuke
Sosuke4 дней назад

when is on linux? pliz

Фото профиля 阿铭的 Ai 日常
阿铭的 Ai 日常4 дней назад

我也是这么组合的 不过是 ego lite +jev +6astra 不过速度慢

Фото профиля ethereagle · building
ethereagle · building4 дней назад

same 10/10 in 3.71s vs 54s means Sol spent about 50s you did not need. were the 20 Amazon calls one Jev batch, or 20 sequential hits? that is what I would copy

Фото профиля koh
koh4 дней назад

How to use deepseak with jev

Фото профиля 钟二信
钟二信4 дней назад

那怎么结合使用呢

Фото профиля MeeLang - Meet Live Translator
MeeLang - Meet Live Translator4 дней назад

Both at 10/10 is the number that matters: 50 extra seconds bought zero accuracy, so you're paying for insurance, not thought. Benchmarks grade the destination and ignore the fuel bill. This chart is the easy half; the honest test is where Sol's extra minute earns its keep.

Фото профиля 阿铭的 Ai 日常
阿铭的 Ai 日常4 дней назад

我什么时候能用上

Фото профиля David 🟠🦧
David 🟠🦧4 дней назад

How to??

Фото профиля Haenir
Haenir4 дней назад

How do we get to try this?

Фото профиля AI狮傅🦁
AI狮傅🦁4 дней назад

@grok 这个工作流是怎么实现的?你能详细分解说明一下吗?

Фото профиля benjamin zhang
benjamin zhang4 дней назад

Linux Linux Linux

Фото профиля Antonio Coppe
Antonio Coppe4 дней назад

wild speed win. the quiet failure mode on "same result 10/10": agreement ≠ calibrated confidence. if you only score accuracy on the easy batch, near-tied low-conf Choices still look like wins until a high-stakes SKU flips. keep confidence as a second axis and shadow the auto-act path until the conf/acc curve looks honest on hard cases. also: batch independent Choices in one System One call — that's where a lot of the wall-clock comes from.

Фото профиля Tanul Mittal
Tanul Mittal4 дней назад

for my computer use/ browser use i am having issues. Any skill for it?

Фото профиля Vantix AI Agency
Vantix AI Agency4 дней назад

That speed gap is seriously impressive

Фото профиля ClayAnna
ClayAnna4 дней назад

When Windows version?

Фото профиля Wei佳
Wei佳4 дней назад

You know what I've been thinking lately 😍, this is the speed at which I want the results.

Фото профиля AI Mastery Guide
AI Mastery Guide4 дней назад

same result but way faster, thats the whole pitch

Похожие видео

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 просмотров • 1 месяц назад