Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Pi Agent vs OpenCode token usage A lot of people recommended Pi Agent so I decided to check Pi Agent took 1.1k tokens in first turn OpenCode took 11.5k Setup: 1) Trimmed OpenCode (from usual 30k first turn to 11.5k) - 0 MCPs - 2 lightweight plugins (opencode-env-protect and...

209,100 görüntüleme • 4 ay önce •via X (Twitter)

35 Yorum

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I'm also disabling all my MCPs and rebuilding my workflow with CLI-based skills Token economics are changing fast, with subscriptions becoming more token constrained 50-100k tokens per turn from MCPs alone is a recipe to hitting usage limits fast As an MCP maximalist I'm going to miss how predictable MCPs are. Models are going to get one CLI argument wrong most of the time

ProxySoul profil fotoğrafı
ProxySoul4 ay önce

@pseudokid now try SoulForge

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

wow thanks for sharing!

Agentic Up profil fotoğrafı
Agentic Up4 ay önce

pi has less number of tools and extensions available, so its going to be more token efficient

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I agree

Artificial Guy / João Vitor A. profil fotoğrafı
Artificial Guy / João Vitor A.4 ay önce

i think need to compare and understand where this tokens going and if its help or not the daily coding... i dont think its useful to just compare raw vs modded... maybe pi with 5k prompts can be even better than pi 1k or opencode 11k (but yeah, 11k is a lot )

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

Totally. I agree. I'm about to figure out if less tokens with Pi and a more capable but slightly expensive model is enough to do the things that I can do with the cheapest in OpenCode Go which is deepseek v2 flash

Prashanth (Manohar) Velidandi profil fotoğrafı
Prashanth (Manohar) Velidandi4 ay önce

Now try the same model from different providers. Like @InferXai and someone else.

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

@InferXai $20/mo for @InferXai right?

Prashanth (Manohar) Velidandi profil fotoğrafı
Prashanth (Manohar) Velidandi4 ay önce

@InferXai Yes. While we are in beta and limited spots very limited available.

Ahmad Awais profil fotoğrafı
Ahmad Awais4 ay önce

have you tried @CommandCodeAI $1 Go plan? also i think a great system prompt helps you!

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

@CommandCodeAI saw their ad a couple of days ago and i might actually try is system.md better than agents.md?

Albert Gao profil fotoğrafı
Albert Gao4 ay önce

I have the same experience for local models. Was trying Gemma 4, qwen 3.6 locally on my M5 Max 128GB with Opencode+Ollama, found it is pretty slow, until I accidentally tried pi with the same repo…damn, it is night and day…is it because it has a gigantic system message? No MCP, no skills, just factory default for both. Also, Codex CLI is awesome to use👍

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I'm seeing a lot of people using Qwen 3.6 locally. I also find Pi being too light and fast. Users start from nothing so you they control what they put in it. Saving tokens is just a side effect of that imo. Which one you find better, Gemma 4 or Qwen 3.6?

FrenchToasters profil fotoğrafı
FrenchToasters4 ay önce

Just go ahead and run some of those longer tasks with context in the 140k range. You’ll really see who is cheaper then

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

About to do that, that's the real test

A. F. ♪ profil fotoğrafı
A. F. ♪4 ay önce

Have you compared it with @CommandCodeAI yet? I haven't tried it yet (hopefully before the end of the month) but I have been seeing a lot of wonders about them and their process. Like

Kris 👾 profil fotoğrafı
Kris 👾4 ay önce

I agree with you! Don’t know much about stats. But I used an agent workflow in opencode and then just gave a try in pi, there’s slight usage difference that I noticed, then optimized in pi itself now for each workflow run it hardly cost 10 cents to me with MiniMax M2.7

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I feel like MiniMax M2.7 and even M2.5 will run better on Pi just because of Pi's minimalist approach I have problems with these models for long-running tasks in OpenCode. Easy to lose key and recent details that even older gens Mimo V2 and Kimi K2.5 will never lose My theory is it easily gets distracted so it has to be given one single coherent point of focus Can't have system prompts competing with system wide AGENTS.md and task brief

ffswunnd profil fotoğrafı
ffswunnd4 ay önce

Because opencode it's a coding harness with longer instructions.

Shreyash profil fotoğrafı
Shreyash4 ay önce

@grok is Pi Agent better than Opencode in terms of perf and cost ?? What's opinion of devs on X ??

Rambone profil fotoğrafı
Rambone4 ay önce

How does no system prompt work? It jut gets tool schemas and runs with it?

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I'm still new to Pi, but I simply didn't use a SYSTEM.md for it I also noticed that Pi relies on the terminal, just like Codex Desktop. OpenCode has its own Read and Write tool and bash.

Atharva Verma profil fotoğrafı
Atharva Verma4 ay önce

Try with @justsisyphus too

Alex profil fotoğrafı
Alex4 ay önce

So you put two different models into two different harness and concluded?

Andrew R profil fotoğrafı
Andrew R4 ay önce

Did you save the first turn and look at it? Bc I am skeptical the pi is including your Agents file

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I thought so too. So I tried probing it and it understood the AGENTS.md I have to test more and learn about the internals of Pi agent I only have the usage record from my OpenCode Go dashboard. Tokens used and price and that's it.

Nazmus Sayad profil fotoğrafı
Nazmus Sayad4 ay önce

in output which seems better?

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

still testing it!

Nati profil fotoğrafı
Nati4 ay önce

This is kinda misleading in longer tasks.

raymel 👋 profil fotoğrafı
raymel 👋4 ay önce

I'm about to figure it out, the real answer lies there

武蔵 profil fotoğrafı
武蔵4 ay önce

@thdxr

CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO profil fotoğrafı
CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO4 ay önce

@Teknium

fecolinhares profil fotoğrafı
fecolinhares4 ay önce

there's support for rules.md, commands and agents on pi besides just the skills and agents.md ?

Storm profil fotoğrafı
Storm2 ay önce

Same pattern I've seen, most of the burn isn't the model, it's re-reading everything every turn. Building ThinkForge, search-first indexing (pull only relevant chunks, cache what's read) cut first-turn cost more than any model swap did.

Benzer Videolar

I respectfully disagree for several reasons. Calling a customer, whether free or paying, an idiot is simply wrong. OpenCode, like any other coding agent, clearly tries to preserve the prompt cache as much as possible. Otherwise, it would be painfully slow. The “Stop Using OpenCode” article, which I believe Dax is referring to, focused on specific cases that can invalidate the cached prefix. These issues become much more visible and painful when using local models. Modifying AGENTS.md is one example, as shown in this video. During this coding session, cache efficiency is extremely high. But the moment I modify AGENTS.md, boom: 62K tokens need to be processed again as new input. The date change is another valid point raised in the article. I understand that both issues may sound trivial, and personally I've never been impacted by them, but their impact on the OpenCode end-user experience can be significant when using local models. I agree that both the feedback and the article were too direct and were probably written by the author out of frustration. However, they also contained constructive points that could help improve the product. I use OpenCode alongside several other tools, and I like it overall. But I really dislike this kind of reaction to user feedback, regardless of how that feedback is expressed. I would have preferred a response focused more on listening, understanding the problem, and improving the product.

Ivan Fioravanti ᯅ

37,594 görüntüleme • 2 ay önce

HERMES AGENT WITHOUT THESE 3 FILES IS A CHATBOT. WITH THEM IT KNOWS WHO IT IS, WHO YOU ARE, AND WHAT IT LEARNED. SOUL.md — who the agent is. first thing in the system prompt. defines personality, voice, values, how it operates, what it can and can't do. structure yours like this: → identity (name, role, relationship to you) → values (what matters, what principles guide decisions) → voice (how it communicates, tone, style) → operations (autonomy level, ground rules) → restrictions (what it never does) → failure protocol (how to operate when things break) lives in ~/.hermes/SOUL.md. auto-seeded on first install. edit anytime. scanned for prompt injection on every load. keep it concise. SOUL.md injects into every turn. a 200-line soul burns tokens on every message. aim for 50-80 lines. one paragraph per section. MEMORY.md — what the agent remembers. persistent facts, insights, preferences. survives across sessions and restarts. capped at ~800 tokens by default: memory: memory_char_limit: 2200 the agent writes to this automatically as it learns about your work. USER.md — who you are. your profile, preferences, context. capped at ~500 tokens by default: memory: user_char_limit: 1375 injected every turn so the agent always knows who it's working for. bonus: AGENTS.md — project-specific instructions. drop one in any project folder. subdirectory AGENTS.md files load lazily during tool calls, not at startup. keeps your system prompt light. prompt stack order on every turn: SOUL.md → tool guidance → MEMORY.md + USER.md → skills index → AGENTS.md → platform hints skills come preloaded. 60+ built-in tools. the agent creates more skills from completed work. you focus on these 3 files. each profile gets its own copy: ~/.hermes/SOUL.md (default profile) ~/.hermes/profiles/researcher/SOUL.md ~/.hermes/profiles/ops/SOUL.md different agent, different soul, different memory, same machine. full breakdown of Hermes as a Personal AI OS in the article 👇

YanXbt

28,924 görüntüleme • 3 ay önce

hy3 vs mimo-v2.5 vs deepseek v4 flash vs minimax m3 the four models on top of the openrouter leaderboard by tokens this week: #1 hy3 (Tencent Hy) – 7.5t #2 mimo-v2.5 (Xiaomi MiMo) – 6.56t #3 deepseek v4 flash (DeepSeek) – 5.24t #4 minimax m3 (MiniMax (official)) – 4.21t so we tested them. 3 prompts, single-file html, Three.js from a cdn, fully procedural, no external assets. all run via AI/ML API each prompt is a transparent cutaway machine that has to be mechanically correct, not decorative: • 4-stroke engine with full oil circulation – slider-crank kinematics, cam at 2:1, valve lift driven by lobes, oil loop from sump to gallery to big-end • watt walking-beam steam engine – four-bar vector-loop closure, eccentric-driven slide valve, steam events synced to real port position • francis reaction water turbine – 20 guide vanes on a regulating ring, 17 lofted runner blades, gpu particle advection, precessing vortex rope at part load the takeaway up front: none of the four cleared all three scenes on the first attempt. but the price spread between them is roughly 70x – hy3 fixed included costs less than two cents overall results (summed across all 3 scenes): cost #1 hy3 – $0.016 #2 deepseek v4 flash – $0.025 #3 mimo-v2.5 – $0.97 #4 minimax m3 – $1.17 tokens #1 hy3 – 19,326 #2 deepseek v4 flash – 63,126 #3 mimo-v2.5 – 322,523 #4 minimax m3 – 702,900 lines of code #1 hy3 – 1,047 #2 mimo-v2.5 – 2,759 #3 deepseek v4 flash – 3,273 #4 minimax m3 – 3,354 scenes needing a second attempt #1 hy3 – 1 (engine) #1 mimo-v2.5 – 1 (turbine) #1 minimax m3 – 1 (turbine) #4 deepseek v4 flash – 2 (steam engine, turbine) observations: 1. the token spread is the real story – minimax burns 36x hy3's tokens and lands in the same place, one retry, ~3.3k lines 2. hy3 is the outlier on density: 1,047 lines total, fewest tokens, cheapest run, and only one scene needed a second pass. deepseek is the opposite trade – near-hy3 pricing but the most retries 3. mimo and minimax seem to overthink instead of writing the code. minimax spent 359.1k tokens on the steam engine and produced 1,346 lines – the tokens are going somewhere other than the file 4. the francis turbine broke three of the four. the spec that separates them is the one with 20 linked guide vanes and gpu particle advection, not the one with the most parts overall impression: none of these models excelled at any of the tasks we gave them. but they were close, and they were extremely cheap. the gap that matters isn't quality anymore – it's that hy3 ran all three scenes for less than two cents while the frontier labs charge dollars for the same work right now you pick these because they're good for the zero price you pay. soon that's something openai and anthropic will have to think about follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

17,145 görüntüleme • 2 ay önce

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time I feel like there is an open-weight model that can reason at the level of Claude and Codex. And it does this in a cost-effective way with support for 1M context length. To be clear, I am using DeepSeek-V4-Pro inside of Pi without any special configuration. It works out of the box. It's exciting that there is a model that can just be plugged into a basic harness like Pi, and it just works. I've never seen that before. Most models require lots of configuration and setup. DeepSeek's DeepSeek-V4-Pro is clearly good at agentic coding (probably the best from the open-weight models), but the model is also great on knowledge-intensive tasks where reasoning matters. The agent pulled agentic engineering best practices from different company docs (Anthropic, OpenAI, Google, Stripe, Meta, Modal, DeepSeek, Mistral, Cohere), searched and digested Reddit and HN threads, summarized arxiv papers, and surfaced trending GitHub repos. Then it distilled everything into actionable tips across categories. I love the Wiki it built. The quality is really good. Here is a snapshot of what the wiki looks like: DeepSeek-V4-Pro handled the task without breaking stride. Multi-step research queries, code generation for scaffolding, context-heavy reasoning across disparate sources. For coding specifically, this is the first open-weight model that genuinely feels like a Codex or Claude Code experience. It compares in capability and actual multi-turn agentic work. What made the loop feel so responsive was Fireworks' inference speed (the fastest in the market) and the fact that they actually validate models at the systems level before shipping. No corrupted reasoning traces. Just fast, reliable iteration. The hybrid CSA and HCA attention design cuts KV cache to just 10% and inference FLOPs by nearly 4x at 1M-token context. This is what makes the agent loop actually fast and cheap enough to run in practice. For devs who've been watching open-weight models close the gap but haven't found one that actually delivers in practice, this is the closest I've seen. Try it here:

elvis

60,281 görüntüleme • 5 ay önce

gemini 3.7 flash vs deepseek v4 pro 0813 vs muse spark 1.2 – on voxel city dioramas three models each built three crossy road-style 3d scenes – a construction site, a nyc intersection, a river with a drawbridge – as single self-contained html files the setup: Nous Research's hermes agent cli on OpenRouter, three.js skills preloaded, identical prompts per scene tasks: 1. construction site – tower crane on a working lift loop, paver laying fresh road, roller compacting it behind 2. nyc crossing – four-way intersection with a traffic light state machine, queuing cars, pedestrians crossing on the walk signal 3. river drawbridge – double-leaf bascule that lifts for tall boats, cars queuing at the barriers, animated water every scene: Three.js r185, box geometry only, a locked 20-color palette, four camera presets, and a day/night mode with bloom. one file, no build step, no assets models: Google DeepMind gemini 3.7 flash, DeepSeek v4 pro 0813, AI at Meta muse spark 1.2 muse and gemini finished every scene in two to three minutes. deepseek took 15 to 41 minutes per scene - build time, all three scenes #1 gemini 3.7 flash – 6m 43s #2 muse spark 1.2 – 7m 20s #3 deepseek v4 pro – 91m 25s - total tokens #1 muse spark 1.2 – 440,279 #2 gemini 3.7 flash – 713,855 #3 deepseek v4 pro – 20,957,568 - total price #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57 - agent calls across the three builds #1 muse spark 1.2 – 12 #2 gemini 3.7 flash – 18 #3 deepseek v4 pro – 143 observations: • muse won two of the three scenes on looks with the smallest files in the test – 887 to 1,042 lines against gemini's 1,934 to 2,377. cheapest, fastest to a good frame, and shortest turned out to be the same column • deepseek burned 20.96m tokens – 29x gemini, 48x muse – across 143 agent calls. prompt caching is the only reason that cost $4.57: the cache discount absorbed roughly $30 of resent context • gemini was the only model whose files needed zero fixes to render – and the only one whose night mode is cosmetic. the sky never darkens and one camera button does nothing. clean code for a scene it never looked at follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

29,033 görüntüleme • 1 ay önce

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 görüntüleme • 1 ay önce

I cut Fable 5 token usage 2.5x with just one change! - Before: 5.5 M tokens · 7 errors · $8.94 - After: 2.3 M tokens · 0 errors · $4.17 The final build was the same for both, but the path the agent took wildly differed. In both runs, the agent started with the same thing, i.e., it understood the backend before building anything, like: - Permission policies - Available storage buckets - Auth providers configured - How edge functions are deployed The first run used Firebase, which was built for a human dev using a dashboard. While the dev can read the above state by clicking through tabs, an agent has no dashboard. So it gathered the same info through API calls. And there's no single Firebase call that returned this info. The agent required to query multiple times, and each query over-returned. For instance, when the agent asked how sign-in is configured, Firebase also returned the entire auth surface and every method it supported. This was far more context than what it needed. And it repeated across every part of the backend it inspected. Some states (like which auth providers are active) weren't queryable at all. I provided it myself. Otherwise, the agent would have guessed. Errors further compounded the token usage. When a dev sees "permission denied," they can look at the console and figure out whether it's a rule, a path, or an unauthenticated request. Firebase returned the same string to the agent as well, and it had none of that surrounding context to debug. So it guessed again, picked the most likely cause, and rewrote code, utilizing more tokens. This Firebase setup cost me 5.5M tokens and 7 manual interventions during errors on a full-stack RAG app. But I brought that down to 2.3M tokens and 0 manual interventions by using InsForge as the backend context engineering layer (open-source and self-hostable via Docker). It provides the same primitives as Supabase/Firebase, but structures the entire information layer for agents, instead of dashboards. In one CLI call that consumed ~500 tokens, the agent saw the full backend topology before writing a single line of code. This included auth, database, storage, edge functions, model gateway, micro VMs, and deployment. Also, instead of loading the entire product surface into context on every task, four narrowly scoped skills activated only when relevant to keep cognitive load minimal. And to ensure efficient retries if needed, every CLI operation returned structured JSON with meaningful exit codes, so the agent never guessed what to do next. Here's the InsForge GitHub Repo: (don't forget to star it ⭐) The video below depicts the final build, comparing Firebase and InsForge. To dive deeper, I recently published a full walkthrough building the same RAG app on both backends and inspected them end-to-end. Read it below.

Avi Chawla

113,307 görüntüleme • 3 ay önce

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 görüntüleme • 4 ay önce

I designed a new test specifically for multimodal models: fill out a paper form. And it's much harder than it sounds. This isn't typing into an electronic field that captures your text. The form is just an image. The model has to place each form element: text, checkmarks — at the correct pixel position on the canvas itself. Results: 🟢 Kimi K2.6 → done in 3:45, 16.7k output tokens 🟡 Step 3.7 Flash → half the fields, 57k output tokens 🔴 Gemini 3.5 Flash → 489k output tokens, never finished. I had to kill it. Gemini burned ~29x more output tokens than Kimi on the exact same task, and Kimi's was the only form that actually looked filled out. The test, a mocked application form, contains some challenging parts, such as one-character-per-box fields. I provided every model the same set of tools: > get canvas size > drop probe markers to find coordinates > add text > add checkmarks > move elements > take a screenshot anytime to check their own work > ... etc So it's vision + spatial reasoning + tool use + long context, all at once. Small models (Qwen, Gemma) can't really complete this test, so I skipped them. What happened: > Kimi nailed name, DOB, ID, gender, marital status, nationality, email, phone, address, postal code — placement slightly loose, but content correct. 15 turns. Clean. > Step got maybe half right — fields dropped, "United States" landed in the email line, data floating outside boxes. Burned 1.24M input tokens doing it (81 turns of re-reading the canvas). > Gemini almost got there visually... then spiraled. By turn 40 it was issuing a delete_elements call wiping element IDs 365–425, basically erasing its own work. 31 minutes, 489k output tokens, still streaming. Terminated. The takeaway isn't "Gemini bad." This test is indeed difficult. But token efficiency is capability now. A model that needs 30x the tokens and still can't converge is going to be 30x the cost in production. Kimi K2.6 just quietly did the thing.

stevibe

25,494 görüntüleme • 4 ay önce

Dr. Fan video translation: First of all, I would like to congratulate more than 16 million Pioneers around the world for transitioning to the open network. This is the result of our joint efforts over the past six years. We should celebrate this historic moment in the development of the network. Looking back at the development of Pi now, it has always been a unique project. Some people will feel that it is different, and yes, it is. Pi is a non-conformist. Many innovations come from non-conformists who challenge conventions, conventions, and established practices. In the case of Pi, in the early days of cryptocurrency development, many projects raised funds through initial coin offerings (ICOs), but Pi never did so. Pi never sold tokens through ICOs and ensured that everyone could get them for free. 80% of Pi tokens belong to the public and the community. When a cryptocurrency project can go online with just a white paper, a smart contract, or even just an emoji today, the Pi community spent six years building its infrastructure, ecosystem, and ensuring that it has usable practical functions. Before opening, most of its addresses were unverified for most blockchains. Pi chooses to verify the identities of millions of users through identity verification (KYC) and business identity verification (KYB) to ensure that the identities of individuals and businesses on the mainnet are authentic and reliable. Because Pi firmly believes that true decentralization is not inconsistent with authenticity and legitimacy. Some cryptocurrency critics worry that Pi's users are too mainstream. Is this a problem? We think not. This is precisely the biggest advantage that the Pi network has been working hard to build and build over the past six years. Ordinary people like Pioneers are the driving force behind Pi's development because Pi aims to solve the problems of large-scale applications and real-world practicality. Pi welcomes both cryptocurrency enthusiasts and mainstream people. History has shown that only by meeting the real needs of the mainstream population can technology develop for a long time. With this development trend, the cryptocurrency industry as a whole will also benefit from the participation of a large number of mainstream audiences to jointly create the real utility of blockchain technology. What does it mean to open the network today? This means that the firewall of the Pi blockchain has been removed and connected to the outside world. This makes it easier for merchants to sell goods and services in the local market; developers can further develop applications and improve business model logic; creators can gain greater influence; and ordinary Pioneers can also connect and trade with each other. For new developers, welcome to this widely distributed crypto network. If you already have a business model, come to the Pi network to test and sell your products. For developers who don't have a business model but have an excellent application experience, the Pi application network has prepared a business model for you. The platform will bring you traffic and make you profitable. At the same time, it creates the use value of Pi for all pioneers, which is very beneficial to the Pi network as a whole. So, Pioneers, in the end, don't let external noise distract you. Focus on key matters, focus on things that can have an impact in the crypto field and even the world. Keep working hard, keep creating, and everything else will fall into place. MY MESSAGE Please listen carefully to this video so that you understand, the value is determined by the existing ecosystem, the current ecosystem is GCV $314,159 What small traders do under the auspices of the ambassadors 🥰 Eagle woman 🦅 Nonny Padja NTT 🇮🇩 Believe it or not, it's up to you, the point is we've won with a GCV barter value of $314,159 If you want to succeed Barter with Pioneer and small merchants with GCV value of $314,159❤️

NONNY PADJA NTT ❤ Eagle woman 🦅

45,669 görüntüleme • 1 yıl önce

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 görüntüleme • 1 ay önce

I dug into Pi because the greatest projects are usually built on top of one very simple building block. 𝗣𝗶 𝗵𝗮𝘀 𝟳𝟳,𝟬𝟬𝟬 𝗚𝗶𝘁𝗛𝘂𝗯 𝘀𝘁𝗮𝗿𝘀 𝗮𝗻𝗱 𝗲𝘅𝗮𝗰𝘁𝗹𝘆 𝗳𝗼𝘂𝗿 𝘁𝗼𝗼𝗹𝘀: 𝗿𝗲𝗮𝗱, 𝘄𝗿𝗶𝘁𝗲, 𝗲𝗱𝗶𝘁, 𝗯𝗮𝘀𝗵. 𝗦𝘂𝗽𝗲𝗿 𝘀𝗶𝗺𝗽𝗹𝗲. No MCP, no sub agents, no plan mode, no to-dos, no permission popups. Their site has a whole section called "what we did not build." So here's the whole logic by Mario Zechner: prompt in → agents.md + system prompt become one forkable JSON → loop starts → read / write / edit / bash → session saved as a tree → loop ends agent-loop.ts has only 792 lines. I think it's beautiful. Initially I was wondering if it's gonna be a little bit too basic. And yes out of the box it's weaker than Claude Code. But skipping MCP isn't purity, it's a context budget. One MCP can sit tens of thousands of tokens in your window every single turn, for a tool you use maybe 10% of the time. Pi moves that weight from always on to on demand. A skill is just markdown. An extension is one .ts file. A package ships both. I ran it in the terminal, then again as a coding sub agent inside my own harness Waku-Agent ( through a delegate_task tool. Same binary both times. You're paying setup time for a harness that stays yours. Full 22 min breakdown in the comments. Save the diagram 🔖 You Can Build Anything. You Can Learn Anything. 💪

Shen Sean Chen

59,186 görüntüleme • 2 ay önce

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 görüntüleme • 6 ay önce

Qwen3.8-Flash-Next is still going strong at 364.7K tokens of context on an M5 Max. And this isn’t just a static long-context test. The model was reasoning about how to speed up its own workflow while using tools, and the tool calls kept working without misses. Setup: • Qwen3.8-Flash-Next • M5 Max • 128GB unified memory • MLX-Serve PR #363 • OpenCode 2 • 364.7K context The interesting part isn’t simply getting hundreds of thousands of tokens into memory. It’s what happens once the context gets this large. Long-context inference usually comes with a painful tradeoff. As the KV cache grows, memory pressure increases and generation can slow down. But this setup is still pushing through 364K tokens while maintaining a usable agent workflow. The model can reason, call tools, inspect results, continue working, and keep the session moving. And the tool calls reportedly haven’t missed so far. That’s important for agentic coding. A huge context window is only useful if the model can actually operate reliably inside it. A 400K-token context that constantly breaks tool calls isn’t very useful. A 364K session that can keep reasoning and executing tools is a different story. And the test isn’t finished yet. The current run is approaching 400K tokens, with the expectation that it can keep going. This is also another interesting example of why Apple Silicon keeps showing up in local LLM experiments. The M5 Max’s unified memory gives a large model and its growing KV cache access to one shared memory pool. With MLX-Serve continuing to improve, these machines are becoming surprisingly capable long-context inference boxes. The bigger takeaway: Context length is becoming a workload, not just a model specification. Running a model at 256K is one thing. Keeping an agent alive at 300K+ while it reasons and uses tools is much more interesting. And Qwen3.8-Flash-Next is showing that this can be pushed surprisingly far on a single 128GB Mac. 364.7K and counting. Next stop: 400K.

FHILY👑

39,982 görüntüleme • 23 gün önce