Загрузка видео...

Не удалось загрузить видео

На главную

I tested GLM 5.3 Flash vs Kimi K3. Same prompts, both capped at $2 per scene. > Same cost, but GLM Flash couldn't build realistic scenes with proper physics. > Kimi K3 nailed it. It actually seemed to understand how objects should move and interact. What are your thoughts?

27,723 просмотров • 5 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 просмотров • 12 дней назад

anthropic will sell you opus 5 at $200 a month. openai will sell you gpt-5.6 at $200 a month. neither will tell you stanford and berkeley published the 5 principles to build a $100k/mo ai company on kimi k3 for $10 stanford and berkeley spent years figuring out what actually separates ai systems that work in production from ai systems that die in demos. they published the findings. anthropic and openai priced their frontier subs like nobody would read the papers. the papers are free this is dspy plus verifiers plus decomposition plus skills plus mcp. five principles from stanford, berkeley and moonshot that turn a $10/mo kimi k3 sub into an ai analyst that runs unattended. the model is public. the system is the moat five moves that turn kimi k3 into the $100k/mo company: P1 don't prompt, program (stanford dspy) -> stanford proved hand-tuned prompts don't scale. define a pipeline as modules, let the optimizer tune them -> the compiled pipeline beat expert few-shot on multi-step tasks. one line of dspy replaces a month of prompt engineering P2 don't trust the model, build verifiers (berkeley 2026) -> a compiler either accepts or rejects. a test either passes or fails. that is a verifier -> berkeley: test-suite reward hit 42.2% pass@1 on swe-bench. hybrid verifiers hit 51.0% best@26. no bigger model, just a real check P3 don't scale agents, decompose them (stanford ai index 2026) -> stanford found multi-agent gains only 2-4 percentage points. two coding agents sometimes did worse than one -> the win is role decomposition, not count. researcher, writer, reviewer, verifier, clear input, clear output, no overlap P4 don't repeat expertise, encode it as skills (kimi code) -> every session starting from zero is institutional knowledge you lost. a skill.md file makes kimi activate the workflow automatically -> week one you write the skill. month six it encodes more institutional memory than most junior employees carry P5 don't keep ai in chat, connect it to tools (mcp) -> a model that only sees what you paste is a consultant working blindfolded. mcp connects kimi to your crm, db, github, linear, slack -> the model is public. the data is yours. the connections are your moat my position, and it is the arguable one: the next $100k/mo ai company will not win because it got early access to a frontier model. it will win because it followed 5 papers that anthropic and openai are quietly hoping you never read drop your $200/mo ai sub to $10. the swarm above is what 300 kimi k3 agents look like running those 5 principles. the full playbook is in the article below

starmex

31,358 просмотров • 14 дней назад

UPDATE: Charlie Kirk 🚨 Muzzle Flash: Second Shooter Location, Reflection, or Both? This video captures a flash on a window right as Charlie Kirk was shot. Let’s break this down… follow the numbers on the videos. #1. This video shows what appears to be a muzzle flash, just as Charlie is shot, leading people to believe the shooter was to Charlie’s right.. Notate the people to the left of the flash #2. This photo shows the location of the flash. Notate people to the left, on a middle-landing on a staircase. #3. This photo shows a clearer image showing both the middle-landing and the same location of the flash. It is a window. #4 . This video shows the building behind Charlie, which is a long hallway. This debunks any claims the shot came from here. — I zoom in where Charlie was — I zoom in on staircase — I zoom in on the window/flash — I zoom in on where suspected shooter was #5. I notate the flash reflection angles. — Video angle #1 notated — two reflection paths notated in relation to Video angle #1. 1. Reflection angle shows where we were told the shooter was. 2. Reflection angle shows mirrored angle, obviously where there is no reports of a shooter being in. 🔻 Final conclusion: This to me is very likely a muzzle flash reflection from a far off location. I am unsure of the reflection angle however, but this can be 100% proven if someone were to recreate the angles… maybe with a flash camera. If anyone is willing to, or has the means to do this (safely), this will either prove the shot came where the FBI said it came from, or it will prove there was a second shooter, in the direction of the mirrored angle.

MJTruthUltra

10,665,164 просмотров • 11 месяцев назад

A lesson for every Polymarket bot developer: I built a strategy that looked perfect on paper. Backtested it. Looked like a winner. Almost went live. Then i actually measured real costs. Strategy was dead before the first trade. And this is what bot building on Polymarket actually looks like. Here is what happened (and what you MUST know): Backtested mean reversion on crypto dips. SOL came back at +44% return and 70.5% win rate. Beautiful clean curve. Looked ready to ship. Then i measured real round-trip costs on SOL flash dips. Backtest assumed 0.45% in fees and slippage. Reality was 1.44%. Strategy stops working at 0.70%. Starting again. But the lesson was worth more than any profit the strategy could have made. Here is what i actually learned: Taking a dip with a market order means you eat the spread the dip just created. The volatility making your signal is the same volatility destroying your fill. You see the opportunity. You enter. You already lost. But resting a limit order below market and letting the dip come to you? You collect the maker rebate instead. Same thesis. Completely opposite execution. One bleeds money, one prints it. That one realization changed how i think about bot strategy entirely. 180 strategies tested to get there. 179 dead. That is not failure. That is how you find the 5 that actually work. Building a bot on Polymarket is not about finding a magic strategy. It is about eliminating every wrong answer until only the right one is left.

Oracle Boar

14,379 просмотров • 4 месяцев назад