正在加载视频...

视频加载失败

sonnet 5 vs sonnet 4.6 vs opus 4.8 vs glm 5.2 – frontend tasks dropped sonnet 5 into a quick test today. same three prompts to all four models, single-shot html/canvas, no edits: • objects falling on a trampoline • rockets playing tennis • a slingshot breaking bottles ranked...

14,518 次观看 • 2 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

hy3 vs fable 5 vs opus 4.8 vs sonnet 5 Tencent Hy just dropped hy3 – their new open-weight model under apache 2.0. following the april preview they scaled up post-training, and it now rivals flagship open models with 2-5x the params. api pricing: ~$0.15 in / ~$0.59 out per 1m tokens. built for coding, office work, frontend, agentic tasks so we ran a test: hy3 vs fable 5 vs opus 4.8 vs sonnet 5 three prompts, one-shot each: • ocean wave crumbling a sand castle (canvas) • looping factory assembly line (html/css/js) • interactive 3d city with three.js + orbitcontrols self-contained files, no libraries beyond the cdn where asked totals across all three prompts: 1. hy3 – 1231 loc / 14m34s 2. sonnet 5 – 1373 loc / 18m55s 3. fable 5 – 1546 loc / 18m32s 4. opus 4.8 – 1904 loc / 27m21s hy3 is the fastest and the leanest by a wide margin we had opus 4.8 analyze hy3's code. the read: - sand castle: checklist-complete but the crumble is parametric, not physical. it shrinks and slumps the towers and fades alpha instead of dissolving into grains. the cheap-but-plausible interpretation. the tell of a smaller model - factory line: the arm-to-part sync is actually causal, not faked. it triggers each robot early by exactly the arm's descent time, so the tap lands right as the part arrives. it also pre-seeds the belt so it never cold-starts empty. clean state machine. one latent bug – a part gets marked processed before checking if the robot is free, so at a faster spawn rate a "laptop" could ship missing a part. never fires at current timing, but the invariant isn't enforced - 3d city: genuinely frontier-adjacent. correct modern setup (pcfsoft shadows, srgb, aces tone mapping, damped orbit + auto-rotate pause). clones the window texture per building and scales the uv repeat to each building's dimensions so windows don't stretch. downside: no instancing – ~800 texture clones across 200 buildings. runs fine, not optimized. roads are implicit gaps, not explicit planes our observations: • hy3 is quite fast • its animations are simple but you can see it trying – it adds detail, and the 3d render sits at the same level as the frontier models • sonnet 5 is weak here. hy3 beats it on the sand castle and the 3d render, level on the conveyor • opus 4.8 is anthropic's best model after the fable 5 nerf – it beats fable on the conveyor and the 3d render net: hy3 runs clean and well-formed across all three with zero syntax errors, even version-matching the three.js core and examples build. it's economical rather than ambitious – it does the minimum viable version of each hard requirement well, and only reaches for the expensive interpretation on the 3d task a very coherent profile for a cost-optimized open-weight model follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

34,961 次观看 • 2 个月前

fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past Three.js from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through OpenRouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: Anthropic fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

18,509 次观看 • 16 天前

Someone ran Claude Code on a beach where any device overheats and that spot suddenly turned out to be the best home for the most powerful AI in the world. This is the reMarkable Paper Pro. A paper tablet for notes with no browser and no social media and not a single app. He sat down right on the sand in the open sun and brought up Claude Code on Opus 4.6 over the Claude API on the paper screen and opened his project ~/repos/webs while the waves broke a few steps away. For years every device had the same trouble outside. In direct sun the screen glares and washes out and heats up and instead of your work you see your own reflection. But e-ink does not blast its own light into your face. It reflects the sunlight like the page of a book. And here is what came out of it. The very thing that kills any normal screen outside turned into fuel for this one. The brighter the sun the sharper the picture because it has nothing to glare with and nothing to wash out. And then comes the thing no laptop on a beach will give you. Your eyes do not get tired. You can watch Opus think on max effort for an hour and it reads like a book in the sun and not a backlight you squint into. The picture only comes alive. In bright light it does not fade but turns sharper and higher in contrast than it ever was in a room. The charge lasts for days. E-ink barely touches the battery so there is no outlet anywhere on the sand and the tablet does not care. It weighs as much as a notebook. The whole setup folds into a beach bag like a pad with a pen on top. Everything on the screen is for real. Claude Code v2.1.110 and Opus 4.6 on the Claude API and the project ~/repos/webs open right on the e-ink in the middle of the sand. In my opinion this is the most unexpected home for an AI this year. Not an office with the blinds drawn and not a monitor cranked to full brightness but a quiet sheet of paper on the sand that open sun only makes better and on it the most powerful Claude writes code right on the page like a pen.

Blaze

89,575 次观看 • 2 个月前

Today, we’re releasing Athena (mvrko-sim-1), the flagship Large Event Model from markopolo.ai that beats GPT-5.6, Claude Opus 4.8, and Claude Sonnet 5 in OPeRa public benchmark for Shopping behaviour prediction! Athena introduces a completely different approach to understanding digital behavior. Most systems record actions after they happen. Athena models the sequence behind them to predict what a shopper will do next and exactly where they will do it, before the action even takes place. Athena vs. frontier models: We evaluated Athena on the full OPeRA test set alongside GPT-5.6, GPT-4.1, Claude Sonnet 5, and Claude Opus 4.8. Athena achieved 24.5% strict exact match scoring the highest among the five models evaluated. Athena is not a general-purpose language model. It is a Large Event Model built to understand behavior as a connected sequence. It combines the page someone is viewing, the actions that brought them there, and the goal they are trying to complete and then predicts the exact next browser action and target element. That distinction matters. Today’s digital systems largely react after someone clicks, abandons, purchases, or leaves. Athena creates the foundation for systems that can understand what is likely to happen before the outcome. Together with Rubaiyat, and the team at markopolo.ai, we built Athena from years of work on shopper behavior, digital journeys, and behavioral intelligence. Commerce was our first proving ground. Great thing is, we’re now releasing Athena publicly so builders can adapt the model to new datasets, industries, and problems. We’re seeing early usecases in Cyber Security, Gaming, Mobile App Ecosystem, Retail and more! Athena is now live on Hugging Face! We built the foundation for relevance and predictability! Now, let’s see what the world builds on top of it. Links in the first reply.

Tasbin

55,854 次观看 • 1 个月前