Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

2/It's a step up on reasoning and coding. compared to Muse Spark 1.2, Muse Spark 1.3 wastes fewer turns, uses ~20% fewer tool calls, ~25% fewer tokens, and holds onto requirements well during long-horizon tasks.

35,480 görüntüleme • 1 ay önce •via X (Twitter)

6 Yorum

Alexandr Wang profil fotoğrafı
Alexandr Wang1 ay önce

1/ today we’re releasing muse spark 1.3—available in muse code & the meta model api. this is our most capable model yet—frontier performance almost too cheap to meter. much stronger at agentic and coding with better usability. we think users will really notice the jump.

Alexandr Wang profil fotoğrafı
Alexandr Wang1 ay önce

3/ available now in Meta Model API and Muse Code. we are cooking! bigger and more capable muse models are close behind 🍉, along with Muse Spark open weights release, and exciting new products. more here →

Vector profil fotoğrafı
Vector1 ay önce

Thats a solid upgrade. Fewer tool calls and tokens usually means less wasted time, so Im curious how it feels in real use. Coding especially needs that efficiency.

Mustapha profil fotoğrafı
Mustapha1 ay önce

I can't wait to test it soon

The transfer Account profil fotoğrafı
The transfer Account1 ay önce

this is observed directly through my usage, amazing work by engineers

Hamza 🌙 profil fotoğrafı
Hamza 🌙1 ay önce

this is the metric I want every coding-model launch to expose: accepted task / tool call accepted patch / token +2 benchmark points is nice. ~20% fewer tool calls + ~25% fewer tokens can matter more if acceptance stays flat. agent loops pay for wandering.

Benzer Videolar

gemini 3.7 flash vs deepseek v4 pro 0813 vs muse spark 1.2 – on voxel city dioramas three models each built three crossy road-style 3d scenes – a construction site, a nyc intersection, a river with a drawbridge – as single self-contained html files the setup: Nous Research's hermes agent cli on OpenRouter, three.js skills preloaded, identical prompts per scene tasks: 1. construction site – tower crane on a working lift loop, paver laying fresh road, roller compacting it behind 2. nyc crossing – four-way intersection with a traffic light state machine, queuing cars, pedestrians crossing on the walk signal 3. river drawbridge – double-leaf bascule that lifts for tall boats, cars queuing at the barriers, animated water every scene: Three.js r185, box geometry only, a locked 20-color palette, four camera presets, and a day/night mode with bloom. one file, no build step, no assets models: Google DeepMind gemini 3.7 flash, DeepSeek v4 pro 0813, AI at Meta muse spark 1.2 muse and gemini finished every scene in two to three minutes. deepseek took 15 to 41 minutes per scene - build time, all three scenes #1 gemini 3.7 flash – 6m 43s #2 muse spark 1.2 – 7m 20s #3 deepseek v4 pro – 91m 25s - total tokens #1 muse spark 1.2 – 440,279 #2 gemini 3.7 flash – 713,855 #3 deepseek v4 pro – 20,957,568 - total price #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57 - agent calls across the three builds #1 muse spark 1.2 – 12 #2 gemini 3.7 flash – 18 #3 deepseek v4 pro – 143 observations: • muse won two of the three scenes on looks with the smallest files in the test – 887 to 1,042 lines against gemini's 1,934 to 2,377. cheapest, fastest to a good frame, and shortest turned out to be the same column • deepseek burned 20.96m tokens – 29x gemini, 48x muse – across 143 agent calls. prompt caching is the only reason that cost $4.57: the cache discount absorbed roughly $30 of resent context • gemini was the only model whose files needed zero fixes to render – and the only one whose night mode is cosmetic. the sky never darkens and one camera button does nothing. clean code for a scene it never looked at follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

29,033 görüntüleme • 1 ay önce

meta muse spark 1.1 vs gpt 5.6 sol vs fable 5 vs grok 4.5 meta recently dropped muse spark 1.1 – a multimodal reasoning model from meta superintelligence labs built for agentic tasks. key facts: • 1m token context with active self-management – the model compacts its own history and keeps only the steps needed for later work • trained to orchestrate multi-agent systems: as main agent it plans and delegates to parallel subagents, as subagent it sticks to its job and knows when to escalate back • computer use trained to pick between scripting and clicking – writes automation when it's faster, clicks when it's simpler, batches actions per step • first public api from meta: the meta model api is now in preview • benchmarks: sweeps the agent column – mcp atlas 88.1 (opus 4.8: 82.2), jobbench 54.7 (opus: 48.4), humanity's last exam 62.1 (1st). loses coding – deepswe 1.1 53.3 vs gpt 5.5's 67.0, swe bench pro 61.5 vs opus's 69.2 our test – 3 prompts, single-file html, three.js, fully procedural, no assets: 1. norwegian house cantilevered over a fjord in a snowstorm – transmissive glass wall, fully modelled interior 2. beijing siheyuan courtyard house in dawn fog – instanced roof tiles, dougong brackets, glowing paper windows 3. new mexico adobe pueblo in an approaching dust storm – deep window reveals, windward grit accumulation we ran the test on AI/ML API platform results: - cost #1 muse spark 1.1 – $0.20 #2 grok 4.5 – $0.51 #3 gpt 5.6 sol – $1.93 #4 fable 5 – ~$5.20 - output tokens #1 muse spark 1.1 – 41,868 #2 gpt 5.6 sol – 49,139 #3 grok 4.5 – 64,954 #4 fable 5 – 81,849 - lines of code #1 muse spark 1.1 – 1,799 #2 gpt 5.6 sol – 2,377 #3 fable 5 – 3,088 #4 grok 4.5 – 4,216 observations: • muse spark is the cheapest of the four by a wide margin – 2.5x under grok, ~26x under fable per run. output quality tracks the price • only 7.4% of its output tokens are reasoning (3,104 of 41,868) – the model barely thinks before writing. economic, not pedantic: it commits to the first plan and ships it • the low loc is not compression, it's omission – all three prompts demanded instancing, muse spark delivered it in one muse spark's code quality – reviewed by fable 5: upsides: 1. all three files run 2. the adobe grit effect is legit – shader injection via onbeforecompile, windward faces detect storm direction through a normal-dot-wind term and darken procedurally 3. the fjord glass is real meshphysicalmaterial with transmission and ior, not a transparent quad 4. the siheyuan properly instances barrel tiles, dougong blocks and courtyard pavers downsides: 1. in the fjord file the strafe vector is negated – press a, you move right; press d, you move left. exactly the key mix-up we kept hitting with this model 2. all three files ship the model's self-doubt as comments: "// actually yaw orientation: need correct" sits above a direction vector that gets computed, abandoned and recomputed – dead vectors allocated every frame, 60 times a second 3. the siheyuan registers two separate keydown listeners, one containing an empty if-block 4. snow "accumulation" on the norway roof is a sine wobble on a scale value, not accumulation 5. "instanced snow" became 3,500 plain points. zero dispose calls anywhere pattern: minimal reasoning, minimal code, minimal price. it nails the flashy requirements – shaders, transmissive glass – and quietly drops the boring ones: instancing, controls, cleanup. you get a demo that mostly runs and a control scheme you can't trust follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

135,556 görüntüleme • 2 ay önce

union alpha (unbiased pareto) vs deepseek v4.1 flash vs muse spark 1.3 – three paintings in three.js the setup: one four-line prompt plus the painting as an image, through OpenRouter. no agent loop, no renders, no feedback – the model writes one html file blind and we open it. Three.js from a cdn, every texture generated in code. when a provider cut the stream early we sent the partial back and said continue exactly where you stopped tasks: 1. the starry night – van gogh, 1889 2. the persistence of memory – dalí, 1931 3. poppies at argenteuil – monet, 1873 two rules in every brief: keep the painting's palette, brushwork and mood, and reply with the code only models: Unbiased AI union alpha (stealth, free), DeepSeek deepseek v4.1 flash, AI at Meta muse spark 1.3 total cost, three scenes #1 union alpha – free (list price: $1.04) #2 muse spark 1.3 – $0.117 #3 deepseek v4.1 flash – $0.162 generation time, three scenes #1 muse spark 1.3 – 4m 45s #2 deepseek v4.1 flash – 11m 30s #3 union alpha – 29m 13s total completion tokens #1 muse spark 1.3 – 26,502 #2 union alpha – 123,756 #3 deepseek v4.1 flash – 156,028 lines of code shipped #1 muse spark 1.3 – 862 #2 union alpha – 1,715 #3 deepseek v4.1 flash – 2,816 observations: • union alpha reads the painting like an art historian. it named every work unprompted, then broke each into parts: dalí's watches deformed along a bezier curve, monet's poppies as instanced brush dabs under a wind shader. no other model went that deep on a one-shot • it is the only model that made the paintings move the way they were painted. in the monet, the woman and the child walk the field on catmull-rom paths, pollen drifts, poppies are brush dabs in a point shader that sway in the wind. deepseek and muse left the figures standing • its dalí is an inventory of the canvas: three soft clocks draped along one parametric curve, a drip falling off the hanging one, a fly, ants on the pocket watch as an instanced mesh. 17 named parts in all. nobody else drew the drip • its starry night is shader work end to end: shared glsl noise, a vortex field for the sky, billboarded shader quads for the moon and stars, a painterly surface shader for the hills, and windows that flicker on their own timers. the file reads like a demoscene entry, not a model output what union alpha is: • we asked it. the stealth window had closed a day after launch and the api answered: "this model was unbiased's pareto". pareto is from circuit & chisel, an ex-stripe team that raised $19.2m in sep 2025 per fortune, and now sells "frontier intelligence for 75% less" • pareto is not one model. per unbiased's site it "runs a mix of frontier and open source models against each other on every request" and keeps the best answer. that is the 300-second first token, the tokenizer listed as "other", and the missing reasoning field – a race, not a model • listed at $2.50 in and $7.50 out per million. our three scenes would have cost $1.04 – 6.4x deepseek, 8.9x muse. asked for its cutoff, it dated nothing past may 2025 follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

23,566 görüntüleme • 18 gün önce