2/It's a step up on reasoning and coding. compared... to Muse Spark 1.2, Muse Spark 1.3 wastes fewer turns, uses ~20% fewer tool calls, ~25% fewer tokens, and holds onto requirements well during long-horizon tasks.show more

Alexandr Wang
35,480 次观看 • 1 个月前
Muse Glimmer one-shot 5 games! AI at Meta dropped... Muse Glimmer. So we ran it against Muse Spark 1.2, on five one-file game demos, each one written to play itself: Tetris, a top-down pixel street race, a rooftop web-slinger, a blue hedgehog platformer, and Flappy Bird. the setup: • Muse Glimmer 30B — via aimlapi[.]com. cost: $0.02 • Muse Spark 1.2 — via aimlapi[.]com. cost: $0.11 what we saw: • every Glimmer scene runs and plays itself. it's a reasoning model, so it burns more tokens — but the output is clean. its Tetris even scores each placement to keep lines clearing, and the blocks burst into pixels on every clear. • Glimmer beat Spark at Flappy Bird — its bird flew through the pipes cleaner. • but it's not perfect: in the racing game its cars sometimes slip off the road, and its graphics are simpler and blockier than Spark's. glad to see Meta back in open source — and we're waiting for Spark 1.2 weights next 🙌show more

AI/ML API
17,476 次观看 • 1 个月前
One of the best uses for Muse Spark: generating... HTML explainers. Dense design docs are painful to review. Point Muse Spark at one and get back an explainer you can click through — before/after states, the technical reasoning, all of it. Works on PRs too. Muse Spark strikes a good balance of speed, quality, and cost for this kind of task. Free for a limited time in OpenCode — which works inside Orca.show more

Orca ADE
18,921 次观看 • 1 个月前
Muse Spark 1.2 supports a broad range of multimodal... tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities. Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck. 🧵👇show more

AI at Meta
58,163 次观看 • 1 个月前
gemini 3.7 flash vs deepseek v4 pro 0813 vs... muse spark 1.2 – on voxel city dioramas three models each built three crossy road-style 3d scenes – a construction site, a nyc intersection, a river with a drawbridge – as single self-contained html files the setup: Nous Research's hermes agent cli on OpenRouter, three.js skills preloaded, identical prompts per scene tasks: 1. construction site – tower crane on a working lift loop, paver laying fresh road, roller compacting it behind 2. nyc crossing – four-way intersection with a traffic light state machine, queuing cars, pedestrians crossing on the walk signal 3. river drawbridge – double-leaf bascule that lifts for tall boats, cars queuing at the barriers, animated water every scene: Three.js r185, box geometry only, a locked 20-color palette, four camera presets, and a day/night mode with bloom. one file, no build step, no assets models: Google DeepMind gemini 3.7 flash, DeepSeek v4 pro 0813, AI at Meta muse spark 1.2 muse and gemini finished every scene in two to three minutes. deepseek took 15 to 41 minutes per scene - build time, all three scenes #1 gemini 3.7 flash – 6m 43s #2 muse spark 1.2 – 7m 20s #3 deepseek v4 pro – 91m 25s - total tokens #1 muse spark 1.2 – 440,279 #2 gemini 3.7 flash – 713,855 #3 deepseek v4 pro – 20,957,568 - total price #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57 - agent calls across the three builds #1 muse spark 1.2 – 12 #2 gemini 3.7 flash – 18 #3 deepseek v4 pro – 143 observations: • muse won two of the three scenes on looks with the smallest files in the test – 887 to 1,042 lines against gemini's 1,934 to 2,377. cheapest, fastest to a good frame, and shortest turned out to be the same column • deepseek burned 20.96m tokens – 29x gemini, 48x muse – across 143 agent calls. prompt caching is the only reason that cost $4.57: the cache discount absorbed roughly $30 of resent context • gemini was the only model whose files needed zero fixes to render – and the only one whose night mode is cosmetic. the sky never darkens and one camera button does nothing. clean code for a scene it never looked at follow thehype. for 24/7 ai news, analysis and breakdownsshow more

thehype.
29,033 次观看 • 1 个月前
Muse Spark 1.1 also excels in perception and multimodal... reasoning, inspecting visual and audio inputs, preserving details across long workflows, and acting on them in real execution environments. It shows particular strengths in visual-to-code generation, rich image/video captioning, and agentic computer use. In this demo, using video shot from a smartphone, Muse Spark 1.1 extracts useful photos and reasons about the product to operate a user's browser and make a Facebook Marketplace listing on the user's behalf.show more

AI at Meta
74,235 次观看 • 2 个月前
Our most intelligent workhorse model yet for coding and... agents has arrived ⚡ Meet Gemini 3.7 Flash. — Crush that seemingly endless to-do list. Gemini Spark in the Google Gemini now uses 3.7 Flash. The new model can equip your personal AI agent to work even smarter for you by seamlessly handling complex, multi-step tasks across your Google Workspace apps like Gmail, Google Calendar and Google Docs — Enjoy a smoother build experience. The model thinks more diligently, putting more effort into multi-step planning and tool calls. A more disciplined execution means less manual oversight and fewer retries across engineering workflows — Build more, spend less. 3.7 Flash is available through the end of the year with an introductory price of half the original 3.6 Flash cost per million tokens ($0.75/1M input tokens and $3.75/1M output tokens)show more

Google AI
129,191 次观看 • 1 个月前
unlimited meta muse spark 1.2 for free until december... 31, 2026 😳 a tool you may not have heard of just did something insane: they released meta's latest model completely free. freebuff copied every feature in Lovable, bolt, replit, and base44, and released it for free with 0 paywalls. full setup guide: 1 / go to and click on Web 2 / create a free account 3 / select Muse 1.2 from the dropdown 4 / type in your prompt. it never even asks for a card. it's that simple. the direction coding tools are moving into are insane. soon, we might see Fable 5 completely free for everyoneshow more

Victor Cheng
50,895 次观看 • 2 个月前
Muse Spark 1.3 is not that good, and Gemini... Flash 3.8 is just pure benchmark-maxxing. Gemini even choked on the first run and needed a second attempt just to produce a working scene. By stripping colliders and enforcing rigorous backend-level logic in 3D space, this benchmark immediately separates true reasoning from trained test shortcuts.show more

OmedTheVibeCoder
29,403 次观看 • 1 个月前
🚀 Anthropic's Claude Opus 5.5 is now generally available... in GitHub Copilot. Our early testing shows it stands out in efficiency, delivering task resolution comparable to Claude Opus 5 while using significantly fewer steps and tokens, with fast error recovery on multistep tasks. Try it out for agentic coding, long-running agentic tasks, and knowledge work. Find it in the GitHub Copilot app, CLI, and Visual Studio Code.show more

GitHub
123,493 次观看 • 14 天前
1/ Karen 11 and this clown are still gaslighting... you on Minneapolis crime stats saying "cRimEz iZ dOwN" on a *** three-year average ***. Uh, crime was at all-time highs during the last 3-4 years. Saying it's down from that is GASLIGHTING. Not to mention, fewer people are reporting, as we've posted numerous times, and there are fewer cops to take reports. See 2/⬇️show more

CrimeWatchMpls
180,441 次观看 • 2 年前
DeepSeek V4.1 Flash is a very interesting model. I... just tested it on the BridgeBench lava lamp test and it took LONGER to complete than Fable 5.1 and GPT 6 Astra. It ran at 344 toks/sec, spent 23.5M tokens with a cache hit rate of 99.7%, and cost $0.33. Even though it runs so fast it ends up completing tasks very slowly. The lava lamp that it produced was better than Muse Spark 1.3 and Gemini 3.8 Flash.show more

BridgeMind
128,202 次观看 • 27 天前
Someone built a browser designed from the ground up... for humans and AI agents to share. It's called ego lite. -Your agents run tasks in their own Spaces while your tabs stay yours -No extra setup - agents access your real logins and tabs directly through ego-browser -Faster task completion on fewer tokens than any existing browser automation framework Every other tool makes you run a separate browser. This one just shares yours. Github:show more

0xMarioNawfal
51,768 次观看 • 2 个月前
Muse Spark 1.3 Ultra Contributor vs Fable 5.1 xHigh... Same prompt, both ran for roughly 2 hours The prompt had a self improvement rule: if the independent judges scored the result below 9.5/10, it had to keep improving and try again, What shocked me most was that Muse Spark just NEVER stopped, It went through 20 self improvement loops, launching 3 agents each time basically 60+ agent runs over 2 hours And the cost?, Less than $1, Yes, under $1… in Ultra mode ,This completely exceeded my expectations Huge credit to Meta, this thing is a work of artshow more

Iam_
150,400 次观看 • 1 个月前
meta muse spark 1.1 vs gpt 5.6 sol vs... fable 5 vs grok 4.5 meta recently dropped muse spark 1.1 – a multimodal reasoning model from meta superintelligence labs built for agentic tasks. key facts: • 1m token context with active self-management – the model compacts its own history and keeps only the steps needed for later work • trained to orchestrate multi-agent systems: as main agent it plans and delegates to parallel subagents, as subagent it sticks to its job and knows when to escalate back • computer use trained to pick between scripting and clicking – writes automation when it's faster, clicks when it's simpler, batches actions per step • first public api from meta: the meta model api is now in preview • benchmarks: sweeps the agent column – mcp atlas 88.1 (opus 4.8: 82.2), jobbench 54.7 (opus: 48.4), humanity's last exam 62.1 (1st). loses coding – deepswe 1.1 53.3 vs gpt 5.5's 67.0, swe bench pro 61.5 vs opus's 69.2 our test – 3 prompts, single-file html, three.js, fully procedural, no assets: 1. norwegian house cantilevered over a fjord in a snowstorm – transmissive glass wall, fully modelled interior 2. beijing siheyuan courtyard house in dawn fog – instanced roof tiles, dougong brackets, glowing paper windows 3. new mexico adobe pueblo in an approaching dust storm – deep window reveals, windward grit accumulation we ran the test on AI/ML API platform results: - cost #1 muse spark 1.1 – $0.20 #2 grok 4.5 – $0.51 #3 gpt 5.6 sol – $1.93 #4 fable 5 – ~$5.20 - output tokens #1 muse spark 1.1 – 41,868 #2 gpt 5.6 sol – 49,139 #3 grok 4.5 – 64,954 #4 fable 5 – 81,849 - lines of code #1 muse spark 1.1 – 1,799 #2 gpt 5.6 sol – 2,377 #3 fable 5 – 3,088 #4 grok 4.5 – 4,216 observations: • muse spark is the cheapest of the four by a wide margin – 2.5x under grok, ~26x under fable per run. output quality tracks the price • only 7.4% of its output tokens are reasoning (3,104 of 41,868) – the model barely thinks before writing. economic, not pedantic: it commits to the first plan and ships it • the low loc is not compression, it's omission – all three prompts demanded instancing, muse spark delivered it in one muse spark's code quality – reviewed by fable 5: upsides: 1. all three files run 2. the adobe grit effect is legit – shader injection via onbeforecompile, windward faces detect storm direction through a normal-dot-wind term and darken procedurally 3. the fjord glass is real meshphysicalmaterial with transmission and ior, not a transparent quad 4. the siheyuan properly instances barrel tiles, dougong blocks and courtyard pavers downsides: 1. in the fjord file the strafe vector is negated – press a, you move right; press d, you move left. exactly the key mix-up we kept hitting with this model 2. all three files ship the model's self-doubt as comments: "// actually yaw orientation: need correct" sits above a direction vector that gets computed, abandoned and recomputed – dead vectors allocated every frame, 60 times a second 3. the siheyuan registers two separate keydown listeners, one containing an empty if-block 4. snow "accumulation" on the norway roof is a sine wobble on a scale value, not accumulation 5. "instanced snow" became 3,500 plain points. zero dispose calls anywhere pattern: minimal reasoning, minimal code, minimal price. it nails the flashy requirements – shaders, transmissive glass – and quietly drops the boring ones: instancing, controls, cleanup. you get a demo that mostly runs and a control scheme you can't trust follow thehype. for 24/7 ai news, analysis and breakdownsshow more

thehype.
135,556 次观看 • 2 个月前
🌌 GPT-6 Astra from OpenAI Developers is now generally... available in GitHub Copilot. This new model is designed for long-horizon, autonomous coding and agentic tasks. Our internal testing shows it plans and validates as it goes, batches diagnosis with verification, and independently confirms its results before declaring a task done. That translated into stronger performance on extensive coding jobs with fewer steps than prior OpenAI models. Try it out in the GitHub Copilot app, CLI, or Visual Studio Code. 🚀show more

GitHub
175,100 次观看 • 1 个月前
Force is arguably the most overlooked ingredient in modern... robot learning. Introducing FACTR 2: it turns *any* commodity robot into a force-aware system with no force sensors required. Train a tiny force network in <1min with <10mins of data and drop it into any existing teleop pipelines: ✅ Free force sensing for both the robot and the operator arm ✅ Makes demos higher-quality → fewer of them needed. ✅ A new force-aware learning algorithm (FIRST) uses those recovered forces to figure out which parts of a demo actually matter, making learning data-efficient. ✅ Strong performance on complex tasks with fewer demos and even no pretraining! More details below.show more

Deepak Pathak
41,055 次观看 • 3 个月前
union alpha (unbiased pareto) vs deepseek v4.1 flash vs... muse spark 1.3 – three paintings in three.js the setup: one four-line prompt plus the painting as an image, through OpenRouter. no agent loop, no renders, no feedback – the model writes one html file blind and we open it. Three.js from a cdn, every texture generated in code. when a provider cut the stream early we sent the partial back and said continue exactly where you stopped tasks: 1. the starry night – van gogh, 1889 2. the persistence of memory – dalí, 1931 3. poppies at argenteuil – monet, 1873 two rules in every brief: keep the painting's palette, brushwork and mood, and reply with the code only models: Unbiased AI union alpha (stealth, free), DeepSeek deepseek v4.1 flash, AI at Meta muse spark 1.3 total cost, three scenes #1 union alpha – free (list price: $1.04) #2 muse spark 1.3 – $0.117 #3 deepseek v4.1 flash – $0.162 generation time, three scenes #1 muse spark 1.3 – 4m 45s #2 deepseek v4.1 flash – 11m 30s #3 union alpha – 29m 13s total completion tokens #1 muse spark 1.3 – 26,502 #2 union alpha – 123,756 #3 deepseek v4.1 flash – 156,028 lines of code shipped #1 muse spark 1.3 – 862 #2 union alpha – 1,715 #3 deepseek v4.1 flash – 2,816 observations: • union alpha reads the painting like an art historian. it named every work unprompted, then broke each into parts: dalí's watches deformed along a bezier curve, monet's poppies as instanced brush dabs under a wind shader. no other model went that deep on a one-shot • it is the only model that made the paintings move the way they were painted. in the monet, the woman and the child walk the field on catmull-rom paths, pollen drifts, poppies are brush dabs in a point shader that sway in the wind. deepseek and muse left the figures standing • its dalí is an inventory of the canvas: three soft clocks draped along one parametric curve, a drip falling off the hanging one, a fly, ants on the pocket watch as an instanced mesh. 17 named parts in all. nobody else drew the drip • its starry night is shader work end to end: shared glsl noise, a vortex field for the sky, billboarded shader quads for the moon and stars, a painterly surface shader for the hills, and windows that flicker on their own timers. the file reads like a demoscene entry, not a model output what union alpha is: • we asked it. the stealth window had closed a day after launch and the api answered: "this model was unbiased's pareto". pareto is from circuit & chisel, an ex-stripe team that raised $19.2m in sep 2025 per fortune, and now sells "frontier intelligence for 75% less" • pareto is not one model. per unbiased's site it "runs a mix of frontier and open source models against each other on every request" and keeps the best answer. that is the 300-second first token, the tokenizer listed as "other", and the missing reasoning field – a race, not a model • listed at $2.50 in and $7.50 out per million. our three scenes would have cost $1.04 – 6.4x deepseek, 8.9x muse. asked for its cutoff, it dated nothing past may 2025 follow thehype. for 24/7 ai news, analysis and breakdownsshow more

thehype.
23,566 次观看 • 18 天前
This is f*cking Dangerous Cline opened Gemini 3.8 Flash,... DeepSeek V4.1 Flash, and three more for free no API key, right inside VS Code and your terminal pick a free model from the dropdown and start coding, debugging, and running terminal commands without ever touching a billing page what it does: - code, debug, and run terminal commands from one editor panel - switch between 5 models mid-session without editing a single config file - rotate models when one hits its context wall or slows down setup (under 1 min): step 1: go to step 2: install the VS Code extension or open the web version email sign-up, no card, no billing step 3: open Model Picker and select any model marked "Free" start coding important: > Gemini 3.8 Flash, DeepSeek V4.1 Flash, MiMo-V2.6-Flash, Space Bunny Alpha, Muse Spark 1.3 > no API key, no billing just log in and pick > free tiers like this rarely stick around, grab it while it's upshow more

Rivon
534,682 次观看 • 9 天前
Anthropic and Andrew Ng built an agent that uses... 90% fewer tokens from scratch: they dropped the entire book of Frankenstein into a prompt - 108,000 tokens asked one question the input dropped from 108,000 tokens to 11 here's how: step 1 → put everything that never changes at the top - tools, then system, then docs step 2 → mark where the static part ends. everything above it gets cached step 3 → one stray space breaks it - and you pay full price again step 4 → the cache dies in 5 min - every read resets the clock step 5 → cached tokens don't count against your rate limits. free headroom most people never touch this - it pays for itself on day one watch & bookmark - this 1-hour brilliant course ↓show more

codila
546,032 次观看 • 2 个月前