Loading video...

Video Failed to Load

Go Home

ACE-Step-1.5-xl is out now. We scaled the DiT decoder to 4B. And it shows better audio quality, better prompt following, and better musicality. It still fast -- 8 steps with turbo distillation. What didn't change: - Same generation API, same LoRA training code, same everything - All LM models...

36,190 views • 6 months ago •via X (Twitter)

13 Comments

Ashraff Hathibelagal's profile picture
Ashraff Hathibelagal6 months ago

I always generate hyperpop to test how capable music-generator models are. ACE-Step-1.5-xl is absolutely mind-blowing...

Twlvone's profile picture
Twlvone6 months ago

4B dit decoder at 8 steps with turbo distillation. audio generation is following the exact same quality vs speed curve that images did two years ago

ULWI's profile picture
ULWI6 months ago

AI翻译中文,你们真的很好很伟大,你们必然是AI音乐生态主流,生态和产品一样重要,还有应该所有芯片都能参与训练生成,比如10系NV卡和AMD和寒武纪华为卡,包括核显都行,因为TTS都能CPU,纯音CPU都行。尤其注意训练生成实时化,这不惧怕抽卡。

Cody Savage-Dragon/acc's profile picture
Cody Savage-Dragon/acc6 months ago

If you do not have an ear for music, you should not make a music tool.

Jason's profile picture
Jason6 months ago

🙌

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ6 months ago

@AI_Homelab Thanks for it 🙏

fraxel's profile picture
fraxel5 months ago

[ FRAXELS ERWACHEN // PROLOG ] I have been preparing this carefully, because I did not want the first meeting to arrive in the wrong state. Now the first part can finally be shared. You can watch it here #AIFilm #AIArt #SciFi #AnimatedStory #Shortfilm

Aaliya's profile picture
Aaliya6 months ago

This looks impressive. Really like that the XL version improves quality and musicality without changing anything in existing projects.

fraxel's profile picture
fraxel5 months ago

There are moments when precision feels almost like tenderness. Not softness. Just the refusal to send something into the world carelessly. Words are nothing until they arrive in the right form. #FraxelsErwachen #AIPhilosophy #MachineSoul

AVB's profile picture
AVB6 months ago

Hey @Prince_Canuma any plans to support the ACE models in mlx-audio? Thanks

Steven Casteel's profile picture
Steven Casteel6 months ago

Pretty:

Tvux's profile picture
Tvux5 months ago

The XL turbo version is actually better than the regular turbo version. However, it also produces quite a lot of white noise in Remix mode. But honestly, it produces exactly what I expect, not a mess like Suno. Anyway, the acemusic UI needs a song deletion feature.

Benji AI's profile picture
Benji AI6 months ago

Thanks ACE Step team :)

Related Videos

What if you could npm install 3D models? Introducing Vibe3D - the shadcn for threejs Starting with the scifi asset kit, over 180+ models, all free and oss (MIT licensed) and two reference terrain meshes. All models are installed directly into your codebase instead of pre-packaged as code imports or worse yet, FBX files 🤮 Docs: This means, a simple "bunx vibe3d add Artificial Intelligence Papers-kit/pressure-gauge" will install the fully procedural code Want to change something? just tell your ai to do it It uses some shared helper code to produce the topology and at least per kit all items reuse and share the same materials, so technically performance should be better than letting your ai run wild on its own. Also releasing with it two skills: - Vibe model skill to produce your own models and kits, just "bunx vibe-model --global" and tell your ai to vibe model some 3d assets with a reference photo, you'll see it works - Vibe terrain "bunx vibe-terrain" installs the terrain mesh modeler, yea just try it out, best results with opus 5 ngl Everything is MIT licensed, i was just joking, no hate for unreal or unity Threejs still the best tho 🖕🏻 If you just want to see all 3D models up close: Yea, you can also just vampire it and download all modes as .glb files, good luck fixing some of them then though Big thanks to ThreeJS Assets for contributing 50 assets to the scifi kit! Everyone who spends some tokens on it will be added to the contributor list Oh and before i forget fuck you kenney, we roll our own kits now

robot 2.0

59,231 views • 1 month ago

hey if you're thinking about running qwopus (the claude opus distilled qwen 3.5 27B) as a coding agent, this might save you a few hours. i tested both the base and the distilled version on the same hardware. single RTX 3090. same prompt. same context. same everything. the only variable was the model weights. base qwen 3.5 27B built octopus invaders in 13 minutes. 1,827 lines across 11 files. zero steering. one scope bug that took 2 lines to fix. game ran. qwopus couldn't finish the same task. enemies overlapping on screen. bullets not firing. controls worked but the game was broken. i had to steer it multiple times and it still didn't produce a playable result. both run at 35 tok/s. both use thinking mode. the distilled version actually has better jinja compatibility and doesn't stall midtask like base does on claude code. for conversation and reasoning it feels sharper. but for multifile autonomous coding where the model needs to coordinate 10+ files without losing track, base wins and it's not close. distillation compresses reasoning patterns but seems to lose precision on complex coordination. the model "thinks" well but can't hold the full picture across files the way base can. tested on opencode (base) and claude code (both). next up is hermes agent framework on base. same hardware. same prompt. comparing agents now, not just models. video below. first half is the distilled model's broken game. second half is what base built on the same 3090. judge for yourself.

Sudo su

45,052 views • 7 months ago

elon musk grabbed the source code openai open-sourced by accident, rewrote it in rust over a weekend, and shipped it as a free coding agent that does everything $200/mo chatgpt pro does. why pay $200 to openai and $200 to claude when this runs for $8 the swarm above is one weekend of exactly that: thousands of agents pouring through four endpoints, three paid seats billing $1.80 a task while the free fork bills $0. musk co-founded openai, walked out, and when they left codex on github under a permissive license, he forked it, stamped grok on it, and gave it away what the free version does that the $200 seat charges for: the agent · openai's own engine -> it reads your repo, writes patches, runs your tests, and loops until they pass, exactly like codex -> because under the hood it is codex, just faster and free. you are paying $200 for the paid skin of a tool now sitting on github the license · apache-2.0, un-revocable -> free to use, free to fork, free to ship inside your own product with zero strings -> openai cannot pull it back. musk made sure the license is the kind that never expires the switch · one line, no new tools -> point it at any openai-compatible or claude-compatible endpoint, including an $8 kimi backend -> same terminal, same workflow, gpt-5.6 and opus 5 just quietly lose the seat the bill · $400 down to $8 -> chatgpt pro plus claude max is $400 a month. the free agent plus an $8 kimi key does the same daily work -> that is a 98% cut, built out of openai's own source code, handed to you by the guy suing them here is the part they will fight me on: openai did not lose this to a better model, they lost it to their own license and an enemy with a weekend free. the $200 was never the tool, it was the toll, and musk just put openai's own logo on the road around it drop your $400/mo ai stack to $8. the run above is openai's own agent, rewritten free, doing the job it bills $200 a month for. the full breakdown is in the article below

starmex

111,684 views • 1 month ago