Loading video...

Video Failed to Load

Go Home

🥳 We present our #ICASSP2024 paper: A diffusion model that generates production-quality (bass) audio stems to any audio input. 😎🎸 According to our experience, that's more useful to artists than generating full mixes. 🙃 📜Paper: 🎶Demo: by Marco Pasini 👈💪SonyCSL(Paris)_Music Team #MusicAI

20,284 views • 2 years ago •via X (Twitter)

3 Comments

Xavier Manuel Mountain's profile picture
Xavier Manuel Mountain2 years ago

Seriously dope! It’s about to be the digital renaissance in music real quick.

SteveRatatata's profile picture
SteveRatatata2 years ago

Thanks for the great work! Are there any plans to release the code?

m҉a҉r҉k҉'s profile picture
m҉a҉r҉k҉2 years ago

this is amazing!

Related Videos

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 views • 2 years ago

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,425 views • 1 year ago

*New Paper on AI & Democracy* Imagine two approaches to democracy. The one we have today, where citizens choose a professional politician to represent them and others. Or an augmented form of democracy, where each citizen controls a personalized AI that helps them participate in thousands of nuanced decisions. This second approach is the idea of Augmented Democracy I introduced six years ago at TED. In our latest paper we explore a simplified version of Augmented Democracy by combining off-the-shelf LLMs, such as ChatGPT, with data collected using a collaborative government program builder. This was an online game where people build a personalized government program using proposals extracted from the programs of the candidates of the 2022 presidential election in Brazil. So how accurate are these augmented forms of democracy? Imagine a user who gave us 40 answers. We can use the first 20 to fine-tune a model that we can test using the 20 answers the model didn’t see. We can then compare the accuracy of these predictions with the ones obtained by a “bundle” rule, which assumes that users that self-reported to be from the left or right always chose the proposals from the candidate that shares their political identity. This showed us that LLMs were more accurate at predicting policy preferences than the bundle rule, meaning that the preferences captured in the participation data were more nuanced than a left-right axis, and that the LLMs can capture some of that nuance. Also, the LLMs can choose among policies coming from the same candidate, which is something that we cannot do using a bundle rule. But can these LLMs help us complete the aggregate preferences of the population? Direct or unbundled forms of participation can result in incomplete data when people answer only a fraction of all questions. In our paper, we simulate this incompleteness by sampling the full dataset. We ask how close we can get to the full dataset by using a random sample, or a random sample augmented by predictions made by these LLMs. Overall, we find that LLM-augmented data gets much closer to the full dataset than a pure random sample. These results do not mean that augmented democracy technology is ready, but they means we are in a much better place to continue exploring this idea than six years ago. This paper was a collaborative effort with Jairo Gudino, PhD student at CCL at the University of Toulouse Capitole and Umberto Grandi from IRIT also at the University of Toulouse Capitole. We hope you find these results insightful!

César A. Hidalgo

26,915 views • 1 year ago

2 years ago, Riffusion was one of the first music generating AI models, by finetuning Stable Diffusion literally on images of waveforms, and then decoding those as audio again I made a song for my gf with it called "Bunny Trouble, Trouble Bubble" about a year ago, extremely catchy and we loved it But it was capped to 12 seconds and I could only imagine how it'd be as a full song Now Suno is finally good enough and I used it as an input so I could extend it from a 12 second clip to 3 minute song, it's essentially the audio version of img2img It took a few times extending it and half the times the output was very bad but the other half it was really great The quality is still a bit low because Riffusion's quality itself was quite low, so that's not Suno's fault, it just extends the low quality into a full song, what I need is an audio upscaler to fix that The fun thing about this, every year music generating AI will be better and I can try inputting the song again to make it better and better If you wanna do the same with your short clip: 1) go to Suno 2) click Create 3) click Custom 4) Upload Audio, upload your short clip here Your existing clip now gets added to the listing on the right (bit confusing) 5) Now hover over it and tap extend 6) More confusingly it will now only play the extended part not the full song 7) To get the full song, click the 3 vertical dots ... -> Create -> Get Whole Song 8) That will stitch your original input clip and the extended clip into one Now you'll probably have a 1 minute clip, so 9) Extend that WHOLE SONG clip again and do the same stitch thing (Get Whole Song) again Depending on the clips you like and not, choose the ones you wanna continue with by stitching them By now you'll have about a 3 minute song!

@levelsio

55,113 views • 1 year ago

We are thrilled to see that our recent Nature Biotechnology paper just became the most accessed publication in the journal within only three months (>225K accesses 🤩)—a fantastic achievement reflecting the power of teamwork! This paper has 7 first authors: two computer scientists, two biologists, and three chemists working together. It perfectly illustrates something I’ve learned deeply over the past 10 years leading my group: great science is mostly the synergy of diverse minds and skills. Building a successful team is like assembling Lego bricks—each person complements the others, fits precisely, and collectively forms something much greater than any single piece. Over the years, we’ve made mistakes and learned valuable lessons, particularly in finding team members whose values and work styles align with our team culture. Now, we’re more intentional about ensuring a mutual fit, making the collaboration enjoyable, reciprocal, and highly productive. This paper is a perfect example: by combining high-resolution imaging and AI, our interdisciplinary team enabled unprecedented visualization and therapeutic development at the single-cell level throughout entire organisms. Proud of the team effort behind this impactful research and deeply grateful to each member who made it possible! What are your experiences working in a highly collaborative team vs. better focusing on your own project? The latter is totally fine and works perfectly for many academic labs.

Ali Max Erturk

19,678 views • 1 year ago