Loading video...

Video Failed to Load

Go Home

We release Diamond Maps💎 unlocking accurate and efficient guidance for diffusion models. Our experiments show that our methods scale incredibly well. Excited to see what people will build with this! Accurate guidance has been a notoriously hard problem, but in this work, we’re bringing TWO (!) solutions to the...

60,526 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

We’re excited to introduce ShinkaEvolve: An open-source framework that evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:

Sakana AI

360,273 views • 11 months ago

And just like that… ‘HIT ME HARD AND SOFT: The Tour’ has officially come to a close after 106 unforgettable shows. We’re feeling overwhelmed with gratitude, replaying a million memories in our heads, and trying to find the words to express how much this entire tour has meant to us. First, to every single person who ever sent us a concert video, a clip, a moment, or a tiny piece of your experience: thank you. You helped us bring this tour to people who couldn’t be there in person, and because of you, thousands of fans around the world got to feel a little closer, a little more connected, and a little more included. We truly couldn’t have done any of this without your generosity and love. To Billie and her incredible team: thank you for crafting a show that felt like magic every single night. The passion, artistry, and heart you put into this tour truly touched people in ways that will stay with them forever. Watching this tour unfold city by city has been an honor, and we’re endlessly grateful we were able to see it in person. And lastly, but so deeply, to all of YOU. This community has become more than we ever imagined it could be. Your support, kindness, excitement, and trust have turned this into something truly beautiful. Even during the nights we barely slept, posting videos until the sun came up, we never once forgot how lucky we are to share this space with all of you. Being able to do this for an entire tour, and to feel your energy every step of the way, has meant everything to us! This page exists exactly because of that. We don’t know exactly what the future holds, but we can promise you this: we’re not going anywhere. As long as Billie keeps touring and honestly, even beyond that we hope to be right here with you, creating, sharing, and celebrating together. We can’t wait for the next chapter, the next show, the next era. All of it. And we hope you’ll stay with us for many more tours to come. Thank you for the memories. Thank you for the love. Thank you for making this community feel like home. Sending each and every one of you so much love, always.

Billie Eilish Tours

10,948 views • 8 months ago

I'm proud to share that Glean has surpassed $300M ARR, just five months after crossing $200M and growing ~3x over the past 15 months. This is an exciting milestone for Glean, and it's a signal about where the enterprise AI market is heading. We’ve long believed the real challenge in enterprise AI is not access to models. It is grounding AI in how a company actually works: its people, knowledge, workflows, permissions, and systems. That’s even clearer now. The companies creating real value with AI are not just adopting better models. They are building systems that understand their business well enough to deliver reliable outcomes at scale. That is the real moat, and it is what we’ve been building at Glean: an unrivaled context layer for enterprise AI. That context has to work across the business, not just inside a single team or use case. We see that in how customers adopt Glean: more than 85% use it across five or more job functions. It also has to meet the security and governance demands of complex enterprises. We see that in who is choosing Glean: our Fortune 500 customer count nearly doubled year over year. And it has to make economic sense as usage grows. In our recent benchmark with Claude Cowork, Glean was preferred roughly 2.5x as often as off-the-shelf MCP tools and used 30% fewer tokens on average. Better context improves both quality and efficiency. I enjoyed talking with CNBC's Deirdre Bosa about this broader shift. In enterprise AI, the winners will not be defined by better models alone. They will be defined by who builds the strongest foundation for enterprise context. Thank you to our customers, partners, and team for helping us build the future of enterprise AI.

Arvind Jain

280,790 views • 2 months ago

The Sabotaging Practice of Over Supply and Sameness in the NFT Space. The current zeitgeist of the NFT space is that the same artists are doing the same kind of work five times a year, with project after project leaving a trail of disappointment and discontent among collectors and all of us watching in disbelief as huge resources are extracted from the space over work that feels like it could be left as an "artist study." I understand that you can do what you want with your money as collectors, but we are killing the whole space with this incestuous practice. No artist is that prolific to be able to do 5 collections of 100+ pieces each every year and actually deliver innovation and some kind of creative evolution. Of course, they can pretend play that the work has something new, but there is no precedent nor proof that that has ever happened in the speed that it happens in the NFT space. Again, people are free to through away their resources on whatever they want but with this way of doing things, we more and more are going to start seeing the consequences. Oh! There are consequences? Yes. Maybe unintended, but there are. Let's see. Let's start with the loss of belief in the NFT space as somewhere where emerging artists can come and find support for their experiments. Why even bother to bring experiments, innovation, and new ways to think of art on the blockchain if the same people have all the collectors hypnotized with their magical flutes? Why even try to come to a space where taking risks and challenging the status quo (the mission of art!!!) is overlooked? This makes the NFT space a social club and not a space for art. I guess it is fine, but IMO it is a recipe for disaster. New collectors stay away because the art will slowly but surely become stale and un-challenging. Why even bother to come and see what is happening here if you can't, as a collector, see new weird and up-and-coming artists? The amount of noise emitted by the same artists doing the same art over and over, drowns out any new voices. Again. A recipe for disaster. The NFT space is becoming a space of disappointment and doubt. We think that collections going to zero one after the other, over and over, is not damaging? I feel we are kidding ourselves. Disappointment piles up, and again, the people who will hurt are the emerging artists, the new blood, the ones who are willing to risk the most and, in return, put fire in this cold space of sameness. I love this space—don't get me wrong—it has changed my life, and I believe it has a ton of potential, but things need to change for it to become a beacon of light in art. But we need to support new voices. We need to support new ideas. The challenge is huge. I hope to contribute all I can to this change. I hope more and more see how exciting it is to go out and try to discover what else is out there and move this space forward. But again, I understand the leaps of faith needed, but if there is a space that is based on that, it's the NFT space...so there is hope. We will see. 📺by Boldtron

alejandro cartagena

98,261 views • 2 years ago

Most recent diffusion language model research (that I’ve seen) seems to be using masking as the noising process. It looks like, however, most closed-source models (Google Gemini Diffusion and possibly Inception Labs’ Mercury) use a different noising process, where instead of masking tokens, they replace them with different tokens (either with a random token or a semantically similar token). I wondered how they were getting such high throughput with the latter noising process, since I believed that optimizing inference with KVCache approximation would be more difficult (for various reasons). I visualized this noising process with tiny-diffusion and compared it to normal unmasking, and was very surprised to see how fast the generation “settles” into a reasonable output, and then only slightly refines afterwards, requiring much fewer steps in total. Unmasking (where tokens are never remasked, the typical implementation) is inherently limited in generation speed by the fact that an increase in tokens decoded per step leads to more errors due to the mismatch between individual and marginal token probability distributions we sample from. The token replacement noising process seems to have a much different set of characteristics. Because we sample each token per step, every token makes “progress” towards the final output each iteration (in addition to *potentially* giving other tokens more information in future steps). Generally, masking has outperformed other noising processes, which is probably why most research focused on it (using smaller models). But the paper referred to in the retweet shows that random replacement as a noising process may scale better as model size increases. Big labs might have noticed these results much earlier (due to having drastically more training resources and being able to test larger models), which may explain the discrepancy in the choice of noising process. I’m gonna test this with larger models, since tiny-diffusion only has 10M parameters.

nathan (in sf)

40,440 views • 7 months ago

It’s more than a little daunting to set out to expand and improve the identity system for a company and brand like Stripe. But we knew we had to — the existing one had served us well, but wasn’t up to the task anymore. Our brand system required new and improved tools to scale with our ever growing audiences, new products, global footprint, and more. This update introduces material improvements to infographics, advertising, type styles, and more. While the wordmark remains unchanged, we’re using the dot of the ‘i’ (called the “tittle”), a parallelogram pointing up and to the right, to serve as our identifying symbol. We’re also using it as an ever evolving storytelling device to use when talking about our many great users (you can see the latest brand campaign in SF and NYC doing just that). Anyone who has ever worked on the refresh and expansion of an existing system for a large company knows that it is no small endeavor. Crafting impactful solutions, building alignment, creating extensible guidelines, building toolkits, and orchestrating rollout requires a ton of resilience. Here’s to the team that continually inspires me with their dedication, rigor, taste, and exceptional vibes. Great work and thank you to the Brand Studio folks, and of course our many many amazing and invaluable friends and collaborators across the company who all helped shape the work. And a special thank you to a handful of creative agencies that helped us along the way.

Michael Jeter

11,567 views • 10 months ago

Let's talk about agentic product design. Every company has its own design process. What has always worked for me is spending long studio hours with our product team, dissecting things into pieces and putting them back together. In those sessions we look at value, usability, simplicity, aesthetics, behavior, storytelling, generics, and emotional mapping. I've been crafting products this way for as long as I can remember. Product work at Lemonade isn't for the faint of heart. This obsession over every detail is hard work, but I believe it yields better results and builds stronger talent. One of the things I love about our design and product team is how this process became a second nature to them. Feedback is fast, professional, and tension free. But in our latest session, something was different. One of our designers used Figma and Cursor to build a mockup that was so advanced, it was almost ready to be shipped. It was an incredible glimpse into a world where a single designer working on top of modern low code infrastructure will be able to launch production grade experiences for products with millions of customers, and with LoCo, I expect this to become a reality at Lemonade in just a few quarters. But there's a problem to watch out for. An interesting phenomenon I've noticed over the years is that the higher the fidelity of the work being reviewed, the more defensive people become. When someone shows up with something polished, they tend to resist feedback. They've already fallen in love with what they built, and it's hard for them to accept rejection. Radical candor feedback works best at an early stage of the project, before people get attached and feel the need to defend their work. This session was no exception. Because the work was so advanced, the review became binary, and its maker became defensive. Happily, we all caught ourselves in time to acknowledge this new dynamic and started figuring out how to go back to obsessing about every corner radius, shade of white, and word. When reviewing agentically coded designs, we'll try having our designers bring in more than one option, as well as the open Cursor project so we can make changes in real time if needed. We'll see how it goes, and if this is of interest, I'll update what we learn.

Shai Wininger

17,558 views • 9 months ago

It’s hard to believe we're on the final day of our campaign to save the farm and keep our special Esther Magic in Campbellville with Operation Angels. In just two short months, we've raised over half a million dollars to help bring this dream to life. Although we haven't hit our goal yet, what we've achieved is incredible. With over 6,100 contributors from 39 countries totaling over $560,000, this is a global affair in typical Esther fashion. This project is designed to generate revenue to support sanctuaries and rescue organizations and serve as a model for revenue-generating programs at existing facilities. What sets us apart is our approach to drawing revenue from a new demographic to ease the pressure on current donors and an educational program to better equip animal organizations for the challenges they face. This project represents everything Esther stood for: kindness, compassion, inclusion, and support. I can't think of a better way to honor her and the community of like-minded people who make our little corner of the internet unique. The campaign will remain active until 4 am ET on August 1st, so we have less than 17 hours to give it our best and finish strong. To anyone who hasn't contributed, I hope you’ll consider doing so today because we could certainly use your help. We may not reach our goal, and that’s okay. Don’t be discouraged. It just means we keep going because come hell or high water, we’re going back to Campbellville. Thank you to everyone who has contributed so far and to everyone who has liked, shared, and commented on campaign posts these last two months. You have all played a critical role, and I am eternally grateful for your support and encouragement. It is an honor and a privilege to be part of this project with all of you. It represents the end of an era and the start of a new and exciting one. We’re going to do amazing things together, and I can't wait to get back to work. See you all live tonight for our final campaign Garden Chat at 6:30 PM ET. xoxo Steve.

Esther TheWonder Pig

11,759 views • 2 years ago