Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

What started as building a personal taste.md skill for myself, turned into building a pipeline to create any taste as a skill. The most important piece is references. This is where you should spend time. If the references suck, so does the skill. I find that references cropped tightly...

61,019 Aufrufe • vor 4 Monaten •via X (Twitter)

41 Kommentare

Profilbild von Jaytel
Jaytelvor 4 Monaten

Connect your OpenRouter account and try it here or clone the repo and run it locally with your own keys

Profilbild von Druids
Druidsvor 4 Monaten

I do something similar, the difference is I manually annotate the things I like on each reference before feeding to the skill.

Profilbild von Jaytel
Jaytelvor 4 Monaten

Yeah that’d probably help here too

Profilbild von Mete Polat
Mete Polatvor 4 Monaten

Dude this is super interesting, thanks for sharing the breakdown. I’m actually planning to write next weeks’s newsletter issue on the need to start collecting more references outside of text and this fits in nicely. 2 questions for you - 1. do you feel like it actually deducts correctly *why* you like something or what makes it good? You mention the “why the reference is successful” but it doesn’t seem like you actually tell it yourself why it’s significant to you? 2. do you feel like this has the potential to broadly capture your taste or is it more about finding a “local valley” in the model’s style that’s aligned with the particular references you provided (and so it’s more like style presets you can tune and switch between vs a larger overarching taste)

Profilbild von Jaytel
Jaytelvor 4 Monaten

When you ask it to identify why it’s successful and ignore the function of the reference- it is surprisingly close to what I would say about each reference myself. I think you need to do this for each general aesthetic you’re looking for. Not broadly my complete taste as a whole

Profilbild von Mete Polat
Mete Polatvor 4 Monaten

Makes sense. Thanks for elaborating

Profilbild von Farhan
Farhanvor 4 Monaten

this is awesome! thanks for sharing. I tested with one sample and tested on both opus 4.7 and gpt 5.5 and results are amazing. I created reference sets so i can work on targeted taste experiments. 1 - local ui wrapper 2 - reference and results

Profilbild von Raffi
Raffivor 4 Monaten

Another banger

Profilbild von Jaytel
Jaytelvor 4 Monaten

🫶🏼

Profilbild von Pratham
Prathamvor 4 Monaten

that's lowkey genius af, thanks for making this 🫡

Profilbild von isle of echo
isle of echovor 4 Monaten

Super cool! I’m exploring something similar with Claude skills right now! trying to build a system/infrastructure that help anyone transform their taste / aesthetic into something can be understood, executed by AI for future projects and interfaces.

Profilbild von Josiah Gulden
Josiah Guldenvor 4 Monaten

🔥 this is great for static refs, but one of the roadblocks I’ve hit in similar experiments is that motion is a big part of why I like what I like and none of the leading vision models can see video. hi-res, HFR sequences are usually too big for the context window. any tips?

Profilbild von Jaytel
Jaytelvor 4 Monaten

there’s a few skills that help. @emilkowalski has a good one but overall still have to spend time finessing motion in general with the models

Profilbild von Jon Moore
Jon Moorevor 4 Monaten

Better yet, use the Mobbin MCP server and give it thousands of references. Cool idea!

Profilbild von Shane Levine
Shane Levinevor 4 Monaten

this is very cool

Profilbild von Michelle Fitzpatrick
Michelle Fitzpatrickvor 4 Monaten

Curious what differences Opus and GPT notice? Is there much duplication from their analysis ?

Profilbild von Jaytel
Jaytelvor 4 Monaten

They both find different things GPT seems to do better from a traditional design perspective

Profilbild von henry
henryvor 4 Monaten

this is so cool!

Profilbild von Wikman
Wikmanvor 4 Monaten

Super great @Jaytel - love this space. We have done a similar thing for design systems that keeps enterprise styles consistent across. Will try this one 👊

Profilbild von Brotzky
Brotzkyvor 4 Monaten

super interesting stuff

Profilbild von Lucas Bacic
Lucas Bacicvor 4 Monaten

Tried something similar (nowhere near that scale), pushing the model to use semiotics as its lens. Adds real depth beyond pure form.

Profilbild von Hemali Tanna
Hemali Tannavor 4 Monaten

Loved it! Starting with those aesthetics can save so much time and tokens!

Profilbild von J 👋
J 👋vor 4 Monaten

Super cool, will give this a go today.

Profilbild von cam
camvor 4 Monaten

gonna try this 🫡

Profilbild von Cog_
Cog_vor 4 Monaten

Brilliant!

Profilbild von Nadine | HealthTech Designer & Developer
Nadine | HealthTech Designer & Developervor 4 Monaten

This is so smart, great explanation too thank you 🙌

Profilbild von Bruno Pardo
Bruno Pardovor 4 Monaten

Love it, thank you so much for sharing. I will try to connect it directly to the Arena API instead and see if I can make it work that way as well, could be quite powerful!

Profilbild von Fazal
Fazalvor 4 Monaten

Very interesting workflow, and if you are using Pi gotta try

Profilbild von Noah Wainwright
Noah Wainwrightvor 4 Monaten

Incredible

Profilbild von betts
bettsvor 4 Monaten

thank you for your work this is amazing

Profilbild von John Allen
John Allenvor 4 Monaten

this is dope

Profilbild von Rahul Singh Bhadoriya
Rahul Singh Bhadoriyavor 4 Monaten

Super intresting, I was thinking on this problem, loved the same image multiple times model call thing

Profilbild von Emily Ritter
Emily Rittervor 4 Monaten

Really nice design! Well done

Profilbild von Built by APE | Austin Product Engineering
Built by APE | Austin Product Engineeringvor 4 Monaten

Any more effective than “make it sexy?”

Profilbild von Jaytel
Jaytelvor 4 Monaten

skill + “make it sexy” ?

Profilbild von Kriss Patel
Kriss Patelvor 4 Monaten

Interesting, I am doing this the other way around, where I am feeding the model all the principles about different styles and building a skill from it so it can refer to it.

Profilbild von Almeida
Almeidavor 4 Monaten

Nicee!

Profilbild von Arslan Iqbal
Arslan Iqbalvor 4 Monaten

Multi-model analysis + chunking = surprisingly scalable design system.

Profilbild von Shellie Hu
Shellie Huvor 4 Monaten

This is super inspiring! Do you envision the skill evolves as your personal tastes change? How would the files capture that?

Profilbild von Farshad Taheri
Farshad Taherivor 4 Monaten

This is cool. Have you tried fine tune a model for this?

Profilbild von J 👋
J 👋vor 4 Monaten

Hey, @Jaytel been playing w. this a bit. Only one model fired in my trial but still got a result. I made some tweaks that maybe are stearing things in the wrong way. It feels more like forced application of images refs rather than an interpretation of an overall aesthetic 🤔.

Ähnliche Videos

Only if education could be this interactive ❤️‍🔥 I've had a looong wish to build something genuinely useful through vibe coding, and I finally did it. A 3D human anatomy application built with Three.js using GPT 5.6 Sol. It all started with a single design image that I created using GPT Image 2.0. I then used it to generate every 3D organ image, one by one. Next, I converted each of those images into 3D models using Tripo (and no, they didn't sponsor this 😄). After that, I opened Codex, wrote a master prompt based on the design, and gave it the prompt, the design image, and all the 3D models. Codex built the first version beautifully, but there was one big problem. Every single 3D model was nearly 120-150 MB. That obviously wasn't practical for the web and was giving a performance of 16fps. After a few iterations, Codex optimized each model down to roughly 2–5.5 MB while preserving the visual quality, reducing the total asset size from ~900 MB to just 28.6 MB. And each model loads on demand. Along the way, Codex also generated those anatomical illustrations showing where each organ sits in the human body, and even created the interactive hotspot markers that explain different parts of every organ. It handled all of that. The process wasn't exactly one shot, but it also wasn't difficult. You just have to do it step by step. It genuinely felt like building something that could make learning anatomy much more engaging. The inspiration came from Dilum Sanjaya's 3D animal plant cell project. I remember seeing it and thinking, "I want to build something like this one day." And I did it :D Live: Code:

The Bugged Dev

2,115,895 Aufrufe • vor 2 Monaten

All these demo videos make HEAD SWAPPING with Nano Banana look so easy, but then you give it a try and you're like... uh... what? Why didn't that work? Here's what I've found. Nano Banana reads your image, almost literally, so if you write on the image, it reads the text. This is how Higgsfield AI 🧩 has capitalized on the tech: "Write on the image" and give it direction, right? Totally true, but you don't need Higgi to write on your image. Nano Banana will understand your direction regardless of where you write on your image. On one hand, Higgi is really smart, because they're hranessing the tech in a unique way, but the whole "Higgsfield's Banana Placement" is a bit of a misnomer. It's more of a "Banana Placement" and Higgi is just giving you a sort of basic Photoshop-type tool to work with (again, pretty smart), but the real tech is the Banana. 🍌 This is how I head swapped heads in Runway, but Nano Banana maintains the aesthetic qualities of your image almost perfectly, whereas Runway Reference spits out a very Gen-4 looking image. I like using Nano in Freepik (now Magnific), mainly because it's fast and I can get 4 gens at a time, and you need to gen a dozen times of so before you get a winner (most of the time). I was pumped when I saw Freepik introduce the @ reference feature, just like Runway has, but it doesn't seem to work for head swapping. My guess is because that's not really how Nano Banana tech works... ideally. Marco is the person I saw using this "A" and "B" method, back when Nano was on LM Arena, and man-oh-man, it just works... like a charm. You need experiment with how much of the face you blot out, and the angle and facial expression of your new head if you want the blend to be perfect. All of the results in this video are 100% Nano Banana. I did not do any Photoshop work to the images after the fact. I really hope this helps. Let me know if you have any questions. I'm happy to help. And I'll keep posting videos like this if you guys find them useful. Let me know! And if you want more serious, one-on-one AI consultation you can throw something on the books here:

Jordan Daniel Chesney

62,235 Aufrufe • vor 1 Jahr

this video is the CLEAREST explanation of how claude skills + AI agents work and how to use them most people set up an AI agent and wonder why it keeps disappointing them. the context window is everything context is what the model assembles before it takes any action. think of it like everything the agent needs to read before it does anything. the quality of what goes in determines the quality of what comes out. the models are genuinely really good right now. claude and gpt are exceptional. the variable is almost always the context you give them. 1. agent.md files are mostly unnecessary every single line you put in an agent.md file gets added to every single conversation you have with your agent. a 1000 line file is around 7000 tokens burning on every run. the model already knows to use react. it can read your codebase. save the agent.md for proprietary information specific to your company that the model genuinely cannot know on its own. 2. skills are the actual unlock a skill.md file works differently. what loads into context is only the name and description, around 50 tokens. the full instructions only appear when the agent recognizes it needs that skill. so instead of 7000 tokens on every run you have 50. and the agent stays sharp because the context window stays lean. the closer you get to filling the context window the worse the agent performs, same way you perform worse when someone dumps 10 things on you at once. 3. here is how to actually build a skill the right way most people identify a workflow and immediately try to write the skill. what you want to do instead is run the workflow by hand with the agent first. walk it through every single step. tell it what to check, what good looks like, what bad looks like. correct it in real time. once you have had a full successful run from start to finish, tell the agent to review everything it just did and write the skill itself. it writes a better skill than you will because it has the full context of what actually worked in practice not in theory. 4. recursively building skills is how you go from frustrated to reliable when the skill breaks, and it will break, ask the agent exactly why it failed. it will tell you specifically what went wrong. fix it together in that same conversation. then tell it to update the skill file so that failure mode never happens again. ross mike did this five times with his youtube report generator. it now pulls from eight different data sources and runs flawlessly every single time without him touching it. 5. sub agents are something you earn not something you set up on day one start with one agent. build one workflow. turn it into one skill. once that works add another. ross mike has five sub agents now covering marketing, business, personal and more. it took months to get there and every single one exists because a workflow proved it deserved to exist. the people who set up 15 sub agents on day one and wonder why nothing works skipped all the steps that make the thing actually run. 6. your workflow is the thing the model cannot get anywhere else the model has been trained on everything. it knows more than you about most things. what it does not have is your specific process, your taste, your way of doing things. that is what skills capture. that is what makes your agent actually useful versus a generic one. downloading someone else's skill means downloading their context onto your setup and it will not work the way you want it to because it was never built around how you work. this is the clearest explanation of how agents actually work i have heard. Micky runs this stuff every single day and the results show it. full episode is now live on The Startup Ideas Podcast (SIP) 🧃 where you get your pods people charge for this sorta stuff i give away the sauce for free i just want you to win watch

GREG ISENBERG

194,524 Aufrufe • vor 6 Monaten

CLIP by hand ✍️ ~ 13 steps walkthrough below CLIP, Contrastive Language-Image Pre-training, is OpenAI's answer to a question that sounds impossible: how do you put a sentence and a picture in the same space? CLIP shipped when OpenAI was still open, and those embeddings were shared far and wide. Almost every multimodal model you use today descends from them. How does it work? Goal: learn one shared embedding space for text and images. = 1. Given = A mini batch of three text-image pairs. OpenAI trained the original on 400 million. = 2. Text to vectors = Let us look up each word with word2vec. = 3. Image to vectors = We cut each image into two patches and flatten them. Now text and pixels are both just numbers. = 4. The other pairs = Repeat steps 2 and 3 for the rest of the batch. = 5. Encode = Let us push both sides through their encoders, a linear layer and a ReLU. In practice these are transformers, but the shape of the operation is the same. = 6. Mean pooling = We average across the columns, so each image and each sentence collapses to a single vector. = 7. Projection = The text vectors are 3D and the image vectors are 4D, so they cannot be compared at all. A linear layer projects both to 2D. That 2D space is the shared embedding space, and getting here is the whole point of the model. = 8. Prepare for matmul = Let us copy the text vectors down and the transposed image vectors across. = 9. MatMul = We multiply, which takes the dot product of every text vector with every image vector. Each cell is one estimate of how well a sentence matches a picture. = 10. Softmax, e to the power = Raise e to each cell. To keep it hand sized we approximate e with 3. = 11. Softmax, sum = Sum each row for image to text, each column for text to image. = 12. Softmax, normalize = Divide, and out come two similarity matrices, one per direction. = 13. Loss gradients = The targets are identity matrices: a pair that belongs together should score 1, every other cell 0. Subtract the target from the similarity and you have the gradients, in both directions. The takeaway: pairing a picture with a sentence comes down to a single dot product. Everything before step 9 is the work of getting them into one shared space, so that the dot product finally means something. 💾 Save this post!

Tom Yeh

20,929 Aufrufe • vor 2 Monaten

how to use Google's NEW open source Design.md + AI Skills to make your startup look like a $100 million company in 1 hour: 1. Design.md is an open source file from Google that captures the soul of a design. Typography, colors, spacing, all in one markdown file. You attach it to your prompt and your agent builds beautiful things every time. 2. Think of it this way. The HTML is the finished dish. The design.md is the recipe. The skills are the ingredients. Put them together and everything you build looks consistent and professional. 3. Don't create a design system from scratch. Find a brand you love. Linear, Stripe, Vercel, whatever resonates. Study it. Use ChatGPT or Claude to help you extract the design language into your own design.md file. 4. Build skills on top of your design.md. A landing page skill. A mobile app skill. A motion design skill. A slide deck skill. Each one references the same design.md so everything looks like it came from the same designer. 5. The biggest mistake people make: they nail one screen and then everything else looks generic. Design.md solves this. One file keeps every page, every format, every medium consistent. 6. Use it across everything. Your landing page. Your app. Your pitch deck. Your promo videos. Same DNA. Same taste. Same system. That's what separates a startup that looks real from one that looks vibe-coded. 7. Build a second brain for design inspiration. When you see something beautiful in the real world or online, capture it. Save it. When you're building something new, reference it. Taste is developed, not downloaded. 8. It's obvious but the difference between a product people trust and a product people bounce from is how it looks and feels. Design.md gives you that edge. you can watch below shoutout to Meng To for coming on The Startup Ideas Podcast (SIP) 🧃 and walking through his full workflow. if you want to use AI to actually build gorgeous designs, you'll want to use see this. watch

GREG ISENBERG

512,753 Aufrufe • vor 5 Monaten

A transformer can learn not just the outcomes of dynamics, but the operator that executes the rules. To show this we trained a transformer on roughly 0.04% of a discrete rule space - 100 of 262,144 possible rules - and it learned to apply unseen rules from the same rule class. The model does not simply memorize specific rules. It learns the operator that maps a supplied rule plus an initial state, including unseen rules from this class, to the correct next state. This is relevant because it is a shift from “neural networks approximate dynamics” to “neural networks can learn to execute symbolic programs within a defined rule class”. The rule itself is supplied at inference time, as data, and the network has internalized how rules act, not which rules to apply. On previously unseen rules, the model achieves 98.5% perfect one-step forecasts and reconstructs governing rules with up to 96% functional accuracy. Two results make this hold up under scrutiny. First, inductive bias decay. As we scaled training rule diversity, the correlation between functional inference accuracy and distance-from-nearest-training-rule collapsed to R² = 0.00. At the largest tested training-rule diversity, the model’s performance on a new rule shows no measurable dependence on how similar that rule is to anything it was trained on. The bias toward training data (the thing we worry most about in compositional generalization claims) is something we can measure decaying, and we find that at scale it is gone. Second, an identifiability theory. We derive a closed-form expression for the number of rules consistent with a single observation. This reframes the inverse problem: failure to recover ground truth is not necessarily a model defect, but can be correct behavior when the data underdetermine the rule. The model is sampling the equivalence class; and identifiability is governed by coverage, not capacity. The methodological move underneath both results is amortization. Classical work on rule inference (e.g. the Santa Fe EVCA program, evolutionary search over CA rule space) was per-instance: search the rule space for each new system. We replace that with a single forward pass of a transformer trained across many instantiations of the rule class. That is what makes symbolic rule inference scalable as a research direction rather than a curiosity. We show that this works in a tightly constrained domain: binary, deterministic, local cellular automata on small grids. The locality-break experiment shows the model fails sharply when target systems violate its structural priors (which is itself a useful diagnostic, but it bounds the operator class). We don't yet know how this scales to multistate, higher-dimensional, or stochastic CA, or whether it transfers cleanly to non-CA systems whose coarse-grained dynamics admit local surrogates. The identifiability framework - what can be inferred from observation, given a hypothesis class - should transfer wherever finite local rules meet sparse data. The amortization argument transfers wherever per-instance symbolic search has been the bottleneck. Those are the pieces I expect to outlive the cellular automata setting. Led by Jaime Berkovich with Noah David, at LAMM@MIT. Out now in Advanced Science Advanced Portfolio (link to paper & code below).

Markus J. Buehler

39,245 Aufrufe • vor 5 Monaten