Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Learn about Google’s new SOTA image model, Gemini 2.5 Flash, its key capabilities, and what’s next on the roadmap with some of the team behind the model Nicole Brichtova Kaushik Shivakumar Mostafa Dehghani Robert Riachi with Logan Kilpatrick. Timecodes: 0:37 New model introduction 01:21 Demo: Image editing 03:44 Text...

31,150 Aufrufe • vor 1 Jahr •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 Aufrufe • vor 11 Monaten

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

412,880 Aufrufe • vor 10 Monaten

Claude Code is a major (and accidental!) hit for Anthropic that surprised even its creator, Boris Cherny. Claude Code, an Agentic AI coding product that lives in the terminal. Most of the new code at Anthropic is created through it today. And in the last 5 months since it was launched publicly, Claude Code went from $0 to $400M in revenue run rate (as per The Information). 00:00 – Intro 01:15 – Did You Expect Claude Code’s Success? 04:22 – How Claude Code Works and Origins 08:05 – Command Line vs IDE: Why Start Claude Code in the Terminal? 11:31 – The Evolution of Programming: From Punch Cards to Agents 13:20 – Product Follows Model: Simple Interfaces and Fast Evolution 15:17 – Who Is Claude Code For? (Engineers, Designers, PMs & More) 17:46 – What Can Claude Code Actually Do? (Actions & Capabilities) 21:14 – Agentic Actions, Subagents, and Workflows 25:30 – Claude Code’s Awareness, Memory, and Knowledge Sharing 33:28 – Model Context Protocol (MCP) and Customization 35:30 – Safety, Human Oversight, and Enterprise Considerations 38:10 – UX/UI: Making Claude Code Useful and Enjoyable 40:44 – Pricing for Power Users and Subscription Models 43:36 – Real-World Use Cases: Debugging, Testing, and More 46:44 – How Does Claude Code Transform Onboarding? 49:36 – The Future of Coding: Agents, Teams, and Collaboration 54:11 – The AI Coding Wars: Competition & Ecosystem 57:27 – The Future of Coding as a Profession 58:41 – What’s Next for Claude Code

Matt Turck

82,327 Aufrufe • vor 1 Jahr

College sports are going through massive changes, so Matt Brown (publisher, Extra Points & one of the top experts on the business of college athletics), joined the show to break it all down. We discuss: - Superconference realignment - Why the transfer portal is chaos right now - Why most NIL deals have terrible ROI - How “bag man” money moves - Which teams make the most money - Will athletes become employees? - The full history of college athletics - The class action lawsuit happening now (00:00) Meet Matt Brown: Expert in College Sports Business (03:09) The Origins of College Sports (06:31) The Evolution of College Sports Broadcasting (14:53) Title IX and Its Impact on College Athletics (17:53) The 1984 Supreme Court Decision and Its Aftermath (20:03) The SMU Death Penalty Scandal (22:19) Conference Realignment and the BCS Era (28:22) The Rise of Conference Television Networks (30:23) The Arms Race in College Sports Facilities (34:41) The Role of Boosters in College Sports (36:03) Financial Breakdown of Major College Sports Programs (37:04) Understanding Nonprofit Accounting in College Athletics (38:20) Revenue Generation in College Sports (40:34) Athletics as Enrollment Management (42:04) The Flutie Effect and University Applications (44:37) Conference Realignment and Financial Instability (48:58) The O'Bannon Case and Video Game Licensing (53:59) The Northwestern Unionization Attempt (58:19) The Alston Case and Educational Awards (01:02:11) Name, Image, and Likeness (NIL) Marketplaces (01:05:51) The Role of Collectives in College Sports (01:12:08) Dependability of Young Campaign Partners (01:13:03) Transfer Portal and Its Impact (01:15:56) Rise of NIL Agents and Handlers (01:17:40) Economic Incentives and Transfer Market (01:20:37) Challenges in NIL Enforcement (01:22:48) House Settlement and Future Implications (01:25:38) Allocation of NIL Funds by Universities (01:44:26) Potential Super Leagues and Investment Challenges (01:48:07) Concluding Thoughts on College Sports

The Logan Bartlett Show

23,210 Aufrufe • vor 1 Jahr

Everyone is talking about Vibe Coding (Using AI to Create Apps Only using AI) This is the most Comprehensive Guide for Vibe Coding with Cursor (By Far) 250 Minutes, All the vibe code basics of cursor, plus 4 Projects in one video! This is how I, as someone who has never written a line of code, approach building apps (every day). Part 1A Intro to Cursor, Composer, and some basics --------------------- 00:00 Intro 03:41 Downloading Cursor 06:09 What the hell is Composer? 10:47 A Note on Context and Keeping Composer Threads Small 11:38 Simple Desings with Cursor Composer From Blank Project 14:04 Editing a Simple Animation With Cursor Composer 16:35 Setting Up The Voice to Talk to Cursor Composer Whispr Flow 17:54 Lets an Early 2000's Landing Page Part 1B AI Image Generator --------------------- 23:59 Using the GitHub Template to Create a NextJS App 26:43 Template is Open, Let's Edit it 28:55 Drawing Out My Idea With Whimsical 30:11 First Prompt Using Place Holders For Image Generation 32:10 Accept All Vs Save All and Restoring in Composer (Saving your work) 33:54 Adding AI Feature (Brief Teaser, Deep Dive Later) 35:15 What is an API 37:22 Perplexity the best place to learn about API's 40:21 Api keys and running prompt for first AI Feature 42:48 Debugging, Woohoo! Learn to love this :) 43:20 Inspect - Console, In Browser Debugging Hack 48:02 AI Image Generation Works! Lets add more Part 2: Landing Page ---------------------- 51:03 Pause and Reflect, What have we done so far? 53:41 Plan for rest of video 54:34 Ok Let's Talk about (1) Designs 56:19 GitHub is like --sref for those who do image gen 58:20 Starting Cursor project from a GitHub Repo we found on Perplexity 01:00:48 Yolo Mode... Wtf is that? 01:02:38 Inspecting GitHub Repo's Examples, to use in our landing page 01:02:58 The Project We're making - A landing page 01:03:56 Landing Page from Screenshot 01:06:17 Making Changes to Landing Page 01:11:42 Making a more epic section 01:13:42 The Essence of Vibe Coding 01:15:17 Creating Cool Testimonials Section From Screenshot 01:18:18 Deploy to Vercel! But First New Repo on GitHub 01:20:45 Ok it's on GitHub... Now lets do vercel 01:21:17 Untechnical Explanation of what Vercel is Lol 01:24:18 Connecting Custom Domain (Bought on Name Cheap) To Vercel Deployment Part 3: App With Database and Authentication ---------------------- 01:27:59 Recap and Prep For The Bigger Project! 01:35:13 Getting Started from Template (Again) 01:38:52 Setting Up Database and Authentication (Firebase) 01:44:01 Back To Cursor, Let's Set up The Auth in the app 01:48:35 Switching to mermaid because compatibility issues 01:51:13 Using AI (Claude) to Generate Mermaid Diagrams 01:52:19 Adding Docs to Cursor to use AI Features over and over again 01:54:38 Let's Troubleshoot 01:56:10 Adding View Button and EDIT WITH AI 02:01:45 AI Diagram Edit Feature is DOPE 02:03:17 Using Search Feature on Cursor to find text in Codebase 02:05:55 Lets add ability to save these to Database 02:09:33 What does saved to Google Firebase even mean? 02:13:00 We can Export as PDF! 02:15:48 GitHub and Vercel Again! 02:17:27 Vercel with CLI From Cursor 02:20:52 Setting Vercel Domain as an Authorized Domain 02:27:34 How To Learn More

Riley Brown

368,431 Aufrufe • vor 1 Jahr

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,687 Aufrufe • vor 1 Jahr