正在加载视频...

视频加载失败

Why choose between autoregressive (fast, streaming) and bidirectional (high-quality) video diffusion when a single model can do both even better? Excited to share Flex-Forcing, accepted as spotlight at #ICML2026! 📍Say hi to Chao Liu at our poster: July 9 | 2:30 - 4:15 PM KST | Hall A #3909...

10,342 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,850 次观看 • 1 年前

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 次观看 • 1 年前

BREAKING: GPT-5.5 "Spud" is out and it is a BEAST We've been testing it Every 📧 for the last 3 weeks on everything from coding, to writing, to knowledge work. Here's our day 0 vibe check: - It's a step change in coding AND it's easy to talk to. It's fast and friendly and quickly became my daily driver. But it's also a coding powerhouse—a really rare combination. - It scored 62/100 on our Senior Engineer benchmark. Opus 4.7 scored only a 33/100. (But GPT-5.5 performed best when using an Opus 4.7 plan). Naveen Naidu used over 900 million tokens during testing—and it let him ship production features for Monologue at both high speed and quality. - It has serious conceptual clarity. It can hold a complex plan in its head over hours of work, without getting distracted by existing code. This makes it the first model that we've tested that can perform well on complex refactors requiring deleting and reimagining an substantial existing codebase. - It's a very good writer. This is the first OpenAI model in about a year that got our writers Every 📧 to switch away from Claude. 5.5 has Katie Parrott's seal of approval—not an easy task. Its writing feels more organic and it's better at mimicking a writing style without going overboard. - It's great for agentic knowledge-work. This is the first OpenAI model that manages to be both a stellar senior engineer AND that can be used for everything from spreadsheets to research. It's crazy fast, and it's amazing inside of the Codex desktop app, and got much of our team to switch away from Claude Code and Cowork during the testing period. However, it's not a perfect model. - 5.5 still loses to Opus 4.7 on plan quality. It's plans are extremely readable but Opus has better attention to detail and sharper insight. - 5.5 still loses to Opus 4.7 by a bit on front-end and full-stack product work. Kieran Klaassen found that it wasn't quite as good when full-stack thinking and design are involved. And it's not great writing Ruby. - 5.5 is a great vibe coder but if you're vibe coding without a plan it's worse than Opus. Mike Taylor found that Opus is better at reading in between the lines on underspecified vibe-coding tasks. Overall GPT-5.5 is a massive achievement from OpenAI and it deserves a serious look as your daily driver. Read our full vibe check on Every 📧 here:

Dan Shipper 📧

130,382 次观看 • 4 个月前

We made a thing! Very happy to announce sqlcoder-pro and the Defog Alignment Platform. Available to use immediately without a wait-list, weights will be open-sourced very soon. The video does a quick show and tell comparison against ChatGPT (with gpt-4o). Read on for more details! TLDR 💪 equal (or better) performance on text-to-SQL as the most capable Claude-3.5 or GPT-4 models 🤝 You can use it today on a free plan/free trial, without a waitlist 🪽 self-hostable on a single RTX4090, with 2 second median generation times for SQL queries 🔁 exactly the same output every time, give the same prompt 👨🏻‍🏫 teachable and steerable: show the model what you want it to do 🛞 debuggable – you can understand WTF is going on inside the model, instead of treating it like a black box Let's dig into each of these one-by-one! Performance SQLCoder-8b-pro significantly exceeds the performance of our previous sqlcoder-8b model on Postgres text-to-SQL (from 88.2% to 90.2% accuracy - gpt-4o is at 87.6%, for reference). It is also better at following instructions. This was done via self-merges, hand crafted fine-tuning data, and adapting the training data to fit our tokenizer. Cost You can host this on the model on a single $3,500 RTX4090, and support ~5 requests/second via VLLM. If you're looking to host on the cloud instead, you can run it on a single L4 GPU that costs $300/mo on GCP Repeatability We have a dense 8b model with no MoE shenanigans. For the same prompt with temperature=0, you'll always get the same answer – which is critical in BI. Teachable In our alignment and feedback modes, you can give the model feedback on how it answered certain questions, and it will automatically adapt to the feedback. Debuggable You can use logprobs and attention scores to determine where, exactly is the model paying attention to inside a prompt + what it's getting confused by when generating outputs. Available today You can use Defog on the cloud today by going to docs[dot]defog[dot]ai, and getting an API key. Excited to hear what you think!

Rishabh Srivastava

13,465 次观看 • 2 年前

At the News18 Town Hall, I addressed a question often asked by critics: Is the #DravidianModel of high industrial growth and high welfarism sustainable? For us, the answer is not just "yes" it is the only way to build a "Good Society." As I shared, the state has three primary responsibilities: improving the quality of life through public goods, creating an environment where entrepreneurs and innovators can thrive, and keeping the gap between the rich and the poor as low as possible to ensure social cohesion. We have always prioritized the social aspects of governance, nutrition, health, and education, as the essential foundation for all growth. This is not just a welfare choice; it is a predicate for economic success. By building high-quality human capital, we ensure that global investment and jobs find us. The results speak for themselves: an 11.19% growth rate and a 700% increase in electronics exports in just the last few years. When critics highlight "better" fiscal metrics in other states, the trade-off is clear. I will always choose a path that invests in the lives and dignity of our citizens, even if it requires borrowing within manageable limits, rather than accepting models where massive capital expenditure coexists with high levels of multi-dimensional poverty. We do not measure success by the creation of a few super-rich billionaires while the average resident lags behind; we measure it by the quality of life and the capacity to aspire for every single citizen. Watch the video here: [English W\ Tamil CC]

Dr P Thiaga Rajan (PTR)

15,792 次观看 • 6 个月前

April 26 It's critical that you have an accurate assessment of your own capabilities, both good and bad. This is another of the many reasons why competition is so important - because it gives you that assessment with no bullshit, sugar coating, or 30 caliber pencils. The overwhelming majority of people who don't compete think they are way better with a gun than they are - and this is especially true among those who carry as part of their profession. My opinion is that if you carry a gun you owe it to yourself and everyone around you to be extremely good with it. Not just better than the average gun owner, but excellent. Most people who carry do not share this opinion, at least in a way that makes them do anything about it. All the information required is readily available at your finger tips. I realize it takes times to build a house - you don't have to be great today, but actively moving in that direction every day is important. And a lot of people who preach training, don't do it nearly as often or intense as they should. I'm not even thinking of anyone specifically when I say that. Everyone reading this knows if it applies to them or not. I have been just as guilty of it as anyone. Warm Up - Targets: Open (10) As many sets as required until thoroughly warm - 10 reloads freestyle, 10 reloads alternating to strong/weak hand. Primary - Targets: Open (15), Partial (15), Open (35) 20 minutes - Match Mode - Draw and engage T1-3 with 2, reload and engage T1-3 with 2. Alter engagement order per rep.

Wash

11,047 次观看 • 4 个月前

Player Profile: Khanyisa Mayo[27]📝 Position: RW/RF/CF Foot: Left Club: Kaizer Chiefs FC Kaizer Chiefs has made another bold move in adding quality to the teams attack. Khanyisa Mayo provides immediate and long-term attacking solutions for the club on the right. 1. Technical Qualities Mayo is a skillful, left-footed attacker with a profile built for modern, high-intensity attacking football. ■ A very direct winger, always looking to advance play instead of recycling possession. ■ Excellent at carrying the ball through pressure, breaking lines, and forcing defenders into retreat. ■ His dribbling and progressive runs add verticality to the team’s attacking structure. Shooting and Passing: ■ Strong shooting range when cutting inside from the right. He provides that unpredictability to the attack with his quality to strike the ball from range. Consistently attempts to feed the box: goals, assists, and key passes. A finisher and creator between the lines. Positional Versatility: ■ In a 4-3-3/4-2-3-1, his best role is as the right forward or right winger, cutting inside to shoot or combine. He can also be effective stretching play and still carry goal threat and creativity. ■ As a No. 9: While not his best role, he is capable of leading the line. He uses his pace to run in behind. He can compete but lacks the physicality to hold off defenders, though his finishing is good. ■ As a Second Striker[442]: A role that plays better to his strengths than as a 9. He is comfortable operating between the lines, combining with midfielders and a traditional box striker. His mobility, creativity, and shooting ability allow him to function as the link player, collecting the ball, supporting the primary striker, and creating or finishing opportunities. This ability to operate both wide and centrally makes him tactically flexible, giving Chiefs options in various attacking structures. 2. Physical Attributes ■ Pace & Explosiveness: Quick acceleration makes him a constant outlet for balls in behind. ■ Strength & Balance: He is strong enough to hold off fullbacks, sustaining attacking sequences under pressure ■ High Intensity: Matches the tempo of Chiefs’ pressing and counter-pressing game. We do most actions with a lot of intensity, and his explosive nature will be very much welcomed in our attack. 3. Tactical Fit at Chiefs Under Coach Nabi, Chiefs are shaping into a side that plays with directness, intensity, and aggression across all phases of play: ■ Pressing & Counter-Pressing: The team applies immediate pressure after losing the ball. Mayo’s speed, defensive work rate, and forward momentum align seamlessly with this system. ■ High-Tempo Attacking: Chiefs don’t rely on slow possession. Attacks are vertical, sharp, and quick. Mayo’s instinct to drive forward, dribble at defenders, and attack space makes him an ideal fit. ■ Attacking Personality: His ability to provide not just progression but also the final goal, final pass, and decisive action elevates Chiefs’ attacking efficiency. Mayo’s game is inherently aligned with protagonist football: high tempo, forward intent, and productivity. 4. Experiential Value ■ Continental Experience: Spent a season in Algeria with CR Belouizdad, scoring 6 goals from wide positions. Notable not just for the numbers but for adapting to a challenging cultural and tactical environment. ■ Tactical Growth: North African teams are disciplined and organized, with compact defenses and structured pressing. Mayo sharpened his decision-making, spatial awareness, and ability to operate against tight blocks, preparing him for high-level CAF competitions. ■ Mentality: At 27, he blends maturity with hunger. His willingness to take risks reflects confidence and attacking intent. Yet to reach his peak, now is the perfect time to step up, showing the quality glimpsed at Cape Town City. Mayo joins Chiefs as both a system player and a game-changer. A top signing by the Glamour Boys! 📝

El Capitano⚪

89,575 次观看 • 11 个月前

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

17,399 次观看 • 18 天前

An automated wallet earned $416K by betting on both sides. Every trade. I found Account88888 with 98% win rate. 9,604 predictions. Expected the usual directional play. ⭢ Account: Then I saw the position breakdown. On 97.7% of markets he holds UP and DOWN simultaneously. Wait. That should not work. If you own both outcomes, you break even. Basic math. Except he does not break even. He made $416K in four months. Here is the trick. When BTC moves fast, Polymarket lags spot by 30 to 90 seconds. Order books thin out. Prices stick. UP sits at 48¢. DOWN sits at 46¢. Combined: 94 cents. But one pays $1 at close. Always. He buys both during the lag. Before the market reprices. Locks 6¢ profit before the window even resolves. I looked at one 15-minute window. 504 trades. Three minutes of execution. Starts when UP=47¢, DOWN=54¢. Market has not figured out direction yet. Spreads are wide. BTC starts climbing. He keeps buying. UP climbs to 92¢. DOWN drops to 9¢. Market reprices in real time. He adjusts both legs. Not static arbitrage. Dynamic harvesting. When volatility hits, spreads blow out to 8-15¢ between outcomes. Retail guesses direction. Bots have not arrived yet. He enters both. Not at once. As the window moves. One leg at 48¢. Other at 46¢. Combined under $1. Hold time: 8-12 minutes average. Market closes. One side resolves. Profit locks. Used to be JaneStreetIndia. Too loud. Same wallet. Just quieter now. Everyone else picks a direction and hopes. This bot made hoping irrelevant.

Blaze

60,373 次观看 • 7 个月前