How much better are the internal, unreleased models at... frontier labs like Google, OpenAI, and Anthropic? We got a glimpse exactly one year ago today, when Google accidentally leaked the “Kingfall” model "Kingfall" was likely an unreleased Gemini 2.5 Ultra-sized model. It was available in AI Studio for only a few minutes but remained accessible through the API for several days At the time, "Kingfall" appeared to be significantly better than Gemini 2.5 Pro at both code generation and creative writing In a recent interview, Sundar Pichai mentioned that Google could have made a better, Ultra-sized Gemini Omni model, but would have had trouble serving it The infrastructure required to serve Ultra-sized models at scale is likely why Google never publicly released models like “Kingfall”show more

AiBattle
12,010 views • 1 month ago
1/ Gemini 2.5 is here, and it’s our most... intelligent AI model ever. Our first 2.5 model, Gemini 2.5 Pro Experimental is a state-of-the-art thinking model, leading in a wide range of benchmarks – with impressive improvements in enhanced reasoning and coding and now #1 on Arena by a significant margin. With a model this intelligent, we wanted to get it to people as quickly as possible. Find it on Google AI Studio and in the Google Gemini for Gemini Advanced users now – and in Vertex in the coming weeks. This is the start of a new era of thinking models – and we can’t wait to see where things go from here.show more

Sundar Pichai
864,374 views • 1 year ago
GOOGLE 🔥: An upcoming Gemini Omni video model from... Google is expected to be much more advanced in video editing, capable of completing tasks like removing watermarks, replacing objects in the video, and more. It is also likely that Google will release 2 versions of this model, including a Pro variant. And I assume what we see isn't Pro? Anime sample 👀show more

🚨 AI News | TestingCatalog
179,313 views • 2 months ago
Announcing Personal Intelligence, a more personalized Google Gemini designed... just for you. How it works: — Customized: With your permission, it reasons across your Gmail, YouTube, Google Photos, and Search apps to share hyper-relevant and context-aware responses — Secure: If enabled, you control which Google apps to connect to. This setting is off by default — Useful: From travel plans based on your Google Photos to gym recommendations based on goals you’ve shared with Gemini, you get help tailored to your world Personal Intelligence in beta is rolling out to Google AI Pro and AI Ultra subscribers in the U.S., with expansions to the free tier, more countries, and AI Mode in Search to come. Take a look at the Gemini app's personalized assistance in the clip below, then let us know what you would use it for!show more

Google AI
320,378 views • 6 months ago
🚨BREAKING: Google just merged Gemini and NotebookLM into one... unified workspace and it changes everything about how you use AI for deep work. It's called Notebooks in Gemini and it's the personal knowledge base that power users have been begging for. You create a notebook for a project, drop in your files, PDFs, and documents, give Gemini custom instructions, and every chat you have stays organized in one place. No more hunting through old conversations. No more re-uploading the same files every session. The wildest part is the sync. Anything you add in Gemini automatically appears in NotebookLM. Anything you add in NotebookLM automatically appears in Gemini. One source of truth. Two powerful apps. Zero friction switching between them. So you can start a research notebook in Gemini, ask it questions all week, then flip to NotebookLM to generate a Cinematic Video Overview from the same material. Next morning, open Gemini and ask it to write a full report on exactly what you just watched. That workflow used to take three apps and a lot of copy-pasting. Now it's one notebook. Rolling out this week to Google AI Ultra, Pro, and Plus subscribers on web. Mobile and free users coming soon. What do you think?show more

Mayank Vora
136,639 views • 3 months ago
AI Is Moving Beyond “Generating Videos” — Toward “Generating... Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:show more

雪踏乌云
112,114 views • 7 days ago
Gemini Omni doesn’t just allow you to render text... more accurately — but to create it in sync with your visuals. 🎥 Choose your type, placement, animation, exposure and more. Prompt for this video: Word by word, one word on the screen at a time: did, you, know, that, this, model, can, do, pretty, good, text!? each word appears with a different animated style, perfect pacing to a rhythm, sizzle reel.show more

152,052 views • 1 month ago
Samsung Galaxy S25 Ultra has completely targeted its competitors... at the iPhone, and no longer competes with Chinese brand phones. This is because young people in South Korea are almost completely occupied by Apple, and Samsung’s strategy is to pull these young people back, so it strives to make Galaxy look like the Apple iPhone. There are several reasons for not competing with Chinese brands: 1. On the surface, although the Chinese Ultra models have powerful cameras and are suitable for those who pursue the ultimate in images, they are only limited to these people. Overall, the sales of Ultra models are not high, and even the sum of all brands of Ultra models cannot be compared with the sales of S24Ultra. 2. Moreover, these models with powerful cameras are relatively thick and heavy, with serious camera bulges, and the design is not perfect. It cannot be perfect. This design may not be suitable for everyone. Samsung will not easily take the risk to adopt this design in the global market. 3. The infinitely enhanced camera configuration will greatly increase the cost, which is difficult for Samsung to accept, and the production of this frequently updated camera may not be able to support Samsung's sales demand of tens of millions. In short, the Chinese brand Ultra model is more like a special non-popular model suitable for geeks. The Galaxy S25 series is defined as a popular model, and it must consider everyone's feelings and sufficient supply. This is why Samsung will not design the S25 Ultra as a Chinese brand Ultra. It remains consistent with the iPhone 16 Pro Max, but subtly. It is always slightly better than it For example, it is a little thinner (8.2mm vs 8.25mm), a little lighter (219g vs 227g), a little narrower bezel, a little stronger performance, a little more camera (retain 3x), a little more ultra-wide-angle pixels (50MP vs 48MP), a little bigger battery, and a little faster charging. Even the most incredible improvement: One UI 7.1 is a little smoother than iOS18 software. This is happening.show more

Ice Universe
347,352 views • 1 year ago
We just dropped a game-changer to slash your Gemini... API costs 🤯 Yesterday we launched Gemini Context Caching, saving you up to 70% on your bill 💸💸 Here's how it works. Gemini let’s you add powerful AI models to create apps like: 📧 Email apps that summarize entire threads for you 💃 Storytelling apps that write the next chapter 💼 HR chatbots that actually understand your company's policies To make these apps, you need to repeatedly send Gemini instructions, which can be pricy. But through the integration of research, software and chips came a breakthrough; 🤖 Context Caching Just tell Google what you want to save money on and we'll take care of the rest. It’s an industry first and it’s going to be a big deal. I'm so proud of the team for always putting developers first 🙌 Cost is always a huge part of the equation and we keep making it better. Try Gemini Context Caching now and tell us what you think! Docs link is below. #AI #LLM #Geminishow more

Liam Bolling
124,644 views • 2 years ago
Most recent diffusion language model research (that I’ve seen)... seems to be using masking as the noising process. It looks like, however, most closed-source models (Google Gemini Diffusion and possibly Inception Labs’ Mercury) use a different noising process, where instead of masking tokens, they replace them with different tokens (either with a random token or a semantically similar token). I wondered how they were getting such high throughput with the latter noising process, since I believed that optimizing inference with KVCache approximation would be more difficult (for various reasons). I visualized this noising process with tiny-diffusion and compared it to normal unmasking, and was very surprised to see how fast the generation “settles” into a reasonable output, and then only slightly refines afterwards, requiring much fewer steps in total. Unmasking (where tokens are never remasked, the typical implementation) is inherently limited in generation speed by the fact that an increase in tokens decoded per step leads to more errors due to the mismatch between individual and marginal token probability distributions we sample from. The token replacement noising process seems to have a much different set of characteristics. Because we sample each token per step, every token makes “progress” towards the final output each iteration (in addition to *potentially* giving other tokens more information in future steps). Generally, masking has outperformed other noising processes, which is probably why most research focused on it (using smaller models). But the paper referred to in the retweet shows that random replacement as a noising process may scale better as model size increases. Big labs might have noticed these results much earlier (due to having drastically more training resources and being able to test larger models), which may explain the discrepancy in the choice of noising process. I’m gonna test this with larger models, since tiny-diffusion only has 10M parameters.show more

nathan (in sf)
40,440 views • 6 months ago
Qualia has been selected for the Google DeepMind Robotics... Program. We train embodied models that put a robot on a real manual task and make it work, on the floor, not in a demo. Foundation models and reasoning are where robotics is heading, and doing that work alongside DeepMind, who are pushing this frontier, is exactly where we want to be. If you are a company looking to see how a new generation of robots can help your manual tasks, contact us at [email protected] More soonshow more

Qualia
87,429 views • 1 month ago
Today we’re introducing Google AI Threat Defense - a... comprehensive AI-powered cybersecurity solution designed to help continuously monitor for and stop AI-powered threats before they can impact your business. Here’s how it works: 1. AI Threat Defense uses our cybersecurity platform Wiz to scan and prioritize what applications and systems have the highest security risk. 2. Gemini and other frontier AI models can then autonomously perform continual deep scanning of your applications - starting with those at the highest risk - to identify security vulnerabilities. 3. The capabilities of CodeMender - a new software repair agent - are then used to verify and accelerate the patching of vulnerabilities. 4. And our Wiz autonomous agents continuously test your systems to find unknown vulnerabilities before adversaries do so that you can remediate them before you are attacked. While other model providers focus on using AI to find and flag vulnerabilities, Google AI Threat Defense actively prioritizes your most critical real-world risks and accelerates their remediation using a variety of models since no single model finds a superset of the vulnerabilities found by all other models.show more

Thomas Kurian
196,916 views • 1 month ago
✨ 3d models are now LIVE on Photo AI... 😊 You can now turn any AI photo you make into a 3d model by pressing [ 📦 Make 3d model ] And then you can view it inside Photo AI or download it as a .GLB 3d model file It's still very early in AI generated 3d model world but it's nice to have this feature working already As always, the models will keep improving, so this feature will keep getting better (like it did with video, it sucked before, now it's getting passable) Next would be nice to switch to .USDZ so you can load it straight into your iPhone with ARKit and put it in your room Available now for everyone on the Premium and Ultra planshow more

@levelsio
112,930 views • 1 year ago
🚨Gemini 3.6 Flash is trash I tested it on... a 3D Golden Gate Bridge, and the results were awful. • I had to re-prompt it three times because it repeatedly ignored the instructions. • First attempt, instead of creating the requested .html file, it first tried to build the experience inside the Gemini app using simulations. • Then second attempt it started placing images from the web into the chat rather than actually producing the file. • Even after getting it to complete the task, the final output was dramatically worse than Gemini 3.1 Pro, which is 5 months old and now not even a top 10 model on leaderboards. This feels like a regression from Gemini 3.5 Flash and honestly, it is one of the weakest models I have tested in the past few months. Has anyone else tested Gemini 3.6 Flash yet, and are you seeing the same thing?show more

Lumina
70,810 views • 2 days ago
Introducing Replit ModelFarm, the fastest and safest way to... build your next Generative AI app. Available for free on Hacker and Pro plans till October 15th. It requires zero setup, zero configuration, and zero API keys. With Replit ModelFarm, you can build a working Gen AI app in as little as 3 lines of code. Get started by installing the Replit AI library in any Python, JavaScript, or TypeScript Repl. The library implements an API for text completion, chat completion, and text embeddings. It supports streaming so your users can see model responses in real-time rather than waiting on a single output. All Hacker and Pro builders will have free access to a selection of Gen AI models offered by Google Cloud Vertex AI through Replit ModelFarm. All models are accessible from the development environment and any deployed app.show more

Replit ⠕
229,285 views • 2 years ago
Something NVIDIA & Google do better than anyone else... is software-hardware-system co-design, and not just optimizing hardware for current model architectures, but predicting future ones. Back in early 2022, when NVIDIA started the design process for NVL72, MoE (Mixture of Experts) models were not yet the standard, and dense models were still dominant for frontier models. However, NVIDIA's strong software-hardware co-design culture enabled them to make a calculated bet that MoEs were the future, and they built NVL72 specifically for best MoE performance per TCO (Total Cost of Ownership). Furthermore, back in 2022, disaggregated prefill and wide expert parallelism (wideEP) MoE inference optimizations hadn't been invented yet, but it turns out that these MoE inference optimizations work best on large-scale systems like NVL72. While most other AI chip companies' in-house AI labs focus on training small 5B models that mainly use data parallelism, NVIDIA and Google's in-house AI labs continuously push the boundaries of model architecture and training recipes, such as NVFP4 training. Just like Super Idol & IShowSpeed, there must be a strong partnership between software engineers and hardware engineers to deliver the best systems that maximize performance per TCO.show more

SemiAnalysis
51,021 views • 8 months ago
Fable 5 comes back!It can now build playable game... prototypes. I think it is actually a signal for where AI coding is going. Making a game is not just “write some code.” Even a small browser game needs: game loop;character movement;collision logic;scoring system;UI states;physics tuning;visual feedback;bug fixing;playtesting This is why game prototyping is a great test for AI models. A model cannot fake it with a pretty answer. Either the game runs, or it does not. What impressed me about Fable 5 is that it is useful for the messy middle: turning an idea into mechanics, turning mechanics into code, debugging broken interactions, and iterating until the prototype feels playable. But here is the practical part: I would not use the strongest model for every step. For game building, I would split the workflow: 1. Fable 5 for game design + architecture 2. a fast coding model for routine implementation 3. a vision-capable model for screenshot/UI feedback 4. a cheaper model for docs, test cases, and small fixes 5. fallback when latency, cost, or output quality becomes a problem That is the real AI coding stack. Not “one magic model does everything.” More like: the right model, for the right task, at the right cost, with fallback when things break. This is why I’ve been looking at ZenMux ZenMux. ZenMux gives developers one gateway to access multiple leading AI models, with OpenAI / Anthropic / Google Vertex compatible APIs, cost tracking, quality benchmarks, auto-routing, and compensation when output quality, latency, or throughput falls short. If AI can now make games, the next question is not just “which model is strongest?” It is:how do we manage the whole model workflow Fable 5 shows the creative ceiling. ZenMux is closer to the infrastructure layer you need when AI coding becomes a real production habit.show more

Rachel🥥
60,942 views • 21 days ago
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

AlphaSignal AI
91,821 views • 3 months ago
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

Jafar Najafov
13,630 views • 3 months ago