Loading video...

Video Failed to Load

Go Home

Today, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewer tokens 💰 ~66% lower cost 📈 ~7% higher accuracy The result: a new Pareto frontier for video understanding across accuracy and cost. Traditional video understanding processes video statically at a...

20,049 views • 2 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

AI has transformed how video is created. We think the next wave is about understanding it. Over the past few years, we've seen remarkable advances in video generation, editing, avatars, and creative tooling. An increasingly important problem is teaching machines to search, analyze, reason over, and extract insight from video - across massive libraries and live streams alike. We're calling this video intelligence, and we're actively looking to back founders building here. We're most excited about companies pushing on the core capabilities: - Video-native models - multimodal embeddings, temporal reasoning, and retrieval built specifically for video rather than adapted from image or text - Real-time and large-scale pipelines - infrastructure for processing, indexing, and querying video at the speed and scale enterprises actually need - Agentic and reasoning layers - systems that don't just retrieve clips but answer questions, surface anomalies, and take action on what they see The models and infrastructure to make this real are appearing to be crossing a capability threshold right now. Multimodal foundation models are maturing, storage costs have collapsed, and enterprises are sitting on years of unstructured video with no way to use it. That infrastructure unlocks a wide range of applications including media and sports workflows, security and physical operations, enterprise knowledge management, advertising analytics, robotics, and consumer products, where video has historically been dark data. If you're building in video intelligence at the model layer, the platform layer, or in a vertical application, we'd love to talk!

Jason Cui

36,205 views • 4 months ago

Gemini-1.5 Pro has its spotlight stolen today, and people are poking fun at Sora vs Google memes. Well, I think it's the biggest boost in LLM capability so far in 2024. v1.5's 10M token context (1) excels at retrieval; (2) generalizes zero-shot to extremely long instructions like full tutorials and codebases; and (3) works across modalities such as text, audio, and video. Here's a stunning example: v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning. I talked about the Myth of Context Length many times before: don't get too excited by claims of 1M or even 1B context tokens. LSTMs already achieved literally infinite context length 25 yrs ago! What truly matters is how well the model actually uses the context to solve real-world problems, and Gemini-1.5 has surpassed the SOTA with flying colors. The paper is also well-written with lots of solid quantitative analysis on in-context memorization and generalization. Paper: “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context” Congrats to Jeff Dean Oriol Vinyals Sundar Pichai and team!

Jim Fan

278,517 views • 2 years ago

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,946 views • 7 days ago

"The future of AI is agentic. That includes browsers!" Imagine having an AI agent in your browser that can help you complete complex tasks, answer your questions, and streamline your workflow. Today I'm thrilled to share a sneak peek at Project Mariner, a cutting-edge research collaboration between Chrome and Google DeepMind, exploring the future of agentic AI within the browser! Building on the power of Gemini 2.0, Mariner envisions AI agents seamlessly guiding users through online tasks, streamlining workflows and enriching browsing experiences. Imagine having an intelligent co-pilot in your browser, anticipating your needs and proactively offering assistance. We're in the early stages of experimentation, focusing on core functionalities like understanding user intent, automating actions, and providing personalized recommendations. This prototype leverages Gemini's advanced natural language understanding and reasoning capabilities to interpret user requests, both typed and spoken. Mariner can then interact with web pages, retrieve information, and even perform actions like filling out forms or navigating to specific sites. For example, a user could simply ask "Find me a job near me," and Mariner would understand the request, navigate to a relevant job search site, and tailor the search based on the user's location and preferences. This is just one example of how we're exploring Gemini 2.0's potential to unlock agentic experiences through a series of prototypes, including: 1. Agents with multimodal reasoning: Project Astra, our research prototype exploring the capabilities of a universal AI assistant, is enhanced by Gemini 2.0. 2. Agents that can help you accomplish complex tasks: Project Mariner itself focuses on the future of human-agent interaction within the browser. 3. Agents for developers: Jules is an experimental AI-powered coding agent that integrates directly into a GitHub workflow. 4. Agents applied across domains: We're exploring agents for navigating video games and even applying Gemini 2.0's spatial reasoning to robotics. We believe that integrating AI agents directly into the browser has the potential to revolutionize how we interact with the web. Project Mariner aims to make browsing more intuitive, efficient, and personalized. By understanding user context and proactively offering assistance, Mariner can simplify complex tasks, save users time, and empower them to achieve more online. This aligns perfectly with the vision of Gemini 2.0 to create more helpful and intuitive AI experiences. We’re currently testing Mariner with a small group of trusted users to gather feedback and refine the user experience. We believe that this technology holds immense potential to transform the way we browse and interact with information online.

Addy Osmani

29,518 views • 1 year ago

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

23,951 views • 6 days ago

Most robotics AI models suffer from the "stop-and-think" problem. They take a static picture, pause to reason, execute an action, and repeat. In the real world, that latency causes spills, collisions, and failed tasks. Google DeepMind just launched Gemini Robotics ER 2: an embodied reasoning model that thinks and acts at the speed of the physical world. Here's why this is a step-change for physical AI engineering: Traditional robotics models rely on static snapshots. But knowing *when* a task is done, such as when to stop pouring coffee into a cup or when a trash bag is securely tied, requires continuous temporal awareness. Gemini Robotics ER 2 integrates directly with the bidirectional streaming Gemini Live API to reason about what comes next while simultaneously executing motor actions. What makes Gemini Robotics ER 2 different: 🎯 91.3% accuracy on live video moment-finding (0.96s mean absolute distance) at 4x the execution speed of frontier models 📈 Continuous progress tracking across 5 completion stages (57.4% accuracy) to self-correct mid-task without restarting 🛠️ Native agentic tool orchestration that commands lower-level VLA models, navigation APIs, and Google Search 🤝 Multi-robot collaboration allowing physically diverse machines (like Apptronik's Apollo 2 humanoid and Franka's FR3 Duo arm) to hand off tasks in shared spaces 🛡️ Built-in physical safety that autonomously halts robots when humans enter a workspace and resumes once clear

Karl Weinmeister

28,313 views • 1 month ago