AshutoshShrivastava's banner
AshutoshShrivastava's profile picture

AshutoshShrivastava

@ai_for_success81,222 subscribers

Post about latest AI news, tools, tutorials and memes.

Shorts

This could become reality one day.

This could become reality one day.

19,139,992 次观看

Google is doing some great AI work in India that is actually helping farmers on the ground and honestly, almost nobody is talking about it. It’s surprising how much we focus on every new model drop, while some of the most meaningful AI work is happening quietly in the real world. Google DeepMind’s AnthroKrishi team built two AI models using satellite imagery: - Agricultural Landscape Understanding (ALU) maps farm boundaries, trees and water bodies, with historical data going back 15 years. - Agricultural Monitoring & Event Detection (AMED) uses multispectral data to monitor crops, sowing and harvesting, identifying 11 major crops. And this is already being used at scale: - 5M+ farmers in Telangana through its agricultural DPI - 140M+ hectares covered by Terrastack - 2.6M hectares of irrigated land in Karnataka - CarbonFarm is using it across 12 countries - FAO is integrating the models into its global agricultural data platform with $2.5M from Google These models are now providing agricultural insights across 6 African countries, and the Agricultural Landscape Understanding layer is one of the most popular layers on Google Earth globally.

Google is doing some great AI work in India that is actually helping farmers on the ground and honestly, almost nobody is talking about it. It’s surprising how much we focus on every new model drop, while some of the most meaningful AI work is happening quietly in the real world. Google DeepMind’s AnthroKrishi team built two AI models using satellite imagery: - Agricultural Landscape Understanding (ALU) maps farm boundaries, trees and water bodies, with historical data going back 15 years. - Agricultural Monitoring & Event Detection (AMED) uses multispectral data to monitor crops, sowing and harvesting, identifying 11 major crops. And this is already being used at scale: - 5M+ farmers in Telangana through its agricultural DPI - 140M+ hectares covered by Terrastack - 2.6M hectares of irrigated land in Karnataka - CarbonFarm is using it across 12 countries - FAO is integrating the models into its global agricultural data platform with $2.5M from Google These models are now providing agricultural insights across 6 African countries, and the Agricultural Landscape Understanding layer is one of the most popular layers on Google Earth globally.

57,874 次观看

My daughter is sleeping on me, my wife is helping our son with his homework, and I’m talking to IRIS through my Bluetooth earbuds while Hermes gets the work done. I built IRIS so I don’t have to be glued to my desk anymore. I can work from another room, from bed, or pretty much anywhere in the house as long as Bluetooth reaches. Honestly, I’m just happy I built something that actually makes my life easier.

My daughter is sleeping on me, my wife is helping our son with his homework, and I’m talking to IRIS through my Bluetooth earbuds while Hermes gets the work done. I built IRIS so I don’t have to be glued to my desk anymore. I can work from another room, from bed, or pretty much anywhere in the house as long as Bluetooth reaches. Honestly, I’m just happy I built something that actually makes my life easier.

114,067 次观看

Google DeepMind is doing some crazy work. This is so good and is going to help so many people. They just launched SL2T (Sign Language to Text), an AI model that translates sign language directly into text. Most spoken language dictation tools rely on audio to text, but sign languages have their own grammars and involve complex 3D movements. Instead of using raw camera feeds or physical gloves, SL2T tracks body landmark coordinates locally with MediaPipe Holistic and translates those spatial points into text. It’s powering sign to text dictation on Pixel 11 in Gboard and Live Transcribe, starting with ASL. Users can sign naturally to search, draft messages, or respond in live conversations instead of typing everything out. This is seriously awesome.

Google DeepMind is doing some crazy work. This is so good and is going to help so many people. They just launched SL2T (Sign Language to Text), an AI model that translates sign language directly into text. Most spoken language dictation tools rely on audio to text, but sign languages have their own grammars and involve complex 3D movements. Instead of using raw camera feeds or physical gloves, SL2T tracks body landmark coordinates locally with MediaPipe Holistic and translates those spatial points into text. It’s powering sign to text dictation on Pixel 11 in Gboard and Live Transcribe, starting with ASL. Users can sign naturally to search, draft messages, or respond in live conversations instead of typing everything out. This is seriously awesome.

27,080 次观看

After Sam Altman, it’s Dario Amodei who’s showing up dead on Google Search.

After Sam Altman, it’s Dario Amodei who’s showing up dead on Google Search.

25,119 次观看

AI ate VFX. Gemini Omni Flash is absolutely 🔥

AI ate VFX. Gemini Omni Flash is absolutely 🔥

89,911 次观看

OpenAI updated the ChatGPT macOS app and literally finished Cursor. ChatGPT for macOS can now work with your coding apps and read content from them. Here’s everything you need to know and how to set it up 👇

OpenAI updated the ChatGPT macOS app and literally finished Cursor. ChatGPT for macOS can now work with your coding apps and read content from them. Here’s everything you need to know and how to set it up 👇

479,221 次观看

ChatGPT Plus users after finding out OpenAI Operator is part of the $200 Pro plan.

ChatGPT Plus users after finding out OpenAI Operator is part of the $200 Pro plan.

432,754 次观看

RIP Privacy! It’s scary to think what could happen if this falls into the wrong hands. This AI tool can pinpoint the exact location based on an image. Check out how it works: 👇

RIP Privacy! It’s scary to think what could happen if this falls into the wrong hands. This AI tool can pinpoint the exact location based on an image. Check out how it works: 👇

407,570 次观看

Meta unveiled Brain2Qwerty v2, an AI system that converts brain activity into text without requiring brain surgery. - Uses non invasive MEG recordings and end to end deep learning. - Achieved 61% word accuracy on average, with the best participant reaching 78%. - Trained on 22,000 sentences, significantly outperforming previous non invasive approaches. - Meta also open sourced the code to accelerate neuroscience research.

Meta unveiled Brain2Qwerty v2, an AI system that converts brain activity into text without requiring brain surgery. - Uses non invasive MEG recordings and end to end deep learning. - Achieved 61% word accuracy on average, with the best participant reaching 78%. - Trained on 22,000 sentences, significantly outperforming previous non invasive approaches. - Meta also open sourced the code to accelerate neuroscience research.

55,421 次观看

China is literally on 🔥 Baidu from China has launched ERNIE 4.5 and ERNIE X1 and it’s freaking cheap . Here is everything you need to know. ERNIE 4.5 - Native multimodal and Outperforms GPT 4.5 in multiple benchmarks at just 1% of GPT 4.5 price - OpenAI GPT 4.5 – Input: $75 / 1M tokens, Output: $150 / 1M tokens; - ERNIE 4.5 – Input: $0.55 / 1M tokens, Output: $2.20 / 1M tokens ERNIE X1 - A deep thinking reasoning model with multimodal capabilities on par with DeepSeek R1 at only half the price See it in action and check out the pricing details👇 📹 source : yiyan[.]baidu[.]com 1/6 ERNIE 4.5 is a multimodal which can take Audio files as well.

China is literally on 🔥 Baidu from China has launched ERNIE 4.5 and ERNIE X1 and it’s freaking cheap . Here is everything you need to know. ERNIE 4.5 - Native multimodal and Outperforms GPT 4.5 in multiple benchmarks at just 1% of GPT 4.5 price - OpenAI GPT 4.5 – Input: $75 / 1M tokens, Output: $150 / 1M tokens; - ERNIE 4.5 – Input: $0.55 / 1M tokens, Output: $2.20 / 1M tokens ERNIE X1 - A deep thinking reasoning model with multimodal capabilities on par with DeepSeek R1 at only half the price See it in action and check out the pricing details👇 📹 source : yiyan[.]baidu[.]com 1/6 ERNIE 4.5 is a multimodal which can take Audio files as well.

324,324 次观看

Using Gemma 4 E2B for audio transcription on my Pixel 10 Pro. It support max 30 sec for now.

Using Gemma 4 E2B for audio transcription on my Pixel 10 Pro. It support max 30 sec for now.

89,037 次观看

Meta just announced Movie Gen, and it’s already blown away Sora, Runway, and other video generation tools! It’s probably the biggest release of the year in Gen AI. - Use simple text inputs to produce custom videos and sounds - Edit existing videos - Transform your personal image into a unique video - Generate audio effects Here are some examples with prompts and everything you need to know: 🧵

Meta just announced Movie Gen, and it’s already blown away Sora, Runway, and other video generation tools! It’s probably the biggest release of the year in Gen AI. - Use simple text inputs to produce custom videos and sounds - Edit existing videos - Transform your personal image into a unique video - Generate audio effects Here are some examples with prompts and everything you need to know: 🧵

301,343 次观看

Send me back, I heard my parents are poor. Made using Sora 2.

Send me back, I heard my parents are poor. Made using Sora 2.

151,564 次观看

Google DeepMind is absolutely on fire 🔥 they have just launched Gemini Robotics-ER 1.5 their first broadly available robotics AI model designed to act as the "high-level reasoning brain" for robots. This is Google's first Gemini Robotics model made available to all developers. - Available in preview through Google AI Studio and Gemini API. - First thinking model for robots interacting with the physical world. Handles complex commands and orchestrates sophisticated robotic behaviors. - Key capabilities: Advanced spatial reasoning, multi-step task planning, Google Search integration, precise 2D pointing, object reasoning, and video analysis. - Flexible thinking budget lets developers tune speed vs accuracy. Can think longer for complex tasks or respond quickly for reactive operations. - Enhanced safety filters refuse dangerous tasks and recognize physical limits.

Google DeepMind is absolutely on fire 🔥 they have just launched Gemini Robotics-ER 1.5 their first broadly available robotics AI model designed to act as the "high-level reasoning brain" for robots. This is Google's first Gemini Robotics model made available to all developers. - Available in preview through Google AI Studio and Gemini API. - First thinking model for robots interacting with the physical world. Handles complex commands and orchestrates sophisticated robotic behaviors. - Key capabilities: Advanced spatial reasoning, multi-step task planning, Google Search integration, precise 2D pointing, object reasoning, and video analysis. - Flexible thinking budget lets developers tune speed vs accuracy. Can think longer for complex tasks or respond quickly for reactive operations. - Enhanced safety filters refuse dangerous tasks and recognize physical limits.

137,226 次观看

Thank you Perplexity and Aravind Srinivas Bhai for the Mac Mini 🤟 Perplexity Computer + Mac Mini. Excited to test it out and see what it can do.

Thank you Perplexity and Aravind Srinivas Bhai for the Mac Mini 🤟 Perplexity Computer + Mac Mini. Excited to test it out and see what it can do.

51,494 次观看

LMAO 🤣 Sam Altman: Another thing that I think Facebook has done exceptionally well is hiring, and I always tell founders that this is the thing you need to get good at. Source: Y Combinator YT : Video Mark Zuckerberg: How to Build the Future.

LMAO 🤣 Sam Altman: Another thing that I think Facebook has done exceptionally well is hiring, and I always tell founders that this is the thing you need to get good at. Source: Y Combinator YT : Video Mark Zuckerberg: How to Build the Future.

135,031 次观看

ChatGPT Agents : Overhyped, underdelivered, and painfully slow compared to competitors - They hyped presentations as its strength - it's actually the worst - Genspark finished the entire report while ChatGPT was still "browsing" - The performance gap is massive More details & Comparison 1/5

ChatGPT Agents : Overhyped, underdelivered, and painfully slow compared to competitors - They hyped presentations as its strength - it's actually the worst - Genspark finished the entire report while ChatGPT was still "browsing" - The performance gap is massive More details & Comparison 1/5

119,935 次观看

Can someone tell Google to calm down Plz.... They have dropped another major update to NotebookLM - Interactive mode (beta) "Call in" and ask the AI hosts questions. This is gonna be awesome 🔥🔥 - New 3-panel" design. - NotebookLM plus ( new subscription plan coming soon) h/t @joshtwoodward

Can someone tell Google to calm down Plz.... They have dropped another major update to NotebookLM - Interactive mode (beta) "Call in" and ask the AI hosts questions. This is gonna be awesome 🔥🔥 - New 3-panel" design. - NotebookLM plus ( new subscription plan coming soon) h/t @joshtwoodward

139,580 次观看

Videos