Загрузка видео...

Не удалось загрузить видео

На главную

GPT-5 fails 59% of the time pulling fresh data. Google's Gemini hallucinates facts. Claude invents sources that don't exist. Companies lose billions on bad AI research. Agrawal's solution broke every rule:

359,143 просмотров • 1 год назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Bob McGrew (Head of Research OpenAI) explains why proprietary data no longer provides companies with a competitive advantage in the AI era. Finance companies once believed their years of accumulated data would give them an edge. They planned to train specialized models on top of GPT or Llama using their exclusive information. The results shocked them. Their industry-specific models performed worse than the next generation of general purpose models. The ability to synthesize new information proved more valuable than memorizing old data. McGrew introduces the concept of "embodied labor" - the human work behind data collection. Companies spent years having employees call customers, analyze case studies, and gather information through manual processes. This accumulated knowledge required massive time and money to build. It represented thousands of hours of human effort that companies thought couldn't be replicated by competitors. But AI changes everything. Instead of years of customer calls, AI can conduct comprehensive surveys instantly. Rather than manual case analysis, AI processes thousands of examples in hours. The core insight is that value wasn't in the data itself but in the labor required to collect it. Since AI makes that labor essentially free, the advantage disappears. Companies can no longer rely on their proprietary data as a protective moat. Any competitor can use AI to replicate years of data gathering almost instantly.

Aish

240,714 просмотров • 1 год назад

Why General AI Fails Tax Professionals — and What Makes TaxGPT State-of-the-Art Tax LLM Every tax professional using general-purpose AI tools like ChatGPT or Claude for research is one hallucinated source away from costly damage to their clients and professional career. So we ran a head-to-head benchmark: TaxGPT vs. OpenAI, Anthropic, and Gemini. Our methodology: 1️⃣ Quantitative testing on CPA + EA tax-specific exam questions 2️⃣ Qualitative testing complex, real-world scenarios across trust & estate, direct and indirect taxes, payroll, multi-state nexus, advisory, and compliance Results (Accuracy): TaxGPT: 96% Gemini 2.5 Pro: 89% Claude Sonnet 4: 87% GPT-4o: 80% GPT5: 92% But here’s the real story 👇 The Source Quality Gap General LLMs are great test takers. But tax research isn’t a trivia game — it’s a liability-driven profession where every position must be defensible. When we examined the sources behind their answers, the gap became a canyon: 🔍 OpenAI included citations only 9% of the time. 🔍 91% of answers had no verifiable source trail. 🔍 When citations did appear, they averaged 3.78 sources — many from Wikipedia, CNBC, AP News, Time, NerdWallet, etc. TaxGPT averages 14 authoritative sources per answer, all from primary/secondary tax law — IRC, Treasury Regs, court cases, IRS rulings, and our proprietary tax knowledge library. More details in our official blog post in the following post.

Kash from TaxGPT.com

59,253 просмотров • 9 месяцев назад

Manish Gupta, Senior Director at Google DeepMind India, sits down with Aakrit Vaish and Pratyush Choudhury at Mumbai Tech Week for a rare on-record conversation about the frontier AI research happening out of Bangalore. Gupta makes a pointed case against the narrative that India lacks AI research talent: a team of roughly 75 researchers, a third of them fresh out of college, producing work on par with the best in the world and feeding directly into Gemini. The conversation goes deep on what DeepMind India actually builds, why Gemini is considered the most efficient model on the planet, and what India needs to become a research leader rather than a fast follower. In this conversation, they go deep on: 0:00 Intro: Manish Gupta of DeepMind India at Mumbai Tech Week 1:23 What DeepMind India does, and why it's a "mystery" 1:59 The three roles: languages, efficiency, continual learning 2:19 The Haryanvi demo at Google I/O and the cultural playbook 2:38 Making models efficient: from mobile to servers 3:25 Matryoshka transformers and why nested models win 4:20 Why Gemini is the most efficient model on the planet 5:11 Continual learning: using Gemini to improve Gemini 5:50 Google's India plans: consumers, agents, enterprise, government 7:44 Government officials live-coding with AI Studio and NotebookLM 8:30 The India AI talent debate and how the team is structured 10:25 Why India lacks courage and R&D investment, not talent 11:16 DeepMind's global labs and India's outsized impact 13:15 On-device AI: Gemma 3n and 4n 14:08 The headline: 75 world-class researchers in Bangalore 14:58 25 of the 75 are fresh out of college If you're a founder, builder or researcher thinking about AI, frontier models, or India's place in global AI research, this one's for you. Aakrit Vaish Pratyush Choudhury (PC) Google DeepMind Google India Manish Gupta

Activate

14,136 просмотров • 2 месяцев назад