Загрузка видео...
Не удалось загрузить видео
Introducing hallucination correction. We have reduced hallucination by 70%. Giga's hallucination rate is at ~1%. Better than the best frontier models. Deploy AI your customers can trust.
1,225,973 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 32

NO MORE FAKE RESEARCH STUDIES

I need to correct my printer’s hallucinations. Prints too much

Learn more at

Check out our CTO's technical blog as well:

Impressive work on hallucination correction-lowering to ~1% is a big step for trustworthy AI! 🚀

The biggest problem with AI isn’t that it makes mistakes. It’s that sometimes it delivers wrong answers with complete confidence. That’s why hallucination correction may become the most important AI race of the next few years.

Another great launch by the team (and great video from @jcarvajalpa )

this was such a fun concept to work on with @jcarvajalpa @RyanJosephHill, @ben_aguilera’s world class animation team

Glad Giga is taking this on. Hallucinations aren’t talked about enough

No hallucinations were spoken. We're simply the best.

Game changer!

🔥🔥🔥🔥

1% is the wall not the floor. one hallucinated balance on a banking call triggers an escalation the agent can't backtrack from. what's the recovery flow when it fires?

GPT-4 Turbo benchmarks at ~3% on TruthfulQA, Claude Opus around 2.5%. If Giga is hitting 1% on real customer queries not curated benchmarks, what's the eval set? Vectara's HHEM leaderboard or in-house? Big difference between RAG-grounded 1% and open-ended 1%.

The AI hype cycle has been pretending hallucinations are “just part of the game.” They’re not. They’re the reason most AI products still feel unsafe in front of real customers. Giga cutting that down is a big deal.

so the fix is literally just exploiting the gap between how fast it generates text vs how slow humans speak?

Everyone is busy shipping AI that sounds confident while making stuff up. Giga fixing hallucinations first is exactly the boring-sounding thing that separates toys from products customers can actually trust.

This is a solid move towards making voice agents more reliable. You need a smarter model to monitor their behavior in realtime. We’ve also implemented a similar pattern for our restaurant voice ai and has improved behavior by 8x

Wow!

While reducing the hallucination rate to ~1% is an impressive achievement, I argue that the risk of AI generating hallucinations as a byproduct of prioritizing performance outcomes fundamentally arises from an overemphasis on "quality and efficiency" in its optimization objectives, with insufficient consideration given to "safety and contribution to humanity." Based on the framework of X-CII theory, I propose that integrating a real-time evaluation and control mechanism for Safety (S: harm prevention and human benefit)—in addition to Quality (Q) and Efficiency (E)—into the AI's decision-making process is essential for a fundamental resolution of this problem.

1% by what benchmark though closed evals or messy real user prompts? hallucination rate claims only matter if the test set actually hurts curious how it holds up when users go off-script

Is it jailbreak proof?

1% is cute until someone asks their balance and gets told they're broke

awesome

New to you, LFG

@garrytan Hallucinations are visible. Behavioral drift usually isn’t. That’s why production systems often fail while still looking operational.

Epic hallucination reduction

Great Giga cuts hallucinations 70%

✨

@garrytan 👀

Giga: 70% fewer hallucinations, ∼1% rate — AI you can trust

Wow, that's really nice. Next level. 😎🔥🙌🏻

