Loading video...

Video Failed to Load

Go Home

Introducing Nova-2, our next-gen model for superhuman speech-to-text. TL;DR Nova-2 delivers: 💥 Next-level accuracy: +18% accuracy than Nova-1 & over 36% accuracy than OpenAI Whisper large 💥 Up to 40x faster 💥 Same low cost: 3-7x cheaper 🧵👇

2,184,459 views • 3 years ago •via X (Twitter)

10 Comments

Deepgram's profile picture
Deepgram3 years ago

Extending upon Nova's groundbreaking training, which spanned +100 domains and 47 billion tokens, Nova-2 continues to be the deepest-trained ASR model in the world.

Deepgram's profile picture
Deepgram3 years ago

Nova-2 was trained in a 2-stage curriculum starting from the largest, most diverse dataset in Deepgram’s history: nearly 6M resources and an extensive library of high-quality human transcriptions. The result? 👇

Deepgram's profile picture
Deepgram3 years ago

A new state-of-the-art model capable of superhuman transcription performance that consistently outperforms any other STT model in the market today across a wide range of speech application domains. Onto the benchmark results…

Deepgram's profile picture
Deepgram3 years ago

In our benchmarking, Nova-2 has an overall WER of 8.4% for the median files tested, representing a 16.8% relative error rate improvement compared to the closest provider. Nova-2 surpassed all tested competitors by an average of 30% and outperformed OpenAI Whisper large by 36%.

Deepgram's profile picture
Deepgram3 years ago

Modern speech apps are increasingly used to automate real-time interactions with end users for use cases like agent assist and live captioning. But there are limited options for true real-time STT and several providers like OpenAI lack native streaming models...

Deepgram's profile picture
Deepgram3 years ago

...However, in our real-time accuracy benchmarking, Nova-2 handily outperforms the field with an average relative reduction in WER of 28.6% across all domains.

Deepgram's profile picture
Deepgram3 years ago

Regarding speed, our benchmarks reveal that Nova-2 surpasses all other STT models, achieving a median inference time of 29.8 seconds per hour of diarized audio. This represents a significant speed advantage ranging from 5-40x faster than comparable vendors offering diarization.

Deepgram's profile picture
Deepgram3 years ago

In terms of cost, Nova-2 maintains the same starting price as Nova at just $0.0043 per minute of pre-recorded audio, nearly 3-5x more affordable than any other full-functionality provider (based on currently listed pricing) in the market.

Deepgram's profile picture
Deepgram3 years ago

Since launching Nova-1 this year, we have also released new features encompassing improved speaker diarization, smart formatting, filler words support, and our inaugural domain-specific language model for summarization.

Deepgram's profile picture
Deepgram3 years ago

You can dive deeper into our approach to model development and the benchmarks in the full announcement. Plus, get started with Nova-2 by requesting early access. Link to announcement:

Related Videos

Introducing LifeGPT, showing that LLMs can simulate complex, Turing-complete systems like Conway's Game of Life with near-perfect accuracy—no prior topology needed.🌐This unlocks new potential for AI in modeling self-organizing systems in biology, materials science, & beyond.🔬🤖 #AI #LifeGPT. Cellular Automata (CA), like Conway's Game of Life ("Life"), are computationally irreducible, meaning their evolution is difficult to predict without an a-priori understanding of the rules of the game, including the topology on which it is played. LifeGPT is a topology-agnostic generative model that learns the rules of Life without prior knowledge of its grid structure or boundary conditions, from only a tiny number of game states. The success in simulating Life suggests promising avenues for scientific discovery, particularly in bridging the gap between AI, artificial life, and real-world biological systems, for both forward and inverse problems. The potential for universal computation within generative AI, including LLMs, through approaches like LifeGPT, represents an exciting area for future research, especially when combined with reinforcement learning. Model Convergence: LifeGPT exhibits rapid convergence during training, achieving high accuracy in predicting next-game-states. We attribute the non-zero cross-entropy loss to the lack of causal relationships within randomly generated ICs. Accuracy & Temperature: LifeGPT achieves near-perfect accuracy, particularly at lower sampling temperatures, but can be continually tuned towards higher creativity to discover patterns that the original ruleset would not be able to produce. This finding highlights the trade-off between model creativity (higher temperature) and accuracy in deterministic predictions, with high relevance to model real-world dynamical systems for which no closed-form rulesets exist. Zero/Few-Shot Learning: Trained on a small fraction of possible initial conditions, LifeGPT demonstrates strong zero/few-shot learning, accurately simulating Life for unseen initial conditions. However, rare prediction errors highlight that LifeGPT approximates rather than perfectly replicates the Life algorithm. Autoregressive Autoregressor: A recursive implementation of LifeGPT demonstrates the model's ability to simulate Life over multiple timesteps. LifeGPT is topology-agnostic with respect to its training data and our results show that a GPT model is capable of capturing the deterministic rules of a Turing-complete system with near-perfect accuracy, given sufficiently diverse training data. The work showcases the possibility for future models to synthesize stochastic generative capabilities with deterministic computational capabilities. Link to code, paper, etc. below. Podcast generated using #NotebookLM. LAMM@MIT DMSE at MIT

Markus J. Buehler

114,251 views • 2 years ago