Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

At Standard Intelligence we’ve been researching scalable cross-modality learning. We’re excited to share some early results in the form of 𝗵𝗲𝗿𝘁𝘇-𝗱𝗲𝘃, an open-source, first-of-its-kind base model for full-duplex conversational audio. 1/

179,831 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

Hertz-dev is an 8.5B parameter transformer trained on 20 million unique hours of high-quality audio data. We’ve released checkpoints and code for both mono and full-duplex generation on our website under the Apache license.

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

Hertz-dev is a base model, without fine-tuning, RLHF, or instruction-following behavior. It can be fine-tuned by researchers for almost 𝘢𝘯𝘺 audio modeling task, from live translation to classification.

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

Base models excel at faithfully modeling their training set, and accurate maps come from contact with reality. From the world’s largest dataset of high-quality real-world conversational audio, hertz-dev learned human-like speech patterns such as pauses and emotional inflections.

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

Hertz-dev has a 80ms theoretical average latency, and benchmarks 120ms real-world latency on a single RTX 4090—1.5-2x lower than the previous state of the art. Low latency is necessary for natural audio, and we're proud to move the field in this direction.

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

We’re currently training a scaled, 70B parameter version of Hertz, and we’ll be expanding to more modalities in the future. We’re excited to see what the research community builds on top of this model.

Profilbild von jian
jianvor 1 Jahr

This is impressive! Seems like the training dataset is mostly podcast? And FYI, I believe there’s also a fully-duplex vision/audio model out there, would be interested in learning more about the implementation!

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

cool project! would love to see our base model used in projects like this one

Profilbild von pranav ⠕
pranav ⠕vor 1 Jahr

i love small business sunday

Profilbild von Standard Intelligence
Standard Intelligencevor 1 Jahr

small. business. sunday.

Profilbild von Nicholas Charette
Nicholas Charettevor 1 Jahr

so happy we got this out. base models are very important research artifacts to have publicly available, and i'm glad to help ensure that they exist further into the timeline:)

Ähnliche Videos