Keep pulling the thread on Russ d'Sa.
Russ d'Sa believes that "turn detection," the ability for an AI to know when a user has finished speaking, is one of the hardest problems in voice AI and a key obstacle to its mainstream adoption.
The "full duplex mode" in Meta AI's app achieves low latency by processing audio directly, bypassing the intermediate step of converting speech to text.
Russ d'Sa predicts that the problem of "turn detection" for one-on-one voice AI conversations will be effectively solved within the next 12 months.
To achieve the fidelity of a human-to-human conversation, voice AI systems need to reduce their end-to-end latency to a threshold of approximately 200 milliseconds.
Russ d'Sa predicts that voice AI could achieve 200-millisecond end-to-end latency and handle complex human-like repartee within 18 to 24 months.
Russ d'Sa believes a fundamental reason for 23andMe's struggles was a timing mismatch, as the company's technology outpaced the slower-moving scientific understanding of genomics.
Russ d'Sa argues that a fundamental failure of 23andMe was its lack of a staged, strategic business model to layer in new value and revenue streams over time, unlike companies such as Tesla and SpaceX.
LiveKit's open-source platform provides the real-time audio infrastructure for OpenAI's ChatGPT voice mode and for Character AI.
The vast majority of voice AI interfaces, including the original OpenAI voice mode, use a "cascade" or "component-based" model for processing.
The "cascaded" model for voice AI processes audio through a pipeline: first to a speech-to-text (STT) model, then the resulting text is fed to an LLM, and finally the LLM's output is converted back to audio by a text-to-speech (TTS) model.
Qtai's Moshi is a 7 billion parameter model featuring a dual-channel architecture with a joint attention mechanism between its audio input and output channels, allowing it to decide when to interrupt.
Russ d'Sa predicts that future voice AI models will process audio directly, but progress is currently limited by a lack of sufficient audio-based conversational training data.