Voice AI Has Not Reached Its ChatGPT Moment Yet, Industry Execs Say
Technologyby Ivan Mehta

Voice AI Has Not Reached Its ChatGPT Moment Yet, Industry Execs Say

Key Takeaways

  • Voice AI has not yet reached its defining ChatGPT moment despite billions in investments.
  • Fast reasoning and full-duplex capabilities are necessary for truly natural conversations.
  • Automatic Speech Recognition errors can break downstream tasks and destroy user trust.
  • Transparency regarding AI usage and call recording is critical for enterprise adoption.

The theory of voice being the next big interface has picked up strong momentum, with investors pouring billions of dollars into voice AI startups working on areas ranging from model makers to enterprise customer service providers, and from meeting note-takers to AI-powered dictation. Every week there is a new model or a tool release that claims to sound human and converse like one. However, in reality, that might not be the case.

Enterprise voice AI platform PolyAI’s CTO Shawn Wen thinks that despite the release of full-duplex models — which can speak while listening to you — voice AI doesn’t have its "ChatGPT moment" yet. He points out that while developing full-duplex models is a major milestone, the next challenge is to make reasoning very fast so that models can fetch answers quickly and make conversations feel truly natural.

Wen also emphasized that AI agents in customer service should not sound robotic and must give callers enough confidence that they can actually solve problems. Once the voice quality improves and customers are willing to engage for the first few turns, they begin to build confidence. Over time, they realize they might not need to talk to a human if the AI agent resolves their issues effectively.

Alex Gay, CMO for meeting notetaker Otter, shared a similar perspective, explaining that speaker identification, intent capture, and connecting those elements with organizational knowledge are key steps for enabling true automation. Otter is also working on digital twins to represent people in meetings. For that technology to succeed, Gay noted that it is paramount for the output voice to convey the same emotive expressions as a real human in a meeting.

Gay highlighted that the best conversations involve debate, strategic discussions, and underlying relationships. If users cannot experience that with an avatar, the tool remains nothing more than a basic question-and-answer chatbot.

Despite these advancements, voice AI models still face significant hurdles in understanding users properly. Wen believes that Automatic Speech Recognition (ASR) models often miss important keywords, creating issues in capturing the full context of a conversation. When a meeting notetaker displays the wrong transcript or summary, the utility of the tool diminishes rapidly.

Otter’s Gay agreed, adding that transcription was never the end point for their company, but rather a foundational layer to drive productivity gains. If the original transcription lacks accuracy, all follow-up actions become flawed, destroying user trust in the platform.

Transparency remains another critical issue for the industry. Tools must clearly declare to customers when they are being recorded or interacting with an artificial intelligence. Both PolyAI and Otter emphasize the importance of notifying participants and building trust to ensure broader adoption of voice technologies in enterprise environments.

Recommended for you

Tools and services we trust to boost productivity and content workflows.

Browse picks
Original source →