Tag: speech-to-text

Blog
>
Tag: speech-to-text

Enhanced speech recognition model is now available

62% Word Error Rate (WER) improvement for US English

ASR speech-to-text

Hot Summer Speech-to-Text Updates

Following Google’s release of new Speech API, we are happy to announce improved quality of call records transcription.

Extend Cartesia Line Agents to SIP, WhatsApp, and Global Phone Networks

Voximplant now includes a native Cartesia Line / Agents connector that connects any Voximplant call to a Cartesia Line voice agent for real-time, speech-to-speech conversations—over PSTN, SIP, WebRTC, or WhatsApp Business Calling—without building custom media gateways or WebSocket streaming infrastructure.

voxengine voice ai MCP Client

Voice AI agents can now act, not just talk: introducing the VoxEngine MCP Client

Voximplant now includes a native MCP Client for VoxEngine, giving developers direct connectivity to any MCP server and full control over every tool call

TTS streaming gemini elevenlabs voice agent

Introducing Gemini 2.0 Flash Live API Client and ElevenLabs Streaming TTS integration

New integrations for Voice AI have arrived: Google's Gemini 2.0 Flash model, featuring seamless voice-to-voice conversation capabilities and ElevenLabs low-latency streaming speech synthesis are now available for Voximplant developers

elevenlabs voice agent voice ai conversational ai

Introducing integration with ElevenLabs Conversational AI

Connect any Voximplant call to ElevenLabs Conversational AI agents

voice ai agent skills

Your AI coding agent can now build on Voximplant

Voximplant AI Agent Skills let your coding agent build and ship voice applications without switching tools

Cartesia Realtime TTS now available in Voximplant

Voximplant now includes a native Cartesia module for streaming, low-latency text-to-speech (TTS). You can use a single VoxEngine API to synthesize speech in real time, connect it to any call (PSTN, SIP, WebRTC, WhatsApp) and control playback from a Large Language Model (LLM) or other source, all inside VoxEngine.

TTS text-to-speech voice ai realtime

Inworld Text-to-Speech now available in Voximplant

Voximplant has new realtime speech generation for voice AI from Inworld, our latest Voice AI text-to-speech (TTS) partner. Together, we combine state-of-the-art TTS with carrier-grade connectivity so you can build voice agents that sound like your brand, not a generic robot.

Grok Voice Agent API now available in Voximplant

Voximplant now includes a native Grok module that connects any Voximplant call to xAI’s Grok Voice Agent API for real-time, speech-to-speech conversations. With a single VoxEngine scenario, you can interact via audio with Grok over phone numbers, SIP trunks and infrastructure, WhatsApp Business, or WebRTC into Grok — all without building custom media gateways or WebSocket streaming infrastructure.

voximplant kit podcast voximplant-kit-cc-news product management voximplant-kit-automation-news web sdk webrtc video kit-updates call center ios sdk sip voximplant pstn api

Tag: speech-to-text

Enhanced speech recognition model is now available

Hot Summer Speech-to-Text Updates

Sign Up for a free Voximplant developer account or talk to our experts

Extend Cartesia Line Agents to SIP, WhatsApp, and Global Phone Networks

Voice AI agents can now act, not just talk: introducing the VoxEngine MCP Client

Introducing Gemini 2.0 Flash Live API Client and ElevenLabs Streaming TTS integration

Introducing integration with ElevenLabs Conversational AI

Your AI coding agent can now build on Voximplant

Cartesia Realtime TTS now available in Voximplant

Inworld Text-to-Speech now available in Voximplant

Grok Voice Agent API now available in Voximplant

Sign Up for a free Voximplant developer account or talk to our experts

Tag: speech-to-text

Sign Up for a free Voximplant developer account or talk to our experts

Sign Up for a free Voximplant developer account or talk to our experts

Contact Us