Voice is quickly becoming a major way people interact with AI agents. Recent research found that 14% of users already prefer speaking with generative AI rather than typing1, signaling that voice interactions are moving beyond early-adopter use cases and toward the mainstream. As adoption grows, users increasingly expect conversations that feel natural, responsive, and capable of carrying out real work.
From customer support and employee assistance to learning companions and industry-specific workflows, organizations want agents that can listen, reason, and respond in real time. Yet many voice solutions still require developers to assemble separate speech, orchestration, deployment, monitoring, and evaluation components before they can reach production. As a result, teams often spend more time integrating infrastructure than improving the agent experience itself.
Today’s announcement introduces voice agents in Microsoft Foundry Agent Service, a voice-native approach that brings real-time speech, agent capabilities, deployment workflows, observability, and evaluation together in a single platform.
Developers can build voice experiences with sub-second latency, multilingual support, integrated knowledge and tools, deployment channels, and built-in lifecycle management, all within Microsoft Foundry. This includes broad model choice for text-to-speech and speech-to-text models from providers like OpenAI, Microsoft AI and Microsoft Azure. Voice agents support the same development workflows developers already use in Microsoft Foundry Agent Service, making it easy to bring real-time voice capabilities to existing agent experiences.
From voice API to voice-native agents – choosing the right path for building voice experiences
Voice experiences have evolved from standalone speech applications into full agent systems. As a result, developers today can choose from multiple architectures depending on how much of the agent lifecycle they want the platform to manage.
The three approaches shown above reflect different stages of voice application maturity. Some teams only need a real-time voice runtime. Others have already built agents and want to make those experiences conversational. Increasingly, however, organizations are looking for a voice solution that spans development, deployment, monitoring, evaluation, and optimization from a single platform.
The important distinction is not simply how speech is added to an application, but where responsibility for the voice experience lives. Earlier approaches give developers maximum flexibility to assemble their own architecture. As requirements grow, however, teams often need deeper lifecycle capabilities, including observability, evaluation, governance, deployment workflows, and operational tooling. While many developers can build a compelling voice agent demo, operating voice agents in production requires visibility into every interaction, the ability to evaluate agent behavior, diagnose failures, and a system to continuously improve quality, latency, and cost. Those requirements become especially important for customer-facing experiences such as contact centers, employee assistants, and business process automation.
Voice agents in Microsoft Foundry were designed around the needs of production applications. Rather than treating voice as a channel layered onto an existing agent, voice becomes a native part of how agents are built, deployed, observed, evaluated, and optimized. Real-time speech, agent capabilities, deployment workflows, observability, and evaluation are brought together through a single platform and SDK. Teams can tailor experiences to their specific scenarios by choosing the models that best fit their workloads, improving recognition for domain-specific vocabulary through speech customization, creating custom voices, and adding photo or video avatars that reflect their brand. Voice agents also include newly expanded standard talking-head avatars that can be enabled directly in the Foundry playground, making it easy to create a visual agent experience without developing custom avatar assets. Developers can start with these out-of-the-box avatars and evolve to custom voice and avatar experiences as their scenarios grow.
All three approaches remain supported and serve different customer needs. For new enterprise voice workloads, however, Microsoft recommends voice agents in Foundry as the preferred path forward because it provides one of the most complete voice-agent lifecycle, from first conversation to production operations.
Build, test, and deploy voice agents with familiar developer tools
Voice agents are integrated into the developer workflows teams already use. With new support across AZD AI and the Foundry Toolkit for Visual Studio Code (coming soon), developers can create, configure, test, debug, source-control, and deploy voice agents without stitching together separate tools or building custom voice testing applications.
AZD AI introduces a dedicated workflow for voice agents. Developers can initialize a voice agent project, manage agent definitions and deployment configuration in source control, and deploy through the standard AZD lifecycle. Voice agents remain a dedicated application type because their models, runtime requirements, testing workflows, and deployment paths are optimized specifically for spoken interaction.
Inside Visual Studio Code, developers can complete the entire development loop without leaving the editor. Teams can create and configure a voice agent, start a voice session, interact through natural speech, inspect transcripts and runtime events, and debug latency or connection issues from the same environment where they author the agent.
Together, these capabilities reduce the time between writing an agent and having a real conversation with it, helping teams move from experimentation to production more quickly.
Operate voice agents with voice native lifecycle management
Building a great voice experience is only the first step. Once voice agents reach production, teams need to know why conversations succeed or fail, where latency is introduced, how users interact with the agent, and how the experience can be improved over time.
Voice interactions introduce challenges that text agents never encounter. Speech-recognition errors, interruptions, turn-taking issues, and tool delays can all impact the conversation experience. Without observability built into the platform, teams often rely on replaying calls, reviewing logs across multiple systems, and manually diagnosing issues.
Voice agents in Microsoft Foundry are built with observability, evaluation, and optimization by default. They use the same Foundry Observability experience as every other Foundry agents, so there is nothing separate to configure. Each conversation is stored as a trace, and the Foundry evaluation framework runs against those traces in the same way it runs for text agents, making it easier to understand agent behavior, diagnose issues, and measure quality over time.
These capabilities feed the continuous improvement loop that runs across Foundry:
- Tracing and monitoring provide visibility into every interaction, helping teams understand user behavior, diagnose failures, and find performance bottlenecks.
- Evaluation measures conversation quality against production traces, validates expected behavior, compares versions, and catches regressions before they reach users.
- Insights and optimization help teams improve quality, reliability, latency, and cost over time
Voice agents are typically built to complete specific, time-sensitive tasks. Developers need to measure whether the agent did its job and followed the rules, and general-purpose evaluations alone cannot tell them that. Foundry rubric evaluators score conversations against criteria written for a specific agent. You list what matters, such as “confirmed the date, time, and party size” or “did not book a party larger than eight,” assign each criterion a weight, and an LLM judge scores every conversation against that list.
Foundry can generate the rubric from the agent’s instructions and production traces, and run it continuously against live traffic, so teams know when an agent skips a step or breaks a policy before customers do. Early customers building voice agents in Foundry are already using rubric evaluators this way.
Together, these capabilities help teams move beyond debugging individual conversations and establish a repeatable process for operating and improving voice agents in production.
Voice Agents in Microsoft Foundry bring together real-time speech, knowledge, tools, deployment workflows, observability, and evaluation into a unified experience. Whether you’re adding voice to customer support, employee assistants, learning applications, or industry workflows, Voice agents provide the recommended path for building production-grade voice experiences on Microsoft Foundry.
Get started with voice agents in Microsoft Foundry
Getting started only takes a few steps:
- Open the Microsoft Foundry portal and create a new agent, select “Voice”.
- Choose a voice-optimized model and configure your agent’s instructions.
- Connect knowledge sources and tools as needed for your scenario.
- Test conversations directly in the built-in voice playground.
- Deploy to your desired channels and monitor performance using built-in observability and evaluation capabilities.
1Study done from Jabra and LSE. The Globe and Mail (October 2025)


