Teams developing voice agents need automated evaluation because human review alone cannot keep pace with thousands of calls and frequent agent updates. Microsoft Foundry addresses this need with multi-turn evaluators that assess agent behavior and outcomes across complete conversations. To understand how these evaluators behave on voice-agent traces, we evaluated Task Completion and Generated Rubric on benchmark conversations and traces from a Foundry voice agent. This article examines call-level accuracy, agent ranking, run-to-run agreement, and rubric-based scoring.
Authors: Sheng Liu, Shivank Goel, Morteza Ziyadi
The Evaluators
Both task completion evaluator and rubric evaluator take a complete multi-turn conversation as input and use a judge model to automatically assess it. They differ in their evaluation criteria and outputs.
- Task Completion: evaluates whether the user’s task was successfully and completely accomplished. It returns a binary success result (true or false) and an explanation of why the task was judged complete or incomplete.
- Generated Rubric: evaluates the conversation against criteria generated from user-provided evaluation requirements. It returns a weighted overall score from 0 to 1, a pass/fail result, and per-criterion scores with explanations.
Study Design
Task Completion
Datasets. We evaluated task completion on two trace datasets:
- τ-Voice benchmark traces. This dataset contains 3000+ simulated calls across airline, retail, and telecom tasks, produced by 12 voice-agent configurations spanning audio-native and cascaded systems. Figure 1 focuses on the nine gpt-realtime, Gemini Live, and cascaded configurations included in this article. We only keep conversations where τ-Voice’s user simulator follows the instruction it receives.
- Foundry agent traces. This dataset contains calls generated by one Foundry voice agent backed by gpt-realtime-2 on τ-Voice tasks. We only keep conversations where τ-Voice’s user simulator follows the instruction it receives.
Judges. We kept the Task Completion evaluator fixed and report results from seven GPT judge configurations. For each call, we compared the evaluator verdict with the reference pass/fail label in the study data. Agent ranking and run-to-run agreement were measured on the τ-Voice benchmark traces.
Generated Rubric
We evaluated Generated Rubric on telecom conversations produced by one Foundry voice agent backed by gpt-realtime-2 on τ-Voice tasks. gpt-5.4-mini generated the rubric from requirements describing the expected behavior. gpt-5.4 and gpt-5.6-sol each scored the same 114 conversations with Generated Rubric, and we compared their predictions with the reference pass/fail labels.
Metrics
Balanced accuracy is our primary correctness metric for both evaluators. For task completion, we also report raw accuracy, ranking agreement, and repeatability.
- Balanced accuracy. The average of success recall and failure recall: (success recall + failure recall) / 2. Each recall is the proportion of calls with that outcome labeled correctly. Equal weighting prevents the majority class from dominating the score.
- Raw accuracy. The percentage of calls whose verdict matches the reference label. Unlike balanced accuracy, it reflects the dataset’s mix of successful and unsuccessful calls.
- Kendall’s τ. Agreement between evaluator and benchmark agent rankings. It ranges from -1 for reverse order to 1 for identical order; 0 indicates no overall rank association. Higher values mean more consistent rankings.
- Pairwise order agreement. The percentage of agent pairs ordered identically by the evaluator and benchmark. The 12 agents form 66 pairs. We report agreement for all pairs and for clearly separated pairs whose benchmark pass rates differ by more than 10 percentage points.
- Repeatability. Consistency when four agents are graded twice under identical settings. We report the percentage of unchanged predictions, absolute changes in balanced accuracy and whether agent rankings are preserved. Higher agreement and smaller accuracy changes indicate more stable judgments.
Task Completion Results
Call-Level Accuracy
In our Microsoft internal study, Task Completion produced approximately 83% balanced accuracy and 85% raw accuracy with gpt-5.6-terra as judge on the kept τ-Voice benchmark calls. Table 1 summarizes one reported value for each metric across the seven judges retained in this article.
|
Metric |
Rounded study value |
|
Balanced accuracy |
83% |
|
Raw accuracy |
85% |
|
Kendall’s τ |
0.818 |
|
Pairwise agreement: all pairs |
91% |
|
Pairwise agreement: gap > 10 pp |
100% |
|
Prediction agreement |
97% |
Table 1. Cross-judge metric summary for the seven GPT judge configurations retained from our Microsoft internal study. Percentage values are rounded to the nearest whole percent; Kendall’s τ is shown to three decimals. Values in different rows can come from different judges and should not be read as the results of one configuration. The study did not test statistical significance.
Figure 1 reports balanced accuracy for each displayed judge-agent combination in our Microsoft internal study. Its nine rows are the gpt-realtime, Gemini Live, and cascaded voice-agent configurations retained for this article; its seven columns are the retained GPT judge configurations. The matrix is descriptive, and the study did not test whether differences between judges were statistically significant or generalize beyond these traces.
Figure 1. Descriptive balanced accuracy for 63 displayed judge-agent combinations from our Microsoft internal study.
Agent Ranking
Many evaluation decisions concern agent ordering, not just whether one call passed. In our Microsoft internal study, we ranked the 12 agents in the τ-Voice benchmark dataset by evaluator-derived pass rates and compared the ordering with the benchmark. Across the seven GPT judges retained in this article, Kendall’s τ ranged from 0.758 to 0.818. Pairwise agreement included an observed value of 90.9% across all 66 agent pairs, with a mean of 89.8% across the seven judges.
|
Judge |
Kendall’s τ |
Pairwise agreement rates (%) |
|
|
Every pair |
Gap > 10 % |
||
|
gpt-5.4-mini |
0.818 |
90.9 |
100.0 |
|
gpt-5.4-nano |
0.818 |
90.9 |
100.0 |
|
gpt-5.6-luna |
0.818 |
90.9 |
100.0 |
|
gpt-5.1 |
0.788 |
89.4 |
100.0 |
|
gpt-5.4 |
0.788 |
89.4 |
100.0 |
|
gpt-5.6-sol |
0.788 |
89.4 |
100.0 |
|
gpt-5.6-terra |
0.758 |
87.9 |
100.0 |
Table 2. Descriptive agent-ranking results for the seven GPT judges retained from our Microsoft internal study. “Every pair” covers all 66 pairs of 12 agents. “Gap > 10 pp” restricts the calculation to pairs whose benchmark pass rates differ by more than 10 percentage points. No significance analysis was conducted for differences among judges.
For pairs whose benchmark pass rates differed by more than 10 percentage points, all seven retained judges ordered every pair in the same direction as the benchmark. More disagreements were observed among pairs with closer benchmark pass rates. These observations apply only to the tested agents, traces, judges, and ranking procedure.
Results Across Judges
With gpt-5.6-terra as judge, Task Completion produced 82.7% balanced accuracy across the τ-Voice benchmark traces. With gpt-5.6-luna, it produced 80.3%, an observed difference of 2.4 percentage points. The corresponding Kendall’s τ values were 0.758 and 0.818, and pairwise agreement values were 87.9% and 90.9%. These descriptive measurements reflect different aspects of the tested configurations. The study did not evaluate pricing, latency, volume, or economic value, and it did not test whether the observed differences were statistically significant or equivalent; the values therefore do not establish a recommended default judge.
|
Agent |
Judge |
Balanced |
|
All agents |
gpt-5.6-terra |
82.7 |
|
Additional measured result |
gpt-5.6-luna |
80.3 |
|
gpt-realtime |
gpt-5.6-terra |
82.0 |
|
Additional measured result |
gpt-5.4-mini |
77.4 |
|
Cascaded |
gpt-5.6-terra |
80.8 |
Table 3. Selected descriptive results for GPT judges by retained voice-agent group in our Microsoft internal study.
Observed scores varied by retained voice-agent group in Table 3. With gpt-5.6-terra, the displayed values were 82.7% across all benchmark agents, 82.0% for gpt-realtime, 78.6% for Gemini Live, and 80.8% for the cascaded configuration. The additional measured values were 80.3% with gpt-5.6-luna across all agents and 77.4% with gpt-5.4-mini for gpt-realtime. These descriptive differences do not establish that one judge is generally superior or that agent architecture caused the observed values. Figure 1 provides the corresponding measurements for the nine displayed voice-agent configurations.
Repeatability
To measure run-to-run agreement in our Microsoft internal study, we graded the same four benchmark voice-agent configurations twice with each of the seven GPT judges retained in this article under unchanged settings. The subset contained exactly 1,109 calls per judge and included audio-native and cascaded configurations. We call an individual verdict that changes between the two runs a prediction flip, and a judge’s flip rate is the percentage of calls with a flipped verdict. Across the seven judges, mean prediction agreement was 95.1%, mean absolute change in balanced accuracy was 0.8%, and the largest observed change was 2.4%.
|
Judge |
Balanced |
Balanced |
Change |
Prediction |
Same agent |
|
gpt-5.6-terra |
82.4 |
80.9 |
1.5 |
6.0 |
yes |
|
gpt-5.6-sol |
81.8 |
82.0 |
0.2 |
3.4 |
yes |
|
gpt-5.4 |
79.2 |
78.9 |
0.3 |
4.0 |
yes |
|
gpt-5.6-luna |
78.9 |
79.9 |
1.0 |
4.2 |
yes |
|
gpt-5.4-mini |
78.1 |
75.7 |
2.4 |
8.8 |
no |
|
gpt-5.1 |
77.8 |
77.6 |
0.2 |
4.1 |
yes |
|
gpt-5.4-nano |
74.6 |
74.5 |
0.1 |
4.1 |
yes |
Table 4. Two runs under unchanged settings on the same set of calls. Changes in balanced accuracy are calculated from the displayed score. Prediction flips are the percentage of individual verdicts that changed.
Six of the seven retained judges reproduced the same agent ranking. gpt-5.6-sol had a 3.4% flip rate and preserved the ranking; gpt-5.4-mini had a 2.4% accuracy change and an 8.8% flip rate and changed the ranking. These cases show why prediction agreement and ranking stability are reported as separate descriptive measures. They do not establish a general repeatability ordering among models.
Foundry Agent Results
We applied the same Task Completion evaluator and judge panel to the 278 Foundry agent traces described in Study Design. These calls were generated by one Foundry voice agent backed by gpt-realtime-2 on τ-Voice tasks and collected through Foundry’s observability pipeline.
|
Judge |
Balanced |
Raw |
|
gpt-5.6-terra |
80.7 |
80.2 |
|
gpt-5.6-luna |
79.4 |
76.4 |
|
gpt-5.4 |
79.3 |
77.5 |
|
gpt-5.6-sol |
78.2 |
77.8 |
|
gpt-5.4-nano |
76.9 |
73.5 |
|
gpt-5.4-mini |
76.7 |
76.3 |
|
gpt-5.1 |
76.6 |
74.7 |
|
Mean (7 listed judges) |
78.3 |
76.6 |
Table 5. Descriptive Task Completion results for the seven GPT judges retained from our Microsoft internal study, measured on our kept τ-Voice task traces from one Foundry voice agent backed by gpt-realtime-2. The mean row averages the seven listed judges. No confidence intervals or significance tests were calculated.
Across the seven retained judges, balanced accuracy on the 278 Foundry agent traces averaged 78.3% and ranged from 76.6% to 80.7%. With gpt-5.6-terra, the observed values were 80.7% balanced accuracy and 80.2% raw accuracy, compared with 82.0% balanced accuracy on the gpt-realtime subset of the τ-Voice benchmark traces. With gpt-5.6-luna, the corresponding balanced-accuracy values were 79.4% and 79.2%. These measurements describe the two tested datasets; without significance or equivalence testing, they do not establish matching performance or a recommended judge.
Generated Rubric Results
Task Completion applies a common definition of successful task resolution. Generated Rubric translates user-provided evaluation requirements into scoring criteria. We measured both evaluators on the same telecom conversations to describe their observed scores under the two tested judge configurations.
Table 6 reports both evaluators on the same set of kept τ-Voice telecom conversations generated by one Foundry voice agent backed by gpt-realtime-2. gpt-5.4 and gpt-5.6-sol each applied both evaluators to the same calls.
|
Judge |
Rubric |
Task completion |
Difference |
|
gpt-5.4 |
74.5 |
76.8 |
−2.3 |
|
gpt-5.6-sol |
78.5 |
76.3 |
+2.2 |
Table 6. Generated Rubric and built-in Task Completion on the same set of kept τ-Voice telecom conversations generated by one Foundry voice agent backed by gpt-realtime-2. gpt-5.4-mini generated the rubric. The difference is Generated Rubric balanced accuracy minus Task Completion balanced accuracy.
In our Microsoft internal study, Generated Rubric produced 74.5% balanced accuracy and Task Completion produced 76.8% with gpt-5.4 as judge, a difference of −2.3%. With gpt-5.6-sol, the observed values were 78.5% and 76.3%, a difference of +2.2%. The study did not run significance or equivalence tests, so these differences do not establish that either evaluator is superior or that their accuracy is equivalent.
Generated Rubric lets teams turn user-provided evaluation requirements into scoring criteria without writing a rubric from scratch. Those requirements can cover task outcomes, expected behavior, or other aspects of quality that matter to an application. On the tested telecom conversations, Table 6 reports the observed scores; this study does not establish equivalent accuracy to built-in Task Completion.
Observed Results and Study Scope
In this Microsoft internal study, Task Completion and Generated Rubric produced the descriptive measurements reported above for the specified τ-Voice benchmark calls, Foundry agent traces, retained GPT judge configurations, and evaluator settings. The findings do not establish general evaluator reliability, judge superiority, or equivalence between the two evaluators. They illustrate how call-level accuracy, agent ranking, run-to-run agreement, and rubric-specific criteria can be examined within a defined evaluation setup.
Get Started
To evaluate your own voice agent, open your project in Microsoft Foundry and select conversations that reflect the tasks and behaviors you want to assess. Follow the conversation evaluation guide to run task completion across complete conversations. To evaluate against your own requirements, use the rubric evaluator guide to generate a rubric and apply it to those conversations. Review the scores and explanations to understand where the agent meets your expectations and where it needs improvement. The voice-agent observability guide explains how to capture and inspect the conversation traces. The voice-agent observability and evaluation workflow discussed here is currently in Preview, is provided without a service-level agreement, and is not recommended for production workloads. Availability and interface paths may vary by region and release.


