Skip to content

Evaluate More, Spend Less: Batching Microsoft Foundry Evaluators for Efficient Evaluation

Authors: Salma Elshafey, Ali Mahmoudzadeh, Kayla Ames, Ahmad Qardahji, Vivek Bhadauria, Morteza Ziyadi, April Kwong


A single agent trajectory may need to be evaluated for groundedness, coherence, instruction following, task completion, and correct tool use. In a conventional pipeline, each evaluator receives the same messages, tool calls, tool results, and tool definitions in a separate model call. As trajectories grow, evaluation repeatedly pays to process the same context. 


Can one LLM judge call apply five or six evaluators without losing the signal that makes evaluation useful? This study tests a simple design: send the shared context once and ask one composite evaluator to score several criteria in the same call.


In summary: On 100-row quality and tool use samples, composite evaluation required 5-6x fewer model calls, reduced total input-token use by 61.25–71.07%, completion tokens by 46.63–63.89%, and measured run wall time by 35.74–46.10%. Across the tested workloads, frontier judge models maintained stable quality across several high-value dimensions. The results support composite evaluation as a cost-efficient default, with targeted individual evaluation for workload-sensitive rubric criteria.


Why Evaluation Becomes Expensive


Evaluating a single conversation rarely involves just one evaluation dimension. A typical assessment examines multiple dimensions of quality, such as whether the response is coherent, follows instructions, completes the user’s task, remains grounded in available evidence, and uses tools correctly. Because each dimension is typically evaluated through a separate prompt and model call, the same conversation trajectory, tool calls, tool outputs, and tool definitions are processed repeatedly. As a result, cost grows with both the number of evaluation criteria and the amount of shared context, making long agent trajectories especially expensive to evaluate.



The Composite-Evaluator Approach 


Rather than evaluating each dimension independently, composite evaluation scores multiple criteria within a single judge-model call, allowing the shared context to be processed once instead of repeatedly. We built two composite evaluators.


The Output Quality Evaluator


The Output Quality Evaluator scores six Microsoft Foundry evaluators in one LLM call:

Evaluator

Question 

Fluency 

Is the response clear and well formed? 

Coherence 

Is it logically consistent with the conversation? 

Intent Resolution 

Did the assistant understand the user’s goal? 

Task Adherence 

Did it follow the user’s instructions? 

Groundedness 

Are its claims supported by the available evidence? 

Task Completion 

Did it complete the requested task? 

Tool Use Quality Evaluator 


The Tool Use Quality Evaluator scores five Microsoft Foundry evaluators in one LLM call: 

Evaluator

Question 

Tool Call Accuracy 

Was the tool call correct overall? 

Tool Call Success 

Did the invocation succeed? 

Tool Input Accuracy 

Were the arguments correct? 

Tool Output Utilization 

Did the response use the tool output correctly? 

Tool Selection 

Was the appropriate tool selected? 

 


Each composite prompt includes the shared trajectory and tool definitions once, followed by the full definition, rating scale, and applicability rules for every evaluator. It instructs the judge to assess each evaluator independently so that one verdict does not bias another. The judge returns structured results for each criterion, including a score, rationale, and applicability status. For multi-turn evaluations, it can also identify the earliest turn where a failure occurred.


How We Validated Evaluator Quality


The study compares individual and composite evaluators across: 



  • Two modes: Single-Turn and Multi-Turn. 



  • Nine judge models: GPT-4o, GPT-5.4, GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, DeepSeek V4 Flash, DeepSeek V4 Pro, and Grok 4.1 Fast Reasoning. 



  • Multiple datasets, as shown in the table below 

Dataset 

Mode 

Rows 

Primary validation target 

Reference labels 

Internal quality set 

Single-Turn and Multi-Turn 

283 

Six quality evaluators

Existing per-evaluator labels 

Internal tool use set 

Single-Turn and Multi-Turn 

200 

Five tool use evaluators

Existing per-evaluator labels 

BFCL v4 

Multi-Turn 

200 

Task Completion, Tool Call Accuracy, and failed-turn localization 

Deterministic state and tool-call checks 

AgentIF 

Multi-Turn 

148 

Task Adherence 

Constraint-level majority-vote reference 

FaithDial 

Multi-Turn 

300 

Groundedness 

Faithful versus hallucinated responses 

FED 

Multi-Turn 

125 

Coherence 

Five-annotator human scores 

Tau-Voice 

Multi-Turn 

278 

Task Completion 

Reward-derived completion labels 

AgentRx tau_retail 

Multi-Turn 

29 

Failed-turn localization 

Human failure annotations 

 


We measured several distinct properties: 



  • Input tokens, output tokens, model calls, and latency

  • Accuracy, macro-F1, and Cohen’s kappa where labels were available 

  • Agreement and correlation structure between individual and composite modes 

  • Repeatability across four evaluation runs 


Finding 1: Composites Cut Input Tokens by 61.25–71.07% and Latency by 35.74–46.10%


We measured cost and latency on matched 100-row Single-Turn samples: one for Output Quality and one for Tool Use, comparing parallel individual evaluators with one composite evaluator per row. For each row, evaluator context includes the query and response, any tool calls and results, and the serialized tool definitions. 


Note: Single-turn means that the evaluator’s focus is the last agent response only, but the entire conversation history is passed to the evaluator as context.


Trajectory Length in the Cost Sample 


Distribution of trajectory lengths, before evaluator reformatting

The log-scale distributions show that most rows clustered around a few thousand tokens, with a smaller number of substantially longer trajectories. In both samples, longer evaluator contexts produced larger absolute token savings.



The Measured Cost Reduction 

Across the sample: 



  • Repeated context is the main saving. Total input-token volume fell from 2,058,930 to 797,843 for Output Quality, a 61.25% reduction, and from 2,036,122 to 589,097 for Tool Use, a 71.07% reduction. Uncached input fell by 57.51% and 53.29%, respectively. For Output Quality, 52.72% of composite input tokens were cached versus 56.88% across the individual suite; for Tool Use, the figures were 46.22% versus 66.69%. However, the composites still processed substantially fewer tokens overall.

  • Call volume collapses. Output Quality fell from 600 individual evaluator calls to 100 composite calls. Tool Use fell from 494 calls to 100. In six rubric-row combinations, native applicability rules legitimately returned no score, such as when there was no tool call to evaluate; those skips explain why the individual Tool Use baseline is below 500.

  • Measured runs finish sooner. Output Quality wall time fell from 504.41 seconds to 324.12 seconds, a 35.74% reduction, while Tool Use fell from 536.95 seconds to 289.42 seconds, a 46.10% reduction. Per sample row, that is 5.044 to 3.241 seconds for Output Quality and 5.369 to 2.894 seconds for Tool Use. Completion tokens fell by 46.63% and 63.89%, respectively.


Parallel individual and composite evaluator token and wall-time comparison

On the quality sample, the composite used 797,843 total input tokens and 36,919 completion tokens, compared with 2,058,930 and 69,180 for parallel individual evaluation. On the tool use sample, the composite used 589,097 total input tokens and 28,358 completion tokens, compared with 2,036,122 and 78,522. These are measured run totals, not prompt-length estimates. 


The broader experiments showed the same benefit. The full six-evaluator rubric Single-Turn study used about 68% fewer input tokens, while the 283-row Multi-Turn quality study used about 77% fewer: 28,715 rendered input tokens per row for a six-call equivalent versus 6,608 for the composite call. Exact currency savings depend on model and deployment pricing, but batching consistently removed duplicated input and round-trip overhead.


Finding 2: Quality Tradeoffs Depend on Rubric and Judge Model


Output Quality Rubric Benchmark Results 


On the internal Single-Turn quality set, composite Task Adherence accuracy improved for every judge, while Task Completion remained close to individual evaluation. External Multi-Turn benchmarks showed a more varied pattern across rubric, judge, and workload. 


External benchmark delta in Cohen’s Kappa

This figure reports Δκ = κComposite – κIndividual: blue favors composite evaluation and orange favors individual evaluation. The mixed directions reinforce the need to validate the selected rubric and judge on production-like data. 


BFCL provided the strongest labeled Multi-Turn example for Task Completion. With GPT-5.6 Luna, the composite evaluator reached macro-F1 0.848 and κ = 0.696, compared with 0.836 and 0.672 for the individual evaluator. Across the other judges, composite and individual Task Completion remained close for Sol, Terra, and DeepSeek V4 Flash, while results for DeepSeek V4 Pro and Grok 4.1 Fast favored individual evaluation. 


Tool Use Quality Rubric Benchmark Results 


The five-rubric tool composite was evaluated on production-style tool traces. Each row carries ground truth for one source rubric, so the table reports accuracy on the available per-evaluator subsets.


Judge 

Tool Call Accuracy 

Tool Call Success 

Tool Input Accuracy 

Tool Output Utilization 

Tool Selection 

gpt-4o 

0.800 

0.902 

0.850 

0.800 

0.878 

gpt-5.4 

0.800 

0.902 

0.775 

0.737 

0.829 

gpt-5.4-mini 

0.750 

0.902 

0.725 

0.763 

0.756 

gpt-5.6-luna 

0.914 

0.902 

0.848 

0.775 

0.901 

gpt-5.6-sol 

0.886

0.902

0.750

0.743 

0.854

gpt-5.6-terra

0.857

0.902

0.750

0.794

0.854

DeepSeek-V4-Flash

0.775

0.902

0.800

0.848

0.854

DeepSeek-V4-Pro 

0.800

0.878 

0.800

0.789

0.902

grok-4-1-fast-reasoning

0.946

0.902

0.925

0.745

0.823


Across the production-style corpus, the composite reached 0.75-0.95 accuracy across the five rubric criteria. The results show model-specific strengths: Grok 4.1 Fast led Tool Call Accuracy and Tool Input Accuracy, DeepSeek V4 Flash led Tool Output Utilization, and DeepSeek V4 Pro led Tool Selection. Tool Call Success remained tied at 0.902 for most judges. Luna remained among the strongest overall, but no single judge dominated each of the rubric criteria.


Classification Accuracy 

On the truly multi-turn BFCL benchmark, Tool Call Accuracy is the only tool rubric criterion with native ground truth. GPT-5.6 Terra led the single-run panel at 0.864 macro-F1, followed by GPT-5.6 Sol at 0.858 and GPT-5.6 Luna at 0.828. Other models showed lower results, with Grok 4.1 Fast at 0.626, DeepSeek V4 Pro at 0.625, and V4 Flash at 0.521. This shows that with the correct model, the composite evaluator can classify overall tool-call correctness on full conversations as well as production-style traces. 


Failure Localization 

The same BFCL run also tested where failures occurred. GPT-5.6 Terra led at 0.78 member-any failed-turn localization: in 78% of failing conversations, the turn predicted by the Tool Call Accuracy rubric criterion matched at least one turn in BFCL’s set of failing turns. It produced 22 false alarms across 75 passing conversations. GPT-5.6 Sol reached 0.75 member-any, and GPT-5.6 Luna reached 0.74. The Tool Selection channel with Luna provided a high-precision secondary signal, identifying a failing turn in 50% of failures with only 3 false alarms. 


To test whether localization generalizes beyond mechanically checked BFCL traces, we also used the 29-row AgentRx tau_retail set with human failure annotations. On the eight in-scope tool-mechanics failures, Sol localized 8 of 8 root causes, while Terra and Luna localized 7 of 8. Meanwhile, Grok 4.1 Fast only localized 2/8, and the two DeepSeek models 1/8. This result is directional because of the small sample. 


Does Batching Preserve Rubric Relationships? 


After testing direct accuracy, we examined whether batching also preserves how rubric verdicts relate to one another. This is supporting structural evidence, not an accuracy measure against ground truth.


Side-by-side heatmaps comparing individual and composite rubric correlations

The heatmaps compare Spearman correlations between GPT-5.6 Luna’s binarized rubric verdicts on rows scored by both modes. The tool use structure was especially stable, including the strongest relationship between Tool Call Accuracy and Tool Input Accuracy (0.76 in both modes). Quality relationships were more mixed. For example, Single-Turn Intent Resolution and Task Completion fell from 0.71 to 0.31, while Multi-Turn Task Adherence and Task Completion fell from 0.53 to 0.23. This shows that batching can preserve the broad rubric structure without preserving every relationship equally. 


Finding 3: Reliability Depends on Both Rubric and Judge


Multi-Turn Quality Reliability 


We repeated the six-evaluator composite call four times on the same 200 BFCL rows. Across the nine-judge panel, DeepSeek V4 Flash had the lowest average flip rate at 3.9%. GPT-5.6 Sol and Grok 4.1 Fast followed at 5.2%, DeepSeek V4 Pro at 5.6%, GPT-5.6 Luna at 6.8%, GPT-4o at 7.4%, GPT-5.6 Terra at 7.8%, GPT-5.4 mini at 10.9%, and GPT-5.4 at 15.6%. 


The rubric criteria-level result is more important than the leaderboard: 



  • Fluency: 0.0% average flip rate

  • Coherence: 2.4%

  • Intent Resolution: 8.6%

  • Task Adherence: 6.7%

  • Task Completion: 11.5%

  • Groundedness: 16.4% 


The same judge can be highly stable on one rubric criterion and noisy on another. Reliability should therefore be monitored at the criterion level, not inferred from a single aggregate score. 


Multi-Turn Tool Use Reliability 


The five-evaluator MT tool study showed the same need for criterion-level monitoring. Across four BFCL repeats, GPT-5.6 Terra and DeepSeek V4 Flash had the lowest average flip rate at 6.0%, GPT-5.6 Sol reached 6.1%, and GPT-5.6 Luna and GPT-4o were at 7.0%. GPT-5.6 Sol achieved the strongest mode-of-four Tool Call Accuracy kappa at 0.759, while Terra led the single-run quality/localization results. Tool Call Success was the most stable criterion, with a 4.7% average flip rate across judges. 


Practical Recommendations 


Based on the results across quality, tool-use, localization, and reliability studies, we can derive practical guidance for both judge selection and evaluator batching.


Recommended Judges: The GPT-5.6 Family


Across the quality, tool-use, localization, and reliability experiments, the GPT-5.6 family consistently delivered the strongest overall results. While individual benchmarks occasionally favored other models or specific GPT-5.6 variants, Luna, Sol, and Terra repeatedly ranked among the top performers and showed the most balanced performance across workloads. Sol was generally the most reliable across repeated quality and tool-use studies, Terra led several BFCL tool-use and failed-turn localization evaluations, and Luna achieved the highest average quality across the tested workloads. For workloads similar to those evaluated here, we recommend starting with GPT-5.6 judges and selecting between Sol, Terra, and Luna based on your specific priorities. When available at a lower deployment cost, Luna is a particularly attractive option because it maintained frontier-level evaluation quality while achieving the strongest average performance in our experiments. By comparison, lower-cost models often involved larger tradeoffs between quality, reliability, and benchmark performance. Exact cost savings depend on model pricing and deployment configuration, so benchmark candidate judges on production-like workloads before optimizing solely for cost.


Which Evaluators to Batch 


These evaluators performed as well in the composite evaluator as they did individually — they’re strong candidates for batching right away: 



  • Fluency and Coherence — highly stable across all nine judges and repeats 

  • Tool Call Success and Tool Selection — robust across the tested judge models 

  • Task Completion — close to individual evaluation across benchmarks; slightly stronger with GPT-5.6 Luna on BFCL 

  • Tool Call Accuracy, Tool Input Accuracy, and Tool Output Utilization — strong on production-style traces, with accuracy ranging from 0.75 to 0.95 depending on evaluator and judge 


Which Evaluators to Watch 


These evaluators showed workload- or judge-dependent results. They can still be batched, but we recommend checking the results on your own data: 



  • Task Adherence — performed well on typical conversations but struggled when a single request contained many independent constraints 

  • Intent Resolution — results varied across datasets even though agreement was strong in some cases 

  • Groundedness — the least consistent evaluator in our tests, with scores changing across repeated runs more often than any other dimension, albeit having better performance on frontier models, like GPT-5.6 Luna and GPT-5.4. If grounding accuracy is critical for your use case, consider keeping the individual evaluator as a fallback 


Where This Approach Fits 


Composite evaluation is most effective as a selective default rather than a universal replacement for every standalone evaluator. GPT-5.6 Sol offered the strongest overall balance, Terra led the single-run BFCL tool results, and Luna is the recommended cheaper high-quality option. Judge and rubric selection should therefore be validated on production-like data. 


The Multi-Turn comparison has one structural boundary: native individual evaluators do not exist for Multi-Turn Fluency and Intent Resolution. The composite evaluator can produce both dimensions, but a direct individual-versus-composite validation is not available for them in that setting. 


Finally, token savings do not translate to one fixed currency percentage. Pricing varies by model and deployment, while output length and retry behavior vary by workload.


Takeaway 


Composite evaluators are a practical way to remove repeated context from evaluation pipelines. In the updated 100-row comparisons, they reduced total input-token use by 61.25% for six quality evaluators and 71.07% for five tool use evaluators, reducing completion tokens by 46.63–63.89%, while reducing measured run wall time by 35.74% and 46.10%, respectively, and preserving comparable quality across many dimensions. 


The strongest deployment pattern is selective rather than absolute: start with a model from the GPT-5.6 family; batch the evaluators that remain stable, monitor them independently, and keep focused fallbacks for the evaluators that do not. 


Composite evaluation is ultimately about removing unnecessary work. By evaluating multiple dimensions in a single judge call, teams can dramatically reduce repeated context processing while maintaining the evaluation coverage needed to monitor production systems.


Get Started


Start with the composite evaluator on a small set of known-good and known-bad agent conversations. Choose Output Quality to assess the six quality criteria or Tool Use Quality to assess the five tool use criteria. Review each criterion’s score and applicability status, then spot-check the results against individual evaluators for the failure modes that matter most to your application.


The GPT-5.6 family is the recommended starting point for judge models. Re-check results when you change the judge model or the domains of conversations being evaluated.


Try these evaluators in the Microsoft Foundry portal now. You can explore the docs on how to use them Agent Evaluators for Generative AI – Microsoft Foundry | Microsoft Learn.


 

Microsoft Tech Community originally posted this article on 2 October 2026 at 5:00 PM.

Leave a Reply