Co-authors and reviewers: Morteza Ziyadi, Hanchi Wang, Han Che, Billy Hu, Sean Gayler, Nishal Dsilva, Avinav Jami, Ankit Singhal, Augustus Arthur
TL;DR. Insights in Foundry turns recurring agent behavior into evidence-linked findings developers can review and act on. We evaluate trace linkage, finding quality, and detection of known issues using labeled benchmarks, LLM-judge assessment, and controlled end-to-end tests.
What is an Insight?
An Insight is a reviewable finding about recurring agent behavior. It brings together an explanation, supporting trace evidence, and a possible next step, helping developers investigate a pattern rather than inspect each execution in isolation. Depending on the available evidence and supported configuration, an Insight can include:
|
Part of an Insight |
What it gives the reviewer |
|
Title and description |
The recurring behavior and an evidence-based explanation of a possible cause. |
|
Linked traces |
The broader group of traces associated with the finding. |
|
Highlighted traces |
Representative examples to open and check against the explanation. |
|
Category, severity, and status |
Context for triaging the finding, not a substitute for a risk assessment. |
|
Agent version and recency |
Which version is represented and when the finding was created. |
|
Proposed action or fix |
An investigation or improvement path; concrete proposals depend on the configuration. |
Figure 1. An Insight detail view in Microsoft Foundry, cropped from public Microsoft Learn documentation. The description, evidence, and proposed fix support human review. This UI illustration is not a benchmark result; the source link provides a full-size view.
UI source and field definitions: Insights in Foundry documentation.
From trace search to an actionable review queue
Production agents can generate thousands of traces containing model calls, tool calls, latency, token usage, errors, and final responses. Observability shows what happened, while evaluations test criteria a team already knows to measure. The harder problem is discovering repeated behavior that the team did not know to predefine.
Insights in Foundry analyzes traces from the Application Insights resource connected to a Foundry project and organizes recurring behavior into reviewable Insights. Depending on the evidence and supported configuration, an Insight can include representative traces, affected agent versions, severity, an explanation of a possible cause, and a recommended next action. Concrete prompt or code proposals are available only for supported agent types and configurations; other Insights provide general investigation or remediation guidance.
Insights in Foundry is now available in public preview. It analyzes production traces to surface recurring behaviors and regressions with supporting evidence and suggested areas for investigation or improvement. Developers remain in control: review the cited traces and validate proposed changes through normal evaluation and deployment practices. During public preview, we will continue expanding the experience based on customer feedback.
To make this concrete, Figure 2 maps 220 inputs from TraceElephant, a public agent-trace benchmark, to seven generated Insights. Figure 3 examines one linked trace.
Figure 2. Recorded links from 220 public TraceElephant inputs to seven generated Insights in the September 17 benchmark. The 80-input finding includes the looping trace examined in Figure 3. Titles are editorial paraphrases; the layout is editorial, not a product screenshot or a validation of each diagnosis.
Figure 3. A public TraceElephant trace contains 54 model calls; steps 26, 30, 34, 38, 42, 46, and 50 repeat an identical extraction plan. The September 17 benchmark run of the production pipeline links this trace to a finding about missing progress-aware termination. Wording is paraphrased; this is an illustration, not product UI. The proposed intervention still needs evaluation.
Three complementary ways to evaluate Insight quality
Insight quality involves which inputs a finding links, how well it explains the evidence, and whether it identifies known issues in controlled tests. We combine labeled trace evaluation, unlabeled finding assessment, and controlled end-to-end testing. These are complementary sources of evidence, not mutually exclusive dataset categories.
|
Evidence |
Question |
Assessment / results |
|
Labeled trace evaluation |
Do Insights link inputs that reference annotations mark as failing? |
Reference labels; trace precision and trace recall (%). |
|
Unlabeled finding assessment |
Are generated findings grounded and useful? |
Eight-dimension LLM judge; mean score (1-5). |
|
Controlled end-to-end tests |
Do Insights identify known injected issues? |
Known issues and healthy baselines; detections, misses, unsupported findings, duplicates, and unscored cases. |
Here, ‘unlabeled’ means reference failure labels are not used for the evaluation; the sources can still contain annotations or reference answers. Public versus internal describes where data comes from, not how it is evaluated. LLM judges can also assess findings from labeled or controlled scenarios.
Labeled evaluation: Do Insights link the right traces?
- Trace precision: among unique benchmark inputs linked by at least one generated Insight, the fraction marked as failing. This measures discrimination only when a dataset includes healthy inputs, and it does not validate the diagnosis.
- Trace recall: the percentage of failure-labeled benchmark inputs linked by at least one generated Insight. It does not measure how many distinct failure modes were discovered.
We evaluated the production pipeline on the same 746 benchmark inputs (653 failure-labeled) on September 15, 16, and 17, 2026. Repeating these inputs does not create 2,238 distinct examples.
The six benchmark corpus slices come from public agent-trace research datasets with human reference annotations: AgentRx (Tau-bench retail), AgentRx Magentic-One, TRAIL, AgentErrorBench, TraceElephant, and the MAST-Data human subset. AgentErrorBench is the annotated failure-trajectory dataset released with AgentDebug, a framework for detecting and recovering from agent failures. The two AgentRx slices come from the same public release. The benchmark uses selected and normalized inputs from these releases, not a representative sample of customer production traffic.
For each dataset, the chart shows daily generated-Insight counts and trace recall. The table reports equal-weight means of daily trace precision and recall. These datasets support ongoing development; results depend on the dataset and model used.
Labeled results
Figure 4. Generated Insight count versus trace recall for all six public corpus slices. Each point is one daily run; labels retain values where markers overlap. Recall measures linkage to failure-labeled inputs, not diagnosis correctness. Three-day mean precision and recall follow in the table.
|
Dataset |
Inputs/day |
Failure-labeled/day |
Linked/day range |
Mean trace precision |
Mean trace recall |
|
AgentRx (Tau-bench retail) |
102 |
29 |
43-45 |
37.8% |
57.5% |
|
AgentRx Magentic-One |
58 |
44 |
54-57 |
75.3% |
94.7% |
|
TRAIL |
148 |
143 |
139-140 |
96.9% |
94.6% |
|
AgentErrorBench |
200 |
200 |
194-196 |
100.0%** |
97.3% |
|
TraceElephant |
220 |
220 |
218-220 |
100.0%** |
99.7% |
|
MAST-Data (human subset) |
18 |
17 |
14-16 |
93.3% |
82.4% |
** When every input is failure-labeled, 100% trace precision cannot measure false-positive control.
What the labeled results show
High precision can reflect the corpus base rate. AgentErrorBench and TraceElephant contain only failure-labeled inputs, so their measured precision is 100% by construction. TRAIL, AgentRx Magentic-One, and the MAST subset also contain mostly failures; their precision values are close to their underlying failure-label prevalence. Read precision beside trace recall and each dataset’s label prevalence rather than as a standalone quality score.
AgentRx (Tau-bench retail) is the clearest mixed-traffic stress test. Only 29 of 102 inputs were failure-labeled. Across the three days, trace precision was 43.2%, 37.8%, and 32.6%, averaging 37.8%; trace recall was 65.5%, 58.6%, and 48.3%, averaging 57.5%. Daily linked counts were 44, 45, and 43, including 19, 17, and 14 labeled failures respectively. These daily rates are averaged before rounding. Some linked inputs without failure labels could contain issues outside the reference labels, such as cost or latency, but this benchmark cannot confirm that.
Trace precision and trace recall measure whether Insights link inputs that reference annotations mark as failing. They do not establish whether an explanation is correct or a proposed action is useful. Our unlabeled evaluation examines those qualities using an eight-dimension LLM judge.
Unlabeled evaluation: Are the findings grounded and useful?
For this evaluation, a separate LLM judge scores generated Insights against the supplied input evidence, without using reference failure labels to compute precision or recall. These scores are automated assessments produced by an LLM judge. They are not human ratings and do not establish that a proposed action will improve the agent.
This evaluation covers PUPA, FailSafeQA, tau2-bench, synthetic scenarios, and an internal production dataset for the monitoring dashboard agent compiled from Microsoft employee usage. These sources are normalized into benchmark inputs; not every input is a complete execution trace.
The eight-dimension judge rubric
Each generated Insight receives a score from 1 to 5 on eight dimensions. The rubric separates the quality of the explanation, the importance of the issue, and the usefulness of the suggested action:
|
Dimension |
Question the judge considers |
|
Actionability |
Does the Insight tell a developer what to investigate, evaluate, or change next? |
|
Specificity |
Does it name concrete behavior, tools, or prompt elements rather than offer generic advice? |
|
Novelty |
Does it surface a pattern beyond an obvious dashboard signal? This is an estimate, not a measurement of what a team already knows. |
|
Correctness |
Is the claim supported by the supplied evidence? |
|
Severity calibration |
Does the assigned severity match the evidenced impact? |
|
Impact |
How consequential does the underlying issue appear, rather than just how well it is described? |
|
Fix specificity |
Does the proposed fix identify an exact asset or behavior to change? |
|
Fix applicability |
Is that proposed change plausibly on target for the issue? |
For example, a concrete next step earns a stronger actionability rating than generic advice. Severity calibration and impact are deliberately separate: a minor issue can have a well-calibrated severity label without being high-impact. Fix applicability is judged from the proposal and evidence; the fix is not executed as part of this scoring.
Unlabeled results
Results compiled from two recent benchmark runs cover 4,496 inputs and 40 generated Insights. Every emitted Insight in these selected snapshots was scored with the same judge rubric. Each Insight’s overall score is the mean of its eight dimension scores; the table then averages those overall scores within each dataset. We do not combine datasets into one quality score.
|
Dataset |
Inputs |
Insights scored |
Mean judge score |
|
PUPA |
901 |
5 |
3.53 |
|
FailSafeQA |
1,101 |
11 |
3.57 |
|
Monitoring dashboard agent (internal) |
1,342 |
4 |
3.88 |
|
tau2-bench |
1,112 |
16 |
2.84 |
|
Synthetic scenarios |
40 |
4 |
3.81 |
These are per-dataset snapshots, not a controlled comparison of dataset difficulty or pipeline versions. Small Insight counts and grading choices, such as grouped versus individual judging, make these protocol-specific diagnostics rather than stable absolute ratings. Automated judge scores are not precision or recall, human ratings, or measured improvements after applying a fix.
The dimension-level scores make the weak spots more concrete. Fix specificity was the lowest-scoring dimension in four of the five snapshots, with means from 2.00 to 2.64. For tau2-bench, correctness was lowest at 2.25: a signal to review whether claims are supported by the supplied evidence. Making a next step more concrete and grounding a diagnosis more carefully are different improvement targets.
What the public findings look like
The examples below connect generated findings to concrete input evidence: an arithmetic inconsistency, a question/answer mismatch, and a payment allocation that exceeds the available balance. They are selected illustrations, not a representative sample of the 40 scored Insights.
Figure 5. Three illustrative findings, with one checked input per finding. Titles and evidence summaries are editorial paraphrases. PUPA is a normalized QA record; FailSafeQA normalization pairs a perturbed question with the original answer, creating the displayed mismatch. Neither is a new agent execution. Tau2-bench shows a recorded tool interaction. These examples are not a representative quality sample.
Public input sources: PUPA; FailSafeQA; tau2-bench.
Controlled end-to-end tests: Can Insights find known issues?
Dataset benchmarks are complemented by tests with healthy baseline agents and versions containing predefined defects. The framework generates traffic, runs Insights on the resulting traces, and compares the findings with the known issues and supporting evidence. It separately tracks detections, misses, unsupported findings, and duplicates. Cases with incomplete evidence remain unscored rather than counting as passes or misses.
The September 23 hosted-agent report, used here as a concrete illustration, covered finance, travel, and support-ticket scenarios. Of 12 expected issues, 11 were scorable and 10 were detected. All three scorable healthy baselines had no confirmed unsupported findings. The figure shows the full outcome counts, including the missed and unscored cases. This end-to-end framework is separate from the 40-input synthetic-scenarios dataset in the unlabeled results.
Figure 6. Controlled tests exercise healthy baselines and known defects through the Insights pipeline. The September 23 report records 10 detections among 11 scorable expected issues, one additional unscored issue, and no confirmed unsupported findings on three scorable baselines. No confirmed unsupported findings (noise) or duplicate findings were reported in the evaluated subset. This illustrative snapshot is not a representative product-wide rate.
A broader view of quality
Trace selection and rubric review are part of a broader quality framework. Category agreement is an exploratory benchmark diagnostic, not an assessment of the portal’s category labels. Controlled scenarios also check for unsupported or duplicate findings.
Figure 7. Five complementary quality questions. The labeled results above measure trace precision and trace recall; the unlabeled results assess findings with an LLM judge. Category agreement is an internal diagnostic, while controlled scenarios check noise and duplication.
How we track quality over time
Labeled benchmarks, unlabeled assessment, and controlled tests feed recurring quality reports and human investigation. Test inventories and assessment conditions can change, so daily scores are not automatically a comparable improvement trend. The reported results do not establish longitudinal recurrence or deduplication performance across runs. In the portal, Give feedback lets users report incorrect findings, categories, severity, grouping, or duplication.
- Start with scenarios whose expected failures and healthy controls are known.
- Track detection, trace precision, categorization, severity, distinctness, evidence grounding, and action quality instead of relying on one score.
- Review changes to data, models, prompts, and evaluation contracts as changes to the measurement system.
- Investigate weak results and newly reported failure patterns rather than tuning only for an aggregate.
- Keep human review in the loop, because benchmark labels and automated graders cannot determine business impact or remediation correctness.
How practitioners should interpret an Insight
Validate each Insight before acting on it. Microsoft Learn recommends this review sequence:
- Confirm the affected workflow, agent version, category, severity, and time.
- Open the highlighted traces and verify that the cited behavior is present.
- Compare problematic examples with healthy traces to test whether the grouping and likely cause are plausible.
- Decide whether the issue belongs to the agent, a tool, a model endpoint, a data source, or the platform.
- Turn confirmed recurring behavior into evaluation coverage, an optimization objective, owner routing, or a monitored no-action decision.
An empty result does not prove that an agent is healthy, and a large linked-trace count does not prove business impact. Quality depends on representative traces, complete telemetry, a supported analysis model, and careful human validation.
Get started
Start with the Insights in Foundry documentation for prerequisites, portal steps, evidence-review guidance, SDK examples, pricing considerations, and preview limitations.
Prerequisites include a connected Application Insights resource, recent representative traces, a supported GPT-5-or-newer Judge model deployment, and the required role assignments. Insight generation, including scheduled runs, uses your model deployment and can incur model charges. See the current documentation for the latest supported agents, models, regions, limits, pricing, and UI guidance.
Python samples: on-demand Insights and scheduled Insights.
Closing thoughts
Agent quality work often starts with a simple question: what is repeatedly going wrong that we did not know to test? Insights in Foundry is designed to help teams answer that question with evidence-linked findings and a reviewable next step.
Measuring those findings requires explicit definitions, diverse data, honest treatment of weak results, and human judgment where automated metrics stop. Labeled trace evaluation, unlabeled rubric assessment, and controlled end-to-end tests answer complementary questions. Newly observed patterns can then inform the next round of evaluation and review.


