In April, I published a set of referential architecture diagrams for Microsoft Purview — Microsoft Purview Referential Architecture Diagrams. They have been useful in workshops and training sessions for a simple reason: one picture is worth a thousand words – attempting to explain the complex in a single image.
Today I’m publishing an update focused on classification and labeling — and on the wave of public previews that make auto-labeling something you can scale across the entire estate with ease.
➡️ Download the interactive architecture on GitHub here. It’s a self-contained HTML file – and it’s a lot more than this picture!
Once in GitHub, click on this icon on the top right bar to download
Why now: auto-labeling scales to the whole estate
The most consequential change is surprisingly small. An auto-labeling policy can now be kept on while you edit it, without re-triggering simulation.
Anyone who has run auto-labeling at scale knows why that matters. Until now, tuning a policy meant turning enforcement off, re-simulating, waiting, and turning it back on — delays every time you refined a rule. That friction made it challenging for organizations from broadening scope as fast as they wanted to. Removing it changes the operating model: you can simulate once with your conditions and expand scope (sites) when you’re ready, without hitting the simulation limit.
It arrives alongside a set of previews that lift every ceiling that used to cap the motion:
- Capacity is up 5x. Auto-labeling for SharePoint and OneDrive moves from 100,000 to up to 500,000 files per tenant per day (Roadmap ID 567890).
- Simulation goes from 4M to 20M items, so you can model against a real estate rather than a sample.
- Policy scope gets much bigger. Select up to 1,000 individual SharePoint sites per policy, adaptive scopes support up to 50,000 sites, and you can filter sites by the SiteTemplate property when configuring a policy (Roadmap ID 570445, MC1469958).
- Reporting finally tells you what happened. Audit and reporting now summarize the sensitive information types detected when labels are applied, and a new 30-day chart shows files processed per day so you can watch policy throughput.
- A per-policy coverage report correlates files processed during enforcement to the latest simulation run — including the files that enforcement processed but simulation never saw (Roadmap ID 568935).
- SharePoint library default labels now reach data at rest, applying the library’s default label to existing files instead of only new or edited ones (Roadmap ID 559105, MC1477181).
Read together, these are one story: classify and label the whole estate — unblocking your Microsoft 365 Copilot deployment by securing the data that matters to you most. The scalability improvements are on by default, and existing policies continue to function unmodified with no end-user impact.
What the architecture shows
The diagram follows content through five stages, each framed as the question it answers:
- User input — where information originates: files, email, Teams messages, files on devices, files in transit, and prompts.
- Information is classified — what is it? SITs, Exact Data Match, named entities, document fingerprinting, trainable classifiers, and AI classifiers, with automatic (in-transit) and on-demand (at-rest) classification as the two entry paths.
- Labeling files — what is the default security posture, and what did the user intend? Five labeling methods — default label from policy, client-side, library default, service-side auto-labeling, and manual — all converging on one node: the label is applied.
- Protection — what do we do about it? DLP (including DLP for Copilot), encryption and usage rights, permissions, label inheritance on Copilot responses, Adaptive Protection, and retention.
- Insights — what did we learn? DSPM, Activity and Content Explorer, the unified audit log, DLP alerts, and Insider Risk alerts.
How to use it – it’s interactive!
It’s a single self-contained HTML file — no dependencies or API call — it works offline.
- Click any node to light up its full upstream and downstream path; everything unrelated dims. The detail panel then shows what it feeds from, what it feeds into, what it sees, what it can do, and — the section I’d argue matters most — what it can’t do.
- Explain runs three guided walkthroughs: Classification vs labels, When is it classified?, and How is it protected? These are my go-to openers for a workshop.
- Isolate layer cuts the diagram by control plane — sensitivity label, SITs, classifiers, DLP, encryption, permissions, retention, audit.
- Licensing lens filters to Included in E3, E5 only — the delta over E3, or Pay-as-you-go. Use it to answer entitlement questions honestly: where Microsoft publishes no SKU, the node says so instead of guessing. Always confirm entitlements on a quote, not from a diagram.
- Overview only strips to titles for a clean read; Fit to width scales the whole thing to the window; light/dark follows your OS or the toggle; it prints cleanly to PDF.
- Keyboard: Tab between nodes, Enter to select, Esc to reset.
Auto-labeling at scale
If you’re standing up auto-labeling at scale, confirm the prerequisites node first — sensitivity labels enabled for Office files in SharePoint and OneDrive, PDF support enabled separately via EnableSensitivityLabelForPDF, auditing on (simulation requires it), labels published to at least one user, and label scope covering Files and/or Emails.
Simulate on your priority sites, review, enforce, then expand scope to more sites or all sites and keep going. Combine that with capabilities already GA — labeling files contextually by site + file types, and overriding manually applied labels — and you can finally clear the backlog of unlabeled files. note: more to come soon for a full deployment guide.
Feedback welcome — if a node is wrong, missing a limitation, or missing a link, tell me in the comments and I’ll fix it in the next revision.


