Guidance Needed: Purview Scanning and Auto-Labeling for Dataverse with Power BI Reporting

Srikanth Yellapragada 0 Reputation points
2026-09-01T05:54:52.29+00:00

We have Microsoft Purview in place and have connected it to our Dataverse UAT environment.

However, there is currently a limitation with the UAT environment:

  • Dataverse UAT does not contain a proper/complete clone of the Production data.
  • The UAT environment contains only a minimal subset of data compared with Production.
  • I have scheduled an initial scan of Dataverse UAT in Microsoft Purview.
  • The scan completed successfully, and I can see the scanned data and the identified classifications in the Unified Catalog.
  • However, because UAT contains very limited data, the scan results do not represent the full range of data and classifications that we expect to have in Production.

    Where I'm stuck

    The challenge I’m facing is that UAT does not contain a representative copy of Production data, so I'm unsure how best to validate the complete Purview scanning, classification, and auto-labeling process. For example, if a particular type of sensitive data exists in Production but is not present in UAT, Purview obviously cannot discover or classify it during the UAT scan.

    Goal

    Our ultimate goal is to have the solution:
    1. Scan the whole Dataverse dataset.
    2. Identify sensitive information through automatic classification.
    3. Based on the classification, automatically apply the appropriate sensitivity labels.
    4. Eventually have this operating against the full Production dataset.
    5. Also how I can connect this to PowerBi to have a graphical representation of data that it scanned, classified, labelled, with other attributes.

    Where I'm stuck

    The challenge I’m facing is that UAT does not contain a representative copy of Production data, so I'm unsure how best to validate the complete Purview scanning, classification, and auto-labeling process. For example, if a particular type of sensitive data exists in Production but is not present in UAT, Purview obviously cannot discover or classify it during the UAT scan.

    Questions

    What would be the recommended approach from here? Should we:
  • Use the existing UAT data and validate the Purview configuration based on the classifications that are currently detected?
  • Obtain/restore a more representative sanitized or masked copy of Production data into UAT and perform the scan against that?
  • Configure the classification and sensitivity-labeling policies in UAT and then validate them against Production?
  • Or is there a recommended Microsoft Purview approach for testing complete classification + auto-labeling when the UAT environment does not contain a full Production data set? Most importantly, I’d like to understand what should be configured/tested in UAT versus what should ultimately be enabled in Production, and how we can confidently validate that the entire Production Dataverse data set will be scanned, classified, and appropriately auto-labeled.
Microsoft Security | Microsoft Purview

1 answer

Sort by: Most helpful
  1. AI answer

    2026-09-01T05:55:15.66+00:00

    Use UAT to validate the configuration and process, not to prove coverage of sensitive data that does not exist there.

    Recommended approach:

    1. Keep the current UAT scan and validate the scan pipeline end to end
      • In Purview, confirm the Dataverse source is registered, the scan runs successfully, and the assets appear in the Unified Catalog.
      • For the Dataverse scan, use Data Map > Sources, select the Dataverse source, create or review the scan, choose the correct integration runtime, credential, collection, run Test connection, and then Save and run.
      • After the run, check Last run status, scanned asset counts, and classified asset counts.
      • This validates connectivity, scan scope, scheduling, and catalog ingestion.
    2. Configure classification deliberately before relying on scan results
      • In Purview Data Map, classifications can be assigned automatically through the scan rule set or manually after metadata ingestion.
      • Before scanning, define which classifications are actually needed. Microsoft recommends avoiding unnecessary labels because they create noise and can mislead data consumers.
      • If the goal is to detect more sensitive patterns, configure the appropriate scan rule set with relevant system or custom classifications before the next scan.
    3. Do not expect UAT to validate discovery of data patterns that are not present
      • Purview classification depends on the data that is scanned. If a sensitive data type is absent from UAT, it will not be discovered there.
      • For complete validation of classification outcomes, a representative sanitized or masked production-like dataset in UAT is the practical approach, because it allows the scan to encounter the same kinds of values and structures expected in production.
    4. Separate metadata classification in Purview from protective labeling
      • For scanned data sources in Purview, such as registered sources in the Data Map, Purview can apply metadata-only sensitivity labels to catalog entries.
      • These labels do not encrypt or restrict the source data. They indicate sensitivity in the catalog and can support governance workflows.
      • For Microsoft 365 data, protective labels are applied through Microsoft Purview Information Protection.
      • The context provided does not establish that Dataverse scan classification directly auto-applies protective sensitivity labels back into Dataverse records. What is supported here is classification in Data Map and metadata labeling of scanned assets.
    5. Plan production validation as a controlled rollout
      • Use UAT to validate:
        • source registration
        • scan scope
        • credentials and connectivity
        • collection placement
        • scan schedules
        • scan rule sets and expected classifications on available test data
        • catalog visibility of assets and classifications
      • Use a production-like sanitized dataset in UAT, if possible, to validate broader classification coverage.
      • Then enable scanning on the full production Dataverse source and review the resulting asset and classification coverage in the Unified Catalog.
    6. Keep classification current with recurring scans
      • Classification should be maintained operationally. Microsoft recommends regular review because stale metadata can lead to incorrect outcomes.
      • For Dataverse, set up scheduled scans so schema and content changes are re-evaluated over time.
    7. Power BI reporting
      • The provided context supports that scan results, assets, and classifications are visible in the Unified Catalog and that Dataverse assets can be browsed and searched there.
      • The provided context does not contain a documented method for connecting Purview scan/classification/label results directly to Power BI for graphical reporting.

    Practical decision for this scenario:

    • Yes: use the current UAT environment to validate the Purview setup and scan behavior.
    • Best next step: obtain a sanitized or masked production-like copy for UAT if complete classification validation is required before production.
    • Then: configure and test scan rule sets and metadata labeling in UAT, and finally run the full scan in Production to validate actual coverage against the real dataset.

    This is the clearest split between environments:

    • UAT: validate configuration, scan execution, rule sets, and expected classification behavior on available or sanitized representative data.
    • Production: validate actual enterprise-wide discovery and final classification coverage on the full Dataverse dataset.

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.