You cannot protect data proportionately until you know which of it matters. Most organisations protect all of it identically, which means none of it well.
A data audit finds where sensitive information actually lives, who can reach it, what is classified, what is overshared and what is leaving. It is the prerequisite for every data control that follows, and it is the work most organisations skip on the way to buying the tooling.

- Where it isIncluding the places nobody expected
- Who can reach itOversharing, quantified rather than assumed
- What is classifiedAnd what carries no label at all
- CIS Control 3Data protection, in the eighteen
Eight questions, and the first four have to be answered before any tooling is bought.
Where sensitive information actually lives
Across SharePoint and OneDrive, mailboxes, Teams, file shares, databases, cloud storage, endpoints and whatever a department set up for a project. The result is reliably wider than expected, because data propagates through export, attachment and convenience rather than through architecture, and every copy carries the same obligations as the original.
What is classified, and what carries nothing at all
Sensitive information types and trainable classifiers identify content by pattern and by trained recognition, and the results surface in reporting and activity explorer. The audit distinguishes what has been labelled deliberately, what has been labelled automatically, and the substantial majority in most organisations that carries no classification of any kind.
Oversharing, expressed as a number rather than a worry
Sites shared with everyone in the organisation, links that grant access to anyone with the URL, permissions inherited from a structure nobody designed, and guests who retain access to material long after the engagement ended. Microsoft notes that generative AI amplifies the problem and risk of oversharing, which is why this finding has become urgent rather than theoretical.
Data outside the platforms anybody governs
File shares that predate the cloud migration, departmental databases, extracts sitting on endpoints, and copies in systems bought by a business function. These carry no labels, are covered by no policy, and are frequently where the most sensitive extracts live because somebody needed to work with them outside the system of record.
What protection is actually applied today
Which content has encryption applied through a label, which has protection applied without a label and therefore does not inherit to new items, and which has neither. That distinction matters operationally, because content protected without a label behaves differently from labelled content in almost every downstream system.
What is moving, and where to
Endpoint activity is visible before any policy is enforced. Microsoft notes that once devices are onboarded, information about audited activities flows into Activity Explorer even before any device-scoped policy exists. That free observation period tells you which egress paths people actually use, which is almost never the set anybody predicted.
Whether the classification scheme is usable
Most organisations have a scheme with four or five levels that nobody can apply consistently because the definitions overlap. The audit tests it by asking several people to classify the same twenty documents. Where they disagree, the scheme is the problem rather than the people, and no amount of automation fixes a taxonomy nobody can apply.
What arrives from outside, carrying somebody else labels
Material received from clients, partners or a parent company abroad. Microsoft states that endpoint data loss prevention cannot detect the sensitivity label from another tenant on a document, which means an incoming label does not carry across as a condition. Your own classification has to do the work, and the audit establishes how much incoming material that applies to.
Oversharing was survivable when finding anything required knowing it existed.
The permissions were always wrong. What changed is that something now searches them at conversational speed on behalf of anybody who asks.
- Microsoft states it directly: because of the power and speed with which AI can surface content, generative AI amplifies the problem and risk of oversharing or leaking data.
- A site shared with everyone in the organisation eight years ago was low risk because finding anything in it required knowing it existed and knowing what to search for. Neither is true when an assistant will summarise it in response to a plain question.
- That is why the data audit and the AI readiness assessment are increasingly the same engagement. Organisations that deploy an assistant first discover the oversharing in the pilot and pause the rollout, which is expensive and avoidable.
- It also reframes the priority. Classification has historically been sold as a compliance activity with a long horizon. It is now the prerequisite for a productivity programme with a board sponsor, which is a considerably easier conversation to fund.
Four things that make a data audit produce protection rather than a report.
We test the classification scheme before applying it
Give several people the same twenty documents and ask them to classify each. Where they disagree, the scheme is the problem, not the training. Most organisations need three levels with definitions written around what somebody would do with the document, rather than five levels defined by abstract degrees of sensitivity that nobody can distinguish.
We prioritise oversharing by content, not by breadth
A site shared with everybody in the organisation containing the staff handbook is not a finding. The same sharing on a site containing salary data is. Sorting exposure by what the content actually is, rather than by how widely it is shared, is what turns a list of thousands of sites into a list of the dozen that matter.
We use the free observation period before enforcing anything
Where endpoints are onboarded, audited activity flows into reporting before any device-scoped policy exists. That period costs nothing and tells you which egress paths people genuinely use. Designing controls from that evidence rather than from assumption is why the resulting policy set survives contact with users.
We look outside the platforms that report on themselves
Cloud platforms report on their own content well, which is exactly why an audit limited to them produces a comfortable picture. File shares that predate the migration, extracts on endpoints, departmental cloud services and supplier systems are where the least governed and frequently most sensitive material sits.
Four phases, and the classification scheme gets simplified in the second.
- 01Weeks 1 to 3
Discover, without classifying anything yet
Where content lives, what patterns and classifiers identify within it, how it is shared, and what is already labelled or protected. Endpoint activity observed where devices are onboarded, since audited activity flows into reporting before any policy is enforced and that period is the cheapest visibility available.
- Content locations mapped, including outside governed platforms
- Sensitive information found by type and by volume
- Sharing and permission exposure quantified
- Existing labelling and protection coverage measured
- 02Weeks 4 to 6
Fix the scheme before applying it
Test the existing classification scheme by having several people classify the same documents. Where they disagree, simplify. Most organisations need three levels rather than five, with definitions written around what somebody would do with the document rather than around abstract sensitivity.
- The current scheme tested with real people and real documents
- A simplified scheme with definitions people can apply
- Mapping from the old scheme where one exists
- Automatic classification rules for the unambiguous cases
- 03Weeks 7 to 12
Remediate oversharing where it matters
Prioritised by what the content is rather than by how widely it is shared, because a site shared with everyone containing the staff handbook is not the problem. Organisation-wide sharing and anyone links reviewed against the sensitive content actually found, and guest access recertified.
- Highest-exposure locations remediated first
- Anyone links and organisation-wide sharing reviewed
- Guest access to sensitive material recertified
- Permission inheritance corrected where structure was the cause
- 04Ongoing
Make classification part of how work happens
Automatic labelling for the unambiguous cases, manual labelling for the rest with the scheme people can actually apply, protection attached to labels rather than applied separately, and a periodic re-scan so the position is measured rather than assumed.
- Automatic classification live for defined content types
- Protection attached to labels rather than applied independently
- Periodic re-scan with a reported trend
- A named owner for the classification scheme itself
Six UAE situations where the data audit is the necessary first step.
An organisation preparing to deploy an AI assistant
Microsoft states that generative AI amplifies the problem and risk of oversharing. The remediation work is data governance rather than AI work, and doing it first is the difference between a rollout that proceeds and one that pauses in the pilot after somebody surfaces a document they should not have seen.
A regulated firm asked where its regulated data is held
The question is straightforward and the answer requires an audit, because the system of record is the easy part and the copies are the hard part. Extracts, reports, attachments and working files carry identical obligations and sit outside every control that protects the original.
A healthcare organisation with patient information across several systems
Clinical systems are usually well controlled. What is less controlled is what leaves them: research extracts, reporting datasets, attachments in mailboxes and files on endpoints. Locating those and establishing who can reach them is where the actual exposure is, and it is rarely what anybody expects.
A professional services firm with client material in shared spaces
Engagement folders, shared channels and guest access that outlived the engagement. Because client material frequently carries contractual confidentiality obligations, the oversharing finding here is a contractual exposure as much as a security one, which usually accelerates the remediation considerably.
A business that has grown by acquisition
Each acquisition brings its own file shares, its own conventions and its own idea of what is sensitive. Establishing a single view across the group, using one scheme, is the only way a group function can hold a consistent position, and it is nearly always the first time anybody has attempted it.
An organisation whose labelling programme has stalled
A scheme was published, training was delivered, adoption is low and nobody can say how low. The audit measures coverage, tests whether the scheme is applicable in the first place, and identifies the content that can be classified automatically, which is usually a larger proportion than the organisation assumed.
How UAE organisations know where their sensitive data is.
| Feature | Discovered and measured | A scheme exists, adoption unknown | No classification |
|---|---|---|---|
Sensitive content located across the estate | Yes | No | No |
Classification coverage measured | Yes | No | Not applicable |
Oversharing quantified | Yes | No | No |
Content outside governed platforms included | Yes | No | No |
Scheme tested with real people | Yes | No | Not applicable |
Automatic classification where unambiguous | Yes | Sometimes | No |
Protection attached to labels | Yes | Partly | No |
Egress paths observed | Yes | No | No |
Ready for an AI deployment | Yes | Unknown | No |
Answer for a regulator | Evidence | Policy | None |
Ten locations, and what each typically holds.
| Location | What it typically holds | |
|---|---|---|
| SharePoint and OneDrive | The largest volume, the widest sharing, and the least classification | |
| Exchange mailboxes | Attachments that are the only remaining copy of something important | |
| Microsoft Teams | Files in channels nobody manages, inheriting site permissions | |
| Legacy file shares | Permissions accumulated over a decade, and no classification at all | |
| Databases and line of business systems | The system of record, usually well controlled | |
| Extracts and reports | Copies of the above, outside every control that protects the original | |
| Endpoints | Working copies, downloads and anything somebody needed offline | |
| Cloud storage outside the tenant | Departmental services procured to solve a specific problem | |
| Third-party and supplier systems | Your data under somebody else controls | |
| AI application prompts and responses | Sensitive content pasted in, and increasingly visible |
Five steps, and the scheme changes in the middle of it.
- 1
Map where content lives, including outside governed platforms
Cloud platforms, mailboxes, collaboration spaces, legacy file shares, databases, endpoints, departmental cloud services and supplier systems. An audit limited to the platforms that report on themselves produces a comfortable and incomplete picture, and the least governed material is reliably outside them.
- 2
Identify what the content is, by type and by volume
Sensitive information types and trainable classifiers applied across the located content, with results reported by type, by location and by volume. Existing classification and protection measured alongside, distinguishing content protected through a label from content protected separately, since the two behave differently downstream.
- 3
Quantify who can reach it
Organisation-wide sharing, anyone links, guest access, permission inheritance and the structural causes behind each. Then cross-referenced against what the content actually is, because exposure only matters in proportion to what is exposed and that cross-reference is what makes the list actionable.
- 4
Test and simplify the classification scheme
Several people classifying the same documents, and simplification where they disagree. Definitions written around what somebody would do with a document rather than around abstract sensitivity. Then automatic classification rules for the unambiguous cases, which is usually more content than anybody expects.
- 5
Remediate by priority and set up the measurement
Highest-exposure sensitive content first, with an owner against each location. Then protection attached to labels rather than applied separately, a periodic re-scan so coverage is measured rather than assumed, and a named owner for the scheme itself so it does not decay back to where it started.
What organisations ask about data discovery audits.
Fifteen questions the audit will answer, and you probably cannot today.
Where and what
- Where does regulated data live?All copies, not the system of record.
- How much content carries no classification?Usually the substantial majority.
- What is on file shares nobody migrated?Frequently the oldest and most sensitive.
- What extracts exist outside the source system?Reports, exports and working copies.
- What sits in departmental cloud services?Procured to solve a specific problem.
Who can reach it
- How many sites are shared organisation-wide?And what is in them.
- How many anyone links are active?And do they expire.
- Which guests can still reach sensitive material?Long after the engagement ended.
- Where is permission inheritance the cause?Structure rather than individual grants.
- Who owns each high-exposure location?Remediation needs an owner.
Scheme and controls
- Can two people classify the same document alike?If not, the scheme is the problem.
- How many levels does your scheme have?Usually more than people can apply.
- Is protection applied through labels?Or separately, which does not inherit.
- What arrives already labelled by others?Those labels do not carry across.
- Who owns the scheme?It decays without one.
Give twenty documents to three people and ask them to classify each.
If they disagree, your scheme is the obstacle rather than your training, and no amount of automation will fix it. That test takes an hour and it usually reframes the whole programme.
Related Services
Explore more solutions that work great with this service
Sensitivity Labels
Classification that travels with the file, and governs what Copilot sees
AI Data Security Posture
Copilot readiness and control of shadow AI use
Endpoint DLP
USB, print, clipboard and browser controls on devices
Data Lifecycle Management
Retention policies, labels and defensible deletion
UAE PDPL Compliance
Federal Decree-Law 45 of 2021 readiness and operations
Microsoft Purview
Data governance and compliance solutions
DLP Solutions
Microsoft Purview DLP and labels
IT Audit Services Dubai
Assessment, technical test or certification, scoped properly