We value your privacy

We use cookies to analyse site traffic and improve your experience. You can accept all cookies or reject non-essential ones. See our Privacy Policy for details.

GR IT SERVICES
  • Contact
Get a quote
  1. Audit and compliance
  2. Data discovery and classification audit
Data discovery and classification audit, UAE

You cannot protect data proportionately until you know which of it matters. Most organisations protect all of it identically, which means none of it well.

A data audit finds where sensitive information actually lives, who can reach it, what is classified, what is overshared and what is leaving. It is the prerequisite for every data control that follows, and it is the work most organisations skip on the way to buying the tooling.

Book a data discovery auditSee what the audit finds
Data discovery and classification audit for UAE organisations
  • Where it isIncluding the places nobody expected
  • Who can reach itOversharing, quantified rather than assumed
  • What is classifiedAnd what carries no label at all
  • CIS Control 3Data protection, in the eighteen
What the audit establishes

Eight questions, and the first four have to be answered before any tooling is bought.

Data protection is control 3 in the CIS Critical Security Controls at version 8.1, and it sits after inventory for the same reason: you cannot protect what you have not located. The audit answers where the data is, what it is, who can reach it and what is happening to it.

Where sensitive information actually lives

Across SharePoint and OneDrive, mailboxes, Teams, file shares, databases, cloud storage, endpoints and whatever a department set up for a project. The result is reliably wider than expected, because data propagates through export, attachment and convenience rather than through architecture, and every copy carries the same obligations as the original.

What is classified, and what carries nothing at all

Sensitive information types and trainable classifiers identify content by pattern and by trained recognition, and the results surface in reporting and activity explorer. The audit distinguishes what has been labelled deliberately, what has been labelled automatically, and the substantial majority in most organisations that carries no classification of any kind.

Oversharing, expressed as a number rather than a worry

Sites shared with everyone in the organisation, links that grant access to anyone with the URL, permissions inherited from a structure nobody designed, and guests who retain access to material long after the engagement ended. Microsoft notes that generative AI amplifies the problem and risk of oversharing, which is why this finding has become urgent rather than theoretical.

Data outside the platforms anybody governs

File shares that predate the cloud migration, departmental databases, extracts sitting on endpoints, and copies in systems bought by a business function. These carry no labels, are covered by no policy, and are frequently where the most sensitive extracts live because somebody needed to work with them outside the system of record.

What protection is actually applied today

Which content has encryption applied through a label, which has protection applied without a label and therefore does not inherit to new items, and which has neither. That distinction matters operationally, because content protected without a label behaves differently from labelled content in almost every downstream system.

What is moving, and where to

Endpoint activity is visible before any policy is enforced. Microsoft notes that once devices are onboarded, information about audited activities flows into Activity Explorer even before any device-scoped policy exists. That free observation period tells you which egress paths people actually use, which is almost never the set anybody predicted.

Whether the classification scheme is usable

Most organisations have a scheme with four or five levels that nobody can apply consistently because the definitions overlap. The audit tests it by asking several people to classify the same twenty documents. Where they disagree, the scheme is the problem rather than the people, and no amount of automation fixes a taxonomy nobody can apply.

What arrives from outside, carrying somebody else labels

Material received from clients, partners or a parent company abroad. Microsoft states that endpoint data loss prevention cannot detect the sensitivity label from another tenant on a document, which means an incoming label does not carry across as a condition. Your own classification has to do the work, and the audit establishes how much incoming material that applies to.

Why this became urgent

Oversharing was survivable when finding anything required knowing it existed.

The permissions were always wrong. What changed is that something now searches them at conversational speed on behalf of anybody who asks.

  • Microsoft states it directly: because of the power and speed with which AI can surface content, generative AI amplifies the problem and risk of oversharing or leaking data.
  • A site shared with everyone in the organisation eight years ago was low risk because finding anything in it required knowing it existed and knowing what to search for. Neither is true when an assistant will summarise it in response to a plain question.
  • That is why the data audit and the AI readiness assessment are increasingly the same engagement. Organisations that deploy an assistant first discover the oversharing in the pilot and pause the rollout, which is expensive and avoidable.
  • It also reframes the priority. Classification has historically been sold as a compliance activity with a long horizon. It is now the prerequisite for a productivity programme with a board sponsor, which is a considerably easier conversation to fund.
Ask us to quantify your oversharing
How we approach it

Four things that make a data audit produce protection rather than a report.

Discovery tooling produces volume competently. What determines whether anything improves is the scheme, the prioritisation and whether the oversharing that matters is separated from the oversharing that does not.

We test the classification scheme before applying it

Give several people the same twenty documents and ask them to classify each. Where they disagree, the scheme is the problem, not the training. Most organisations need three levels with definitions written around what somebody would do with the document, rather than five levels defined by abstract degrees of sensitivity that nobody can distinguish.

We prioritise oversharing by content, not by breadth

A site shared with everybody in the organisation containing the staff handbook is not a finding. The same sharing on a site containing salary data is. Sorting exposure by what the content actually is, rather than by how widely it is shared, is what turns a list of thousands of sites into a list of the dozen that matter.

We use the free observation period before enforcing anything

Where endpoints are onboarded, audited activity flows into reporting before any device-scoped policy exists. That period costs nothing and tells you which egress paths people genuinely use. Designing controls from that evidence rather than from assumption is why the resulting policy set survives contact with users.

We look outside the platforms that report on themselves

Cloud platforms report on their own content well, which is exactly why an audit limited to them produces a comfortable picture. File shares that predate the migration, extracts on endpoints, departmental cloud services and supplier systems are where the least governed and frequently most sensitive material sits.

How we run it

Four phases, and the classification scheme gets simplified in the second.

The technical discovery is fast. The work that determines whether anything improves is agreeing a scheme people can apply and prioritising the oversharing that actually matters.
  1. 01
    Weeks 1 to 3

    Discover, without classifying anything yet

    Where content lives, what patterns and classifiers identify within it, how it is shared, and what is already labelled or protected. Endpoint activity observed where devices are onboarded, since audited activity flows into reporting before any policy is enforced and that period is the cheapest visibility available.

    • Content locations mapped, including outside governed platforms
    • Sensitive information found by type and by volume
    • Sharing and permission exposure quantified
    • Existing labelling and protection coverage measured
  2. 02
    Weeks 4 to 6

    Fix the scheme before applying it

    Test the existing classification scheme by having several people classify the same documents. Where they disagree, simplify. Most organisations need three levels rather than five, with definitions written around what somebody would do with the document rather than around abstract sensitivity.

    • The current scheme tested with real people and real documents
    • A simplified scheme with definitions people can apply
    • Mapping from the old scheme where one exists
    • Automatic classification rules for the unambiguous cases
  3. 03
    Weeks 7 to 12

    Remediate oversharing where it matters

    Prioritised by what the content is rather than by how widely it is shared, because a site shared with everyone containing the staff handbook is not the problem. Organisation-wide sharing and anyone links reviewed against the sensitive content actually found, and guest access recertified.

    • Highest-exposure locations remediated first
    • Anyone links and organisation-wide sharing reviewed
    • Guest access to sensitive material recertified
    • Permission inheritance corrected where structure was the cause
  4. 04
    Ongoing

    Make classification part of how work happens

    Automatic labelling for the unambiguous cases, manual labelling for the rest with the scheme people can actually apply, protection attached to labels rather than applied separately, and a periodic re-scan so the position is measured rather than assumed.

    • Automatic classification live for defined content types
    • Protection attached to labels rather than applied independently
    • Periodic re-scan with a reported trend
    • A named owner for the classification scheme itself
Where this matters most

Six UAE situations where the data audit is the necessary first step.

Almost every data control programme depends on knowing what exists and where. Buying the control first, and discovering afterwards, is the sequence that produces stalled deployments.

An organisation preparing to deploy an AI assistant

Microsoft states that generative AI amplifies the problem and risk of oversharing. The remediation work is data governance rather than AI work, and doing it first is the difference between a rollout that proceeds and one that pauses in the pilot after somebody surfaces a document they should not have seen.

A regulated firm asked where its regulated data is held

The question is straightforward and the answer requires an audit, because the system of record is the easy part and the copies are the hard part. Extracts, reports, attachments and working files carry identical obligations and sit outside every control that protects the original.

A healthcare organisation with patient information across several systems

Clinical systems are usually well controlled. What is less controlled is what leaves them: research extracts, reporting datasets, attachments in mailboxes and files on endpoints. Locating those and establishing who can reach them is where the actual exposure is, and it is rarely what anybody expects.

A professional services firm with client material in shared spaces

Engagement folders, shared channels and guest access that outlived the engagement. Because client material frequently carries contractual confidentiality obligations, the oversharing finding here is a contractual exposure as much as a security one, which usually accelerates the remediation considerably.

A business that has grown by acquisition

Each acquisition brings its own file shares, its own conventions and its own idea of what is sensitive. Establishing a single view across the group, using one scheme, is the only way a group function can hold a consistent position, and it is nearly always the first time anybody has attempted it.

An organisation whose labelling programme has stalled

A scheme was published, training was delivered, adoption is low and nobody can say how low. The audit measures coverage, tests whether the scheme is applicable in the first place, and identifies the content that can be classified automatically, which is usually a larger proportion than the organisation assumed.

Three positions

How UAE organisations know where their sensitive data is.

The middle column is common and it is a reasonable place to have reached. Labels exist, some people use them, and nobody has measured what proportion of the estate carries any classification at all.
Sensitive content located across the estate
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNo
Classification coverage measured
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNot applicable
Oversharing quantified
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNo
Content outside governed platforms included
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNo
Scheme tested with real people
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNot applicable
Automatic classification where unambiguous
Discovered and measuredYes
A scheme exists, adoption unknownSometimes
No classificationNo
Protection attached to labels
Discovered and measuredYes
A scheme exists, adoption unknownPartly
No classificationNo
Egress paths observed
Discovered and measuredYes
A scheme exists, adoption unknownNo
No classificationNo
Ready for an AI deployment
Discovered and measuredYes
A scheme exists, adoption unknownUnknown
No classificationNo
Answer for a regulator
Discovered and measuredEvidence
A scheme exists, adoption unknownPolicy
No classificationNone
Feature
Discovered and measured
A scheme exists, adoption unknown
No classification
Sensitive content located across the estate
YesNoNo
Classification coverage measured
YesNoNot applicable
Oversharing quantified
YesNoNo
Content outside governed platforms included
YesNoNo
Scheme tested with real people
YesNoNot applicable
Automatic classification where unambiguous
YesSometimesNo
Protection attached to labels
YesPartlyNo
Egress paths observed
YesNoNo
Ready for an AI deployment
YesUnknownNo
Answer for a regulator
EvidencePolicyNone
Where we look

Ten locations, and what each typically holds.

The audit covers all of these where they exist. The right hand column is what we most commonly find in each, which is our experience rather than a published figure.
LocationWhat it typically holds
SharePoint and OneDriveThe largest volume, the widest sharing, and the least classification
Exchange mailboxesAttachments that are the only remaining copy of something important
Microsoft TeamsFiles in channels nobody manages, inheriting site permissions
Legacy file sharesPermissions accumulated over a decade, and no classification at all
Databases and line of business systemsThe system of record, usually well controlled
Extracts and reportsCopies of the above, outside every control that protects the original
EndpointsWorking copies, downloads and anything somebody needed offline
Cloud storage outside the tenantDepartmental services procured to solve a specific problem
Third-party and supplier systemsYour data under somebody else controls
AI application prompts and responsesSensitive content pasted in, and increasingly visible
How an engagement runs

Five steps, and the scheme changes in the middle of it.

Typically six to twelve weeks depending on estate size. Discovery is quick. Fixing the scheme and remediating the oversharing that matters is where the time and the value are.
  1. 1

    Map where content lives, including outside governed platforms

    Cloud platforms, mailboxes, collaboration spaces, legacy file shares, databases, endpoints, departmental cloud services and supplier systems. An audit limited to the platforms that report on themselves produces a comfortable and incomplete picture, and the least governed material is reliably outside them.

  2. 2

    Identify what the content is, by type and by volume

    Sensitive information types and trainable classifiers applied across the located content, with results reported by type, by location and by volume. Existing classification and protection measured alongside, distinguishing content protected through a label from content protected separately, since the two behave differently downstream.

  3. 3

    Quantify who can reach it

    Organisation-wide sharing, anyone links, guest access, permission inheritance and the structural causes behind each. Then cross-referenced against what the content actually is, because exposure only matters in proportion to what is exposed and that cross-reference is what makes the list actionable.

  4. 4

    Test and simplify the classification scheme

    Several people classifying the same documents, and simplification where they disagree. Definitions written around what somebody would do with a document rather than around abstract sensitivity. Then automatic classification rules for the unambiguous cases, which is usually more content than anybody expects.

  5. 5

    Remediate by priority and set up the measurement

    Highest-exposure sensitive content first, with an owner against each location. Then protection attached to labels rather than applied separately, a periodic re-scan so coverage is measured rather than assumed, and a named owner for the scheme itself so it does not decay back to where it started.

Straight answers

What organisations ask about data discovery audits.

Because both depend on knowing what you have. A label scheme applied to content nobody has located covers a fraction of the estate. A data loss prevention policy tuned without knowing where sensitive content sits produces either no matches or an unmanageable number. The audit is what makes both of those investments land, and it is cheaper than either.

That is partly your decision and partly determined by obligation, and the audit surfaces both. Pattern-based sensitive information types find the identifiable categories such as identification numbers and payment card data. Trainable classifiers find content by type rather than by pattern. The remaining question, which is what your organisation considers commercially sensitive, is a business conversation the audit informs.

A combination. Scanning where the platform supports it, endpoint activity data where devices are onboarded, network and discovery evidence for cloud services procured outside IT, and interviews with business functions about the extracts and reports they work with. The last of those is unglamorous and consistently the most productive.

Content reachable by more people than anybody intended. Sites shared with everyone in the organisation, links that grant access to anyone with the URL, permission inheritance from a structure designed for something else, and guests who retain access long after their engagement ended. The finding that matters is the intersection of that exposure with content that is genuinely sensitive.

Because search speed changed. Microsoft states that generative AI amplifies the problem and risk of oversharing or leaking data, given the power and speed with which it can surface content. Permissions that were survivable when finding something required knowing it existed are not survivable when an assistant will summarise it on request.

Usually because it cannot be applied consistently. The test is simple: give several people the same twenty documents and ask each to classify them. Where they disagree substantially, the scheme is the problem rather than the training. Most organisations need three levels with practical definitions rather than five levels distinguished by degrees of abstract sensitivity.

More than most organisations assume, and never everything. Content matching clear patterns, content in defined locations serving a single purpose, and content types a trainable classifier recognises reliably can all be handled automatically. What cannot is the judgement content, which is where a scheme people can apply becomes necessary rather than optional.

Their labels do not help you. Microsoft states directly that endpoint data loss prevention cannot detect the sensitivity label from another tenant on a document. So material arriving from a client, a partner or a parent company abroad needs your own classification applied, and the audit establishes how much of your content that describes.

Where endpoints are onboarded, yes, and before any policy is enforced. Microsoft notes that once devices are onboarded, information about audited activities flows into reporting even before a device-scoped policy exists. That observation period is free, it shows which egress paths are actually used, and it is the best available input to designing controls people will accept.

Yes, and it is increasingly part of the picture. Sensitive information types and trainable classifiers can find sensitive data in user prompts and responses when people use AI applications, with the results surfacing in reporting and activity explorer. For most organisations this is the first visibility they have had of what staff are typing into these tools.

Six to twelve weeks for a full engagement including remediation of the highest-exposure findings. The discovery itself is a few weeks. Testing and simplifying the classification scheme, and getting owners to act on the oversharing findings, is what determines the overall timeline and is also where the durable value is.

The discovery does not, because it is read-only. Remediating oversharing does, and that is worth planning rather than absorbing. Removing organisation-wide sharing from a site people rely on produces immediate support calls unless the change is communicated and the legitimate access is preserved through a different mechanism first.

It happens in most audits and it is worth agreeing the handling in advance. Personal data in unexpected places, material subject to obligations nobody had applied, and occasionally content that should never have been retained at all. Agreeing who is told, in what order, and what happens next, before the discovery starts, avoids improvising during an uncomfortable week.

Automatic classification for the unambiguous cases so new content is classified as it is created, protection attached to labels rather than applied separately so it travels with the content, and a periodic re-scan producing a measured coverage figure. Without the last one you have a point-in-time picture that ages invisibly.

We scope by estate size and by how many locations outside the governed platforms are in scope, since that is where the effort concentrates. Where an AI deployment is driving the work, we usually scope the discovery and the highest-exposure remediation together, because that combination is what unblocks the rollout.
Before the audit

Fifteen questions the audit will answer, and you probably cannot today.

These are the questions that come up in a regulator conversation, a customer security review or an AI readiness assessment. Most organisations cannot answer any of them without looking.

Where and what

  • Where does regulated data live?
    All copies, not the system of record.
  • How much content carries no classification?
    Usually the substantial majority.
  • What is on file shares nobody migrated?
    Frequently the oldest and most sensitive.
  • What extracts exist outside the source system?
    Reports, exports and working copies.
  • What sits in departmental cloud services?
    Procured to solve a specific problem.

Who can reach it

  • How many sites are shared organisation-wide?
    And what is in them.
  • How many anyone links are active?
    And do they expire.
  • Which guests can still reach sensitive material?
    Long after the engagement ended.
  • Where is permission inheritance the cause?
    Structure rather than individual grants.
  • Who owns each high-exposure location?
    Remediation needs an owner.

Scheme and controls

  • Can two people classify the same document alike?
    If not, the scheme is the problem.
  • How many levels does your scheme have?
    Usually more than people can apply.
  • Is protection applied through labels?
    Or separately, which does not inherit.
  • What arrives already labelled by others?
    Those labels do not carry across.
  • Who owns the scheme?
    It decays without one.
Related reading

The pages around this one.

Sensitivity labels

The classification control the audit prepares the ground for.

Learn more

AI data security posture

The AI readiness question, which is largely the same work.

Learn more

Endpoint DLP

Controlling what leaves, designed from the observed egress paths.

Learn more
Next step

Give twenty documents to three people and ask them to classify each.

If they disagree, your scheme is the obstacle rather than your training, and no amount of automation will fix it. That test takes an hour and it usually reframes the whole programme.

Book a data discovery auditCall +971 56 613 2743

Related Services

Explore more solutions that work great with this service

Sensitivity Labels

Classification that travels with the file, and governs what Copilot sees

Learn more

AI Data Security Posture

Copilot readiness and control of shadow AI use

Learn more

Endpoint DLP

USB, print, clipboard and browser controls on devices

Learn more

Data Lifecycle Management

Retention policies, labels and defensible deletion

Learn more

UAE PDPL Compliance

Federal Decree-Law 45 of 2021 readiness and operations

Learn more

Microsoft Purview

Data governance and compliance solutions

Learn more

DLP Solutions

Microsoft Purview DLP and labels

Learn more

IT Audit Services Dubai

Assessment, technical test or certification, scoped properly

Learn more
GR IT SERVICES

Leading IT services provider in Dubai,
delivering enterprise-grade solutions
for businesses across the UAE.

Microsoft CSP PartnerCISGuard

Get the Helpdesk app

Raise and track IT tickets from your phone.

Download on the App StoreGet it on Google Play
Learn more about the app

Microsoft 365

  • Microsoft 365 Administration
  • M365 Reporting & Auditing
  • Microsoft 365 Licensing
  • Microsoft Copilot
  • Microsoft 365 Apps
  • Windows 365 Cloud PC
  • Microsoft SharePoint
  • Outlook & Exchange

Security

  • Microsoft Defender
  • Microsoft Purview
  • Microsoft Intune
  • Microsoft Entra
  • Compliance Manager
  • Cybersecurity Audits
  • Copilot for Security
  • Microsoft Sentinel
  • Microsoft Priva

Infrastructure

  • Google Workspace
  • Cloud Migration Services
  • Data Analytics & BI
  • Active Directory
  • Server Management
  • Apple Business
  • Apple Jamf Pro
  • IP Telephone
  • Data Backup
  • Website Development

IT Services

  • Managed IT Services
  • IT Support Dubai
  • IT AMC Dubai
  • New Office IT Setup
  • IT Relocation
  • Remote IT Support
  • On-Call IT Support
  • Startup IT Business Kit
  • Disaster Recovery & BC

Company

  • About Us
  • Careers
  • Contact
  • Blog

Contact

  • Iris Bay Tower, Office 903,
    Business Bay, Dubai, UAE
  • +971 56 613 2743
  • hello@gritservices.ae
  • gritservices.ae

© 2026 GR IT Services. All rights reserved.

Privacy PolicyTerms of UseCookie Policy