Data Governance and Compliance
Know where the sensitive data is, who can reach it, and what left
Most organisations cannot answer three questions about their own data: where the personal information is, who can currently reach it, and who actually did. Those are the questions a regulator asks after an incident, and they are the ones a notifiable data breach assessment depends on. This work is about being able to answer them from a system rather than from memory. It starts with discovery and classification, because you cannot govern what you have not found, and it ends with evidence you can produce on a deadline.
Also called: Microsoft Purview consulting · Data classification and discovery · Data security posture management (DSPM) · Information governance · Data catalogue and lineage · Privacy Act and APP readiness
Choose your platform
Business outcomes
Discovery, classification, access control and audit evidence over the data you already hold, designed against the obligations that actually apply to you.
What changes for the business, and how we agree to measure it before work starts.
- An inventory of where sensitive data actually is
- Automated discovery and classification across the estate, so personal, health and commercially sensitive information is located rather than assumed. Most engagements find it in at least one place nobody expected.Agreed measure: Proportion of in-scope repositories scanned and classified, with the unscanned remainder listed and explained rather than omitted.
- Access you can state precisely
- Entitlements resolved down to who can reach which dataset, including inherited and standing access, so the answer to a permission question is a query rather than an investigation.Agreed measure: Whether effective access to a named sensitive dataset can be produced on request, tested on a sample at handover.
- A breach assessment you could actually complete
- Audit trails retained long enough and in enough detail to reconstruct who accessed which records. Without that, the assessment a notifiable breach requires cannot honestly be finished inside the window.Agreed measure: Whether access history for a named dataset can be reconstructed for the retention period agreed with your privacy officer.
- Retention that matches the purpose
- Data kept for as long as it is needed and not longer, which is both a privacy obligation and the cheapest way to shrink the impact of a future incident.Agreed measure: Share of in-scope repositories with an applied retention rule and a named owner for exceptions.
- A catalogue people use rather than a wiki nobody opens
- Data products with named owners, definitions and quality signals, published where analysts already work. A catalogue without ownership is documentation, not governance.Agreed measure: Proportion of published data products with a named owner and a definition agreed by the business, not by IT alone.
- Evidence assembled rather than reconstructed
- Control evidence produced from the platform on request, so an audit or a client security questionnaire stops consuming a fortnight of somebody's time.Agreed measure: Turnaround to produce an agreed evidence pack, baselined against your current manual process.
Common client problems
What we usually hear first.
These are the sentences that start most engagements, and what we do about each one.
We hold sensitive information and we honestly cannot say where all of it is.
That is the normal starting position and it is the right thing to fix first. We run discovery and classification across the repositories in scope, then report what was found, what could not be scanned and why. The uncomfortable part is usually not the volume, it is finding personal information in a place with no owner and broad access, such as an old file share or an export somebody made for a project in 2021.
Can you make us compliant with the Privacy Act?
No, and neither can anyone else selling you technology. Compliance is an organisational position involving policy, process, training and legal interpretation, and we are not your lawyers. What we can do is build the technical capability the obligations depend on: knowing where personal information is, controlling and reviewing access, applying retention, and being able to evidence all three. We work with your privacy officer rather than in place of them.
Our board wants to know if we are Essential Eight compliant.
Essential Eight is assessed against maturity levels rather than as a pass or fail, so the honest answer is a maturity position with gaps, not a yes. We can assess the cloud-side controls and design uplift against the level you are targeting. We would also flag that Essential Eight is largely about endpoints, patching and administrative privilege, so a data governance programme addresses part of it and the rest sits elsewhere in your organisation.
We already bought Microsoft 365 E5. Do we not already have this?
You have a lot of the licensing, and almost certainly not the configuration. The gap between owning Purview capability and having it usefully applied is where most of the work is: label taxonomy agreed with the business, policies tuned so they do not fire constantly, scanning scoped and paid for, and someone owning the output. We will tell you what your existing licences already cover before recommending anything additional.
Our last classification project produced labels nobody used.
Almost always because IT designed the taxonomy alone. Labels that do not match how the business talks about its own information get ignored, and policies that fire too often get overridden by administrators, which is worse than no policy because it manufactures a false record of control. We keep the taxonomy small, design it with the people who create the documents, and start in audit mode so you can see the noise before anyone is blocked.
We are on AWS. Is there a Purview equivalent?
Not a single product, and pretending otherwise is how AWS governance programmes get underscoped. The same outcomes are assembled from a technical catalogue, a fine-grained access layer, a sensitive data discovery service and a configuration compliance service, each with its own permission model and console. It is achievable and it is genuinely more integration work for the same result. There is also no equivalent of protection that stays attached to a document after export.
Capabilities
What this domain covers.
- Sensitive data discovery and classification
- Data catalogue and lineage
- Sensitivity label taxonomy design
- Data loss prevention policy design
- Fine-grained access control on lake and warehouse data
- Access review and entitlement management
- Retention and lifecycle policy
- Audit trail design and retention
- Data residency assessment
- Privacy Act and APP technical readiness
- Essential Eight maturity assessment
- Control evidence automation
- Data quality and ownership frameworks
- AI data governance and Copilot readiness
How we deliver
From assessment through to the day we are still operating it.
- 01
Assessment and advisory
The assessment establishes which obligations genuinely apply to you, what you actually hold, and the gap between the two. It is deliberately done with your privacy officer rather than presented to them.
- Obligations scoped honestly. Which of the Privacy Act, sector regulation, contractual client requirements and internal policy genuinely bind you, since most organisations are told they need more than they do.
- Discovery across the real estate. Cloud stores, analytics platforms and the forgotten file shares, with anything that cannot be scanned listed rather than quietly excluded from the report.
- Effective access mapped. Who can currently reach each sensitive dataset including inherited and standing access, which is usually broader than anyone expects.
- Breach reconstruction tested. Whether existing audit retention would let you determine who accessed which records, checked against a real scenario rather than assumed.
- Residency verified, not assumed. Where data, metadata and logs physically sit, because metadata and log layers are the ones that quietly leave the country.
- Licence position costed first. What your existing entitlements already cover, since a technically correct design can be commercially unacceptable once the licences are priced.
- 02
Architecture and implementation
Discovery first, because classification drives everything downstream. Policy comes after the taxonomy is agreed with the business, and enforcement comes after a period in audit mode.
- Classification before control. Scanning and classification established first, because a policy written before you know what you hold protects the wrong things.
- A small taxonomy the business agreed. Few enough labels that people apply them correctly, worded the way the business already describes its own information.
- Audit mode before enforcement. Policies run in report-only first so the noise is visible and tuned before anyone is blocked and starts requesting overrides.
- Access bound to one identity plane. Data permissions resolved against the workforce directory rather than local accounts, so a leaver loses access everywhere at once.
- Ownership recorded per data product. A named business owner and definition attached to each published asset, since a catalogue without owners is documentation rather than governance.
- Audit retention set by scenario. Log retention chosen against the breach assessment and dispute windows you actually face, not left on a platform default.
- 03
Security and governance
This domain is where data governance and identity meet. The recurring failure is a control that exists in configuration but has quietly lapsed in effect, so verification matters more than deployment.
- Least privilege proven by test. Access verified by querying as accounts at several permission levels, including one that should see nothing, rather than by reading a policy.
- Standing access removed. Privileged data and compliance roles moved to approved, time-bound elevation so permanent access to sensitive data is the exception.
- Access reviews on a cycle that runs. Recurring reviews with a named reviewer and an escalation path, because a review nobody completes is worse than none at all.
- Encryption keys with a documented owner. Customer-managed keys for governed stores and evidence repositories, with key access separated from the teams being audited.
- AI access governed explicitly. What assistants and agents may read treated as a data governance decision, since an assistant surfaces oversharing faster than any audit.
- Residency pinned deliberately. Data, metadata and log locations constrained on purpose, because aggregation and scanning layers can move content outside Australian regions by default.
- 04
Adoption and enablement
Governance fails on adoption more than on technology. A label nobody applies and a review nobody completes both produce a record of control that is not real, which is more dangerous than an acknowledged gap.
- Taxonomy designed with the business. Workshops with the people who create the documents, because a taxonomy written by IT alone produces labels that go unused.
- Data owners actually appointed. Named individuals accepting ownership of specific data products, with the time commitment made explicit rather than assumed.
- Reviewers trained on the decision. Access reviewers shown how to judge whether access is still needed, so a review is a decision rather than a bulk approval.
- Policy tips explained before enforcement. Staff told what will be flagged and why ahead of enforcement, which reduces the override requests that hollow out a policy.
- Privacy officer inside the process. Involved from assessment rather than consulted at the end, since they own the obligation the technology is supposed to support.
- Evidence pack rehearsed. The audit or breach evidence process walked through once before it is needed for real, which is when the gaps become obvious.
- 05
Managed continuation
Classification drifts as new repositories appear, access accumulates as people move, and platform naming in this area changes often enough to matter. Governance that is not maintained becomes governance on paper.
- Discovery re-run on a cycle. New repositories and stores found and classified as they appear, since the estate grows faster than any one-off project's scope.
- Access review completion tracked. Not just whether reviews ran, but whether reviewers made real decisions, with bulk approvals treated as a signal to investigate.
- Policy override rates watched. Frequent overrides mean the policy is mistuned rather than that staff are careless, and a hollowed-out policy is a false control.
- Residency re-verified after change. Checked again whenever a service, region or aggregation setting changes, because this is where drift is least visible.
- Platform naming and status tracked. Product renames, converged features and retirements assessed before they affect a documented control or a client-facing statement.
- Evidence pack kept current. Regenerated periodically so the assembly process is known to still work, rather than discovered to be broken during an audit.
Questions we get asked
The things people ask before they commit.
Including the awkward ones. If the honest answer is that this is not right for you, that is the answer you will get.
What is data governance, as distinct from data security?
Data security keeps people out. Data governance decides who should be in, what the data means, who owns it and how long it is kept. In practice they overlap heavily and the same tooling often serves both, but the distinction matters commercially: a security programme buys you controls, while a governance programme buys you the ability to answer questions about your own data. Most organisations have more of the first than the second.
Will this make us compliant with the Privacy Act?
No. Compliance is an organisational position covering policy, process, training and legal interpretation, and no technology purchase delivers it. What this work delivers is the technical capability those obligations rely on: knowing where personal information sits, controlling and reviewing access, applying retention, and evidencing all three. We work alongside your privacy officer and legal advisers rather than substituting for them.
Is Celestique certified to assess us against ISO 27001 or IRAP?
No, and we would not claim otherwise. We hold no compliance certifications, and we are not an IRAP assessor. It is also worth knowing that IRAP assessors do not accredit, certify, endorse or register systems on behalf of the Australian Signals Directorate: an assessment produces findings and recommendations, and the authority to operate stays with the system owner. We design and build against these frameworks, and a certifying body is a separate engagement.
How long does a data discovery and classification project take?
Initial discovery across cloud repositories is usually a matter of weeks and produces useful findings early. What extends a programme is everything after: agreeing a label taxonomy with the business, tuning policies out of audit mode, appointing data owners and establishing review cycles. The technology is rarely the long pole. Organisations that treat this as a tooling deployment tend to finish with labels nobody applies.
What does it cost to run ongoing?
Three recurring lines, and the third is the one underestimated. First, licensing, which is genuinely hard to model because capability spans several products and add-ons. Second, scanning and classification consumption, which scales with how much data you hold and how often you rescan. Third, audit log retention, which grows quietly and is the line most often left on a platform default until someone reads the bill.
Can we not just do this in-house with our own data team?
Often yes, particularly if your data team already has the platform access and someone can own the business engagement. Where clients get value from help is the parts that are not tooling: scoping which obligations actually apply, designing a taxonomy that survives contact with the business, and knowing which vendor capabilities are settled versus mid-rename. If your constraint is capacity rather than knowledge, that is a staffing decision more than a consulting one.
Do we need this before we deploy Copilot or AI assistants?
Yes, and this is the most common sequencing mistake we see. An assistant reads with the permissions of the person asking, so it surfaces existing oversharing immediately and at scale. Deploying one across a poorly permissioned content estate does not create a new problem, it makes an old one visible in front of staff. Discovery, permission clean-up and labelling are the cheap prerequisite.
Related industries
Where this work has the most leverage.
Healthcare and Community Services
Administrative automation, policy search, workforce analytics and privacy uplift for healthcare and community providers, with clinical decisions left entirely to clinicians.
Professional Services
Governed enterprise search, document intelligence and secure copilots that respect matter confidentiality and conflict boundaries.
Manufacturing and Distribution
Demand and inventory intelligence, production analytics and supplier automation built on data your planners already trust.
Logistics and Warehousing
Shipment and document automation, proof-of-delivery processing and exception dashboards that let a small team run a large network.
Related services
What usually comes with it.
Security and Governance
Close the ways in, know exactly who can do what, and answer an auditor or a client questionnaire from current evidence rather than memory.
Data and Analytics
Reporting your executives trust, produced once, with a modelled layer underneath it that the next AI or forecasting project can stand on.
Artificial Intelligence
AI that is chosen for a reason, costed before it is built and governed once it is live.
Managed Services
Your platform keeps earning its business case after go-live, with cost, security posture, reliability and adoption reviewed on an agreed cycle rather than left to drift.
Free discovery workshop
Start with a data governance discovery workshop.
Bring one challenge in this area. We will map the opportunity, the readiness gaps and a recommended next step.