Skip to main content

Data and Analytics

One set of numbers, produced once, that the business stops re-checking

Most organisations do not have a data shortage. They have one number reported three different ways, a month-end that runs on a spreadsheet somebody maintains privately, and no agreed owner for the definitions everyone argues about in the meeting. We fix the foundation first, then the modelled layer, then the reporting and forecasting on top of it, so analytics becomes something the business acts on rather than re-checks. The order matters: almost every platform we are asked to rescue was built the other way round.

Also called: Data platform consulting · Modern data platform · Data warehouse modernisation · Lakehouse implementation · Business intelligence consulting · Data engineering and pipelines · Analytics modernisation

Choose your platform

Business outcomes

Reporting your executives trust, produced once, with a modelled layer underneath it that the next AI or forecasting project can stand on.

What changes for the business, and how we agree to measure it before work starts.

One agreed set of operational numbers
Finance, operations and the executive team work from the same definitions of revenue, margin, utilisation and stock position instead of reconciling three versions live in the meeting. The definition is implemented once, in a shared model, rather than restated in each report.Agreed measure: Number of reported metrics with a single owner, a written definition and a traceable source, agreed at assessment.
The reporting cycle stops eating the week
The people who currently rebuild the same workbook every period get that time back for analysis and exception handling. The output stays familiar to its audience while the process stops depending on one laptop and one person's memory.Agreed measure: Hours spent on manual preparation each period, baselined with the team who do the work before anything is built.
Numbers that arrive before the decision
Operations and project managers see yesterday's production, dispatch or site position at the start of the day rather than waiting for a monthly pack. Only the measures that genuinely need to be current get built that way, because currency is the expensive part.Agreed measure: Data latency per measure: how old a number is at the moment someone acts on it, with the target set per measure before build.
Forecasts that can be argued with
Demand, cash and resourcing forecasts come from a documented model with visible assumptions, so a planner can challenge one input instead of distrusting the whole output. A forecast nobody can interrogate gets overridden by instinct within a month.Agreed measure: Forecast error tracked against actuals over agreed periods and reported openly, including the periods it misses.
A foundation an AI project can stand on
When the business is ready for assistants, extraction or prediction, the underlying data is already modelled, permissioned and documented, which removes the most common reason pilots stall between demonstration and production.Agreed measure: Readiness scored against a checklist agreed at assessment covering ownership, lineage, sensitivity and access control.
Platform cost you can attribute
Platform owners see capacity, storage and query cost broken down by workload and business unit, so an over-refreshed model or an expensive dashboard becomes a conversation rather than a surprise on the invoice.Agreed measure: Monthly cost split by workload and business unit, with the utilisation threshold that triggers a review agreed up front.

Common client problems

What we usually hear first.

These are the sentences that start most engagements, and what we do about each one.

  • Every department brings a different number to the same meeting and we spend the first 20 minutes arguing about whose is right.

    We trace each contested metric back to its source field, write down the definition the business will actually agree to, and implement it once in a shared semantic model. Reports then reference that model instead of each team's private calculation. The hard part is not technical: it is getting two executives to accept one definition of margin, and that conversation is part of the engagement rather than a prerequisite for it.

  • Our month-end runs on a spreadsheet that one person maintains, and she is going on long service leave.

    We work through the workbook with the person who owns it and rebuild its logic as documented, version-controlled transformations in the platform. The output stays familiar to the audience while the process stops depending on one laptop. Expect to find undocumented adjustments in there that turn out to be load-bearing, because they always are.

  • We bought a BI tool two years ago and people still export everything to Excel.

    That usually means the model does not answer the question people actually have, or they do not trust it yet. We sit with the heaviest exporters, work out what each export is really for, then rebuild the model and the measures around those decisions before anyone touches a visual. Some exports are legitimate and stay: an analyst doing genuine ad hoc work in Excel connected to a governed model is a good outcome, not a failure.

  • We have a data lake and nobody can tell me what is in it or who is allowed to see it.

    We establish which parts of it a live workload actually reads, assign ownership and a sensitivity position to those, and archive or delete the datasets nobody can justify keeping. Access is then granted through identity groups tied to roles, so who can read a table becomes a configuration item rather than a guess. The deeper cataloguing and classification programme is a separate piece of work and we scope it as one.

  • Our reports are always a month behind, so by the time we see a problem the money is already spent.

    We separate the small number of measures that genuinely need to be current from the many that are fine monthly. Streaming or incremental ingestion is applied only to that subset, which keeps the cost proportionate to the value of knowing sooner. Real-time everything is the most common way a data platform budget gets spent on nothing anyone reads.

  • We want to do something with AI but we keep being told our data is not ready and nobody explains what that means.

    We turn readiness into a specific, testable list: which entities are modelled, where lineage breaks, which fields carry personal information and what access control exists today. You get a remediation plan with sequencing and effort attached, not a verdict. Sometimes the honest finding is that the use case you want needs a well-organised document set rather than a warehouse, in which case we say so and the data work waits.

Capabilities

What this domain covers.

  • Data platform strategy and target-state architecture
  • Data estate and AI-readiness assessment
  • Lakehouse and data warehouse design
  • Data warehouse modernisation and migration
  • Pipeline engineering and orchestration
  • Dimensional and semantic modelling
  • Self-service business intelligence enablement
  • Real-time and streaming analytics
  • Data quality rules and reconciliation testing
  • Master and reference data alignment
  • Lineage and downstream impact analysis
  • Forecasting and predictive modelling
  • Row-level and column-level access design
  • Capacity, storage and query cost optimisation

How we deliver

From assessment through to the day we are still operating it.

  1. 01

    Assessment and advisory

    We start by establishing what your data estate actually is, what it costs to run and where the distrust in the numbers comes from. The output is a costed, sequenced plan rather than an architecture diagram.

    • Every feeding source inventoried. Source systems, extracts, databases and reporting spreadsheets currently feeding a management report, including the ones maintained privately on desktops.
    • Preparers and consumers interviewed. How long each pack takes to produce, where it is corrected by hand, and which figures the audience quietly does not believe.
    • Contested metrics traced end to end. Three to five disputed numbers followed from source field to the figure on the slide, documenting exactly where the versions diverge.
    • Current spend established first. Licensing, capacity and storage cost recorded before any target state is priced, so the business case compares like with like.
    • Readiness scored and ranked by blocker. Ownership, definitions, quality, lineage, sensitivity and access control scored, with each gap ranked by what it actually blocks.
    • Sequenced roadmap with success measures. Effort, dependencies and the measures we will report against after go-live, agreed before the first table is created.
  2. 02

    Architecture and implementation

    We build in increments that each deliver a report someone uses, rather than a platform programme that shows nothing until year two. The first release is normally one business domain taken end to end into production.

    • Layered zones and naming agreed first. Raw landing, cleansed and conformed, and business-ready serving layers, with naming standards settled before the first table exists because renaming later breaks every downstream report.
    • Incremental loads with replay. Ingestion built per source with watermarking and re-execution, so a failed run can be rerun without duplicating rows or a manual clean-up.
    • Transformations as reviewed code. Logic held in version control and promoted through development, test and production by pipeline rather than edited in the production workspace.
    • Modelled once in a shared semantic layer. The business layer modelled dimensionally with documented measures and hierarchies, so a new report extends the model instead of recreating its arithmetic.
    • Quality tests at every boundary. Schema checks, referential integrity, null and range thresholds, and reconciliation back to the source system total on each load.
    • Legacy reports retired, not rebuilt. Existing reports migrated deliberately, with the ones nobody opens switched off rather than reproduced by default in the new platform.
  3. 03

    Security and governance

    Analytics widens who can see what, so the controls have to arrive with the platform rather than after the first awkward discovery. This block covers the controls that ship with the platform; the wider cataloguing, classification and compliance-evidence programme is on our data governance and compliance page.

    • Owner and sensitivity per dataset. Each dataset carries a named owner, a retention position and the basis on which any personal information is held, recorded rather than assumed.
    • Access through identity groups. Role-based access granted to groups, with row-level and column-level restrictions wherever a report crosses business units or reveals individual pay, cost or client detail.
    • Lineage recorded to the source system. Technical lineage registered so a reported figure can be traced back to the system it originated in when somebody challenges it.
    • Non-production data masked. Development, test and production separated, with personal data masked or synthesised in every non-production copy rather than restored intact.
    • Sensitive access logged where it is read. Audit logging enabled on access to sensitive datasets and routed into the monitoring the security team already reviews, not a second console.
    • Controls mapped to your framework. The control set documented against the framework your organisation is aligning to, such as the Essential Eight or the Australian Privacy Principles, with open gaps recorded rather than tidied away.
  4. 04

    Adoption and enablement

    A dataset nobody opens is a cost. Adoption is planned as part of delivery, with named owners on the business side and training pitched at what people actually do each week.

    • Named owners before go-live. Data owners and report owners agreed in writing, with what each one is accountable for recorded rather than implied by job title.
    • Training split by what people do. Separate sessions for report consumers, self-service authors and the small group who will maintain shared measures, because one session serves none of them.
    • A data dictionary in plain language. Each measure published with its definition, refresh schedule and who to contact when it looks wrong, written for the business rather than for engineers.
    • Retirement dates for what is replaced. Agreed dates for switching off the spreadsheets and legacy reports, with reconciliation confirmed by the owner before anything goes dark.
    • A request path for new measures. A short review step that decides whether a request extends the shared model or stays local, which is what stops the model becoming a dumping ground.
    • Usage tracked and acted on. Activity reviewed after launch, with any report showing no use after an agreed period revisited with its owner or retired.
  5. 05

    Managed service continuation

    Pipelines fail quietly, upstream systems change schema without telling anyone and cost drifts upward. We stay on afterwards to run the platform and keep improving it.

    • Refresh and load monitoring. Pipeline and refresh outcomes watched, with alerts on failures, late arrivals and row counts falling outside expected ranges rather than a daily inbox digest.
    • Failures investigated to a cause. Load failures resolved including reruns and backfills, with the cause and the fix recorded each time so the same failure is not diagnosed twice.
    • Monthly cost and capacity review. Capacity, storage and query cost reviewed on a cycle with specific recommendations such as archiving, partitioning or resizing attached.
    • Upstream schema changes absorbed. Source changes applied in a controlled way, with the impact on downstream models and reports assessed before release rather than discovered by a user.
    • Reporting against the agreed measures. Monthly reporting on refresh reliability, data quality test results, usage and cost against the measures set at the start of the engagement.
    • A visible improvement backlog. A prioritised list worked through in agreed increments alongside business-as-usual support, so improvement is funded rather than hoped for.

Questions we get asked

The things people ask before they commit.

Including the awkward ones. If the honest answer is that this is not right for you, that is the answer you will get.

  • What does a data platform project actually cost?

    The build depends on how many sources feed the first domain and how bad the definitions are, and anyone quoting a figure before seeing that is guessing. What we can be firm about is the shape of the cost: an assessment, then an increment that puts one business domain into production, then further increments. The running cost is the part clients underestimate. Capacity or query compute, storage and licences all continue after the project ends, so we cost the run alongside the build and tell you which line grows quietly as usage rises.

  • Can our own BI team do this instead of hiring a consultancy?

    Often yes, and where that is the case we will say so. A team with a competent data engineer and an analyst who understands the business can build a good platform, and they will maintain it better than we would because they are there every day. Where outside help earns its fee is the decisions that are expensive to reverse: the zone and naming structure, whether the semantic layer or the warehouse holds the business logic, and the political work of getting one agreed definition. If what you actually need is two weeks of architecture review and a decision record, we would rather sell you that than a programme.

  • Our last data warehouse project failed. Why would this one be different?

    Usually because the last one delivered a platform before it delivered a report anyone used. The failure pattern is a two-year build with no visible output, a modelling layer designed without the people who argue about the numbers, and a go-live that duplicated every legacy report rather than retiring the dead ones. We deliver one business domain end to end first, and if that increment does not change how a real meeting runs, that is a signal to stop rather than to scale up.

  • Do we have to pick Microsoft Fabric or AWS, and how locked in are we?

    The storage layer is the portable part. Both platforms now keep data in open table formats over object storage, so the raw and curated layers can be read by other tools and are the least locked-in thing you will build. What is not portable is everything above it: semantic models, pipeline definitions, capacity and workspace configuration, and permission grants all have to be rebuilt if you move. The realistic position is that your data stays yours and your platform investment does not, so choose on where your existing skills, applications and identity already sit rather than on a feature comparison.

  • How long before we see anything useful?

    The first production increment is usually a matter of weeks rather than quarters, and it is normally one business domain: one set of numbers, one audience, one legacy report switched off. What extends that is not the technology, it is the definitions. If three people have to agree what utilisation means and their diaries do not align for a fortnight, the build waits. We name that dependency in the plan rather than absorbing it silently and then explaining a slip.

  • Do we need a data warehouse, a lakehouse, or both?

    Most Australian mid-market estates need one modelled serving layer and a place to land raw data cheaply, which is what a lakehouse gives you without running two platforms. A separate warehouse earns its cost where concurrency is high, the queries are relational and predictable, or third-party tools connect over SQL and expect warehouse behaviour. We decide it per workload on observed query patterns rather than on architecture fashion, and we are happy to keep an existing warehouse that works and put the lake beside it.

  • How much of our reporting can the business build itself afterwards?

    More than most organisations get today, but not everything, and it is worth being honest about the ceiling. Once the shared measures are defined and tested, an analyst can build a new report in an afternoon and that covers the majority of requests. What still needs an engineer is a new source, a change to a shared measure, or anything crossing security boundaries. Self-service works when the model underneath is trustworthy and someone owns it, and it becomes sprawl again when nobody does.

Related industries

  • Manufacturing and Distribution

    Demand and inventory intelligence, production analytics and supplier automation built on data your planners already trust.

  • Logistics and Warehousing

    Shipment and document automation, proof-of-delivery processing and exception dashboards that let a small team run a large network.

  • Construction and Property

    Tender intelligence, addenda tracking and project reporting that keep estimators and contract administrators ahead of the documents instead of buried in them.

  • Professional Services

    Governed enterprise search, document intelligence and secure copilots that respect matter confidentiality and conflict boundaries.

Related services

  • Artificial Intelligence

    AI that is chosen for a reason, costed before it is built and governed once it is live.

  • Data Governance and Compliance

    Discovery, classification, access control and audit evidence over the data you already hold, designed against the obligations that actually apply to you.

  • Cloud Modernisation

    Ageing systems become a cloud platform your team can change safely, recover predictably and account for line by line.

  • Managed Services

    Your platform keeps earning its business case after go-live, with cost, security posture, reliability and adoption reviewed on an agreed cycle rather than left to drift.

Free discovery workshop

Start with a data and analytics discovery workshop.

Bring one challenge in this area. We will map the opportunity, the readiness gaps and a recommended next step.