Skip to main content

Data and Analytics on Microsoft Azure

Data and analytics on Microsoft Fabric and Azure

Microsoft Fabric puts ingestion, engineering, warehousing, real-time analytics and Power BI on one governed footprint over OneLake, which removes most of the copying that used to sit between those layers. That matters most if your organisation already runs Microsoft 365, Entra ID and Power BI, because identity, sensitivity labels and a large part of the licensing are already in place. We settle tenant, capacity and workspace structure first, then build the pipelines, the model and the reports on top of it. The capacity decision is the one that quietly determines your bill, so it is made with numbers rather than a default.

Why Azure

When this is the right platform.

  • OneLake gives every Fabric workload one storage layer in Delta format, so a warehouse table, a Spark notebook and a Power BI semantic model read the same data without a separate copy being made for each tool.
  • Where Microsoft 365 and Power BI are already licensed, the identity model, sensitivity labels and sharing behaviour your users know carry straight across, which shortens the adoption work considerably.
  • Power BI remains the strongest fit when finance and operations teams want to build their own reports against a governed semantic model rather than raising a ticket for every question.
  • Direct Lake lets Power BI read Delta tables in OneLake without a scheduled import refresh, which suits large models previously constrained by the length of the overnight refresh window.
  • Mirroring brings operational databases into OneLake as read-only replicas without you writing an extraction pipeline, which removes a category of plumbing that used to be a project on its own.
  • Governance metadata, classification and lineage sit in the same tenancy through Microsoft Purview, so the governance conversation is a scope decision rather than a second product purchase.

Where it is less suited

We would rather say this now than after a migration.

  • Fabric capacity is a shared pool, and that is the design consequence most teams meet first. One badly written notebook or an over-frequent refresh can throttle unrelated workspaces on the same capacity, because bursting and smoothing spread consumption over hours rather than isolating it. Workload separation, capacity metrics and a surge protection threshold need deliberate design, not defaults.
  • The two Direct Lake modes fail differently and the difference matters. Direct Lake on SQL endpoints falls back to DirectQuery when it cannot read the Delta table directly, for example against a view or where SQL-based row-level security applies, which shows up as a report that is suddenly slow. Direct Lake on OneLake has no fallback: exceed the guardrails and the refresh fails outright. Neither is a universal replacement for import mode.
  • The platform moves quickly and not all of it lands in Australia at once. Features change behaviour between releases and some capabilities reach general availability later than the announcement implies, so we build against current behaviour, record the verification date, and avoid depending on preview features unless you accept that risk in writing.
  • The commercial decision has partly been made for you. Microsoft stopped selling Power BI Premium per-capacity subscriptions in early 2025 and set end-of-life dates that differ by agreement type, so organisations are moving to Fabric capacity at renewal regardless of whether they wanted the wider platform. Confirm your own dates against your agreement, not against an article.
  • If your organisation has little existing Microsoft 365 or Power BI footprint, much of the licensing and adoption advantage does not apply and Fabric should be judged on technical fit alone. In that situation we have recommended the other platform, and we would again.
  • Mirroring removes the extraction pipeline, not the modelling. It lands source tables in their source shape, with supported-source and configuration conditions attached, so a mirrored database is a starting point for conformance rather than something a report should be built directly on.

Business outcomes

What Azure delivers here.

Reporting consolidated onto one platform
Teams currently spread across Excel, on-premises reporting services and standalone Power BI files work from shared workspaces and a single semantic model, with personal workspaces no longer quietly acting as production.Agreed measure: Share of production reports resolving to the shared semantic model rather than a private extract, verified at handover.
Refresh windows stop constraining the business
Analysts stop planning their day around a nightly import that takes hours, because Direct Lake reads the Delta tables in OneLake directly and the model is current as soon as the load completes.Agreed measure: Current refresh duration baselined per model before migration, then reported against after the model moves.
Capacity cost you can attribute and defend
Platform owners see which workspaces and workloads consume capacity units, so an over-refreshed model or a runaway notebook can be found and fixed before it throttles somebody else's report.Agreed measure: Monthly capacity report by workspace with the utilisation threshold that triggers a review agreed before go-live.
Row-level security that has been tested
Where a report crosses business units, restrictions are implemented in the semantic model and proven with a representative account for each role rather than asserted in a design document.Agreed measure: Every role tested with a representative account before release, including one that should see nothing at all.
Operational data available the same day
Fabric Real-Time Intelligence carries event streams from plant, fleet or application telemetry so exception dashboards update through the day, while the rest of the estate stays on a batch schedule it does not need to leave.Agreed measure: Latency target agreed per measure before build, then reported against once the stream is live.
A platform your team can change
Workspaces, deployment pipelines and Git integration set up so a change moves from development to production by promotion rather than by somebody editing the production report at four in the afternoon.Agreed measure: Whether a change can be promoted from development to production without manual edits, tested at handover.

Common client problems

What we usually hear first.

  • We have Power BI everywhere and no idea which report is the official one.

    We audit the tenant for workspaces, semantic models and reports, identify duplicates and orphaned content, and consolidate the survivors onto shared models in governed workspaces. Publishing rights are then tightened so personal workspaces stop becoming production by accident. Expect the audit to find reports whose owner left two years ago and which a senior manager still opens weekly, and expect that conversation to be the slow part.

  • Someone told us to move from Power BI Premium to Fabric and nobody can explain what actually changes.

    Commercially the decision has largely been made for you: Microsoft stopped selling Power BI Premium per-capacity subscriptions in early 2025 and set end-of-life dates by agreement type, so most organisations are moving to Fabric capacity at renewal whether or not they wanted the wider platform. Technically what changes is the consumption model. We map your workloads and refresh patterns against capacity behaviour, including how bursting and smoothing affect what you are billed, and confirm the current dates against your own agreement rather than a blog post.

  • Our data still sits in an on-premises SQL Server and reporting runs off a nightly copy.

    We assess whether each workload belongs in Azure SQL Database, a Fabric warehouse or a lakehouse, based on query pattern, latency need and how the application writes to it. Ingestion is then built with Fabric Data Factory, and the on-premises source stays in place until the new reports reconcile against it. The gateway is usually the thing that surprises people: it becomes a single point of failure and it needs an owner, a host and a patching plan.

  • We already run Azure Databricks and we are not throwing that away.

    You do not need to. We keep Databricks where the heavy engineering and machine learning work already lives, surface its curated Delta output into OneLake through shortcuts, and use Fabric and Power BI for the serving and reporting layer. Running both means two skill sets and two cost lines, so we are direct about whether the split is worth keeping or whether one of them is being retained out of habit.

  • Every new report request turns into a three-week project.

    That is usually a modelling problem rather than a capacity problem. We build a governed semantic model with the measures the business argues about already defined and tested, so most new questions become a report an analyst builds in an afternoon. If requests are slow because a single overloaded person is the only one who may publish, that is an operating model problem and no amount of platform work will fix it.

  • Our capacity keeps throttling and nobody can tell us which workload caused it.

    Capacity is shared, so an over-frequent refresh or one badly written notebook can degrade unrelated workspaces. We use the capacity metrics to attribute consumption by item and operation, separate interactive from background use, move the offenders or reschedule them, and set a surge protection threshold so background work is rejected before it starves the reports people are looking at.

How we deliver

Our Azure delivery approach.

  1. 01

    Assessment and advisory

    The Azure assessment is largely about what you already own. Most Microsoft-centred organisations are paying for more analytics capability than they are using, so the plan starts with what can be recovered before anything new is bought.

    • Power BI tenant audited. Workspaces, semantic models, reports, refresh failures, gateway dependencies and content nobody has opened in 90 days, listed with an owner or marked as orphaned.
    • Capacity and licensing established. Current Fabric or Power BI capacity, utilisation patterns, and where throttling or bursting is already happening, recorded before any sizing recommendation is made.
    • Existing pipelines catalogued. Azure Data Factory pipelines, SQL Server Integration Services packages, on-premises gateways and the manual extracts nobody has admitted to yet.
    • Fabric compared against what you run. A Fabric-first design costed against retaining Azure Databricks, existing Synapse workloads or Azure SQL Database, with effort attached to each option rather than a recommendation by default.
    • Direct Lake feasibility checked per model. Table sizes, use of views and row-level security assessed against the Direct Lake guardrails, because the two Direct Lake modes behave differently when those limits are crossed.
    • Australian region and feature check. Capability availability confirmed in Australia East and Australia Southeast with the verification date recorded, since Fabric features do not arrive everywhere at once.
  2. 02

    Architecture and implementation

    We settle tenant, capacity and workspace structure before the first pipeline is built, because retrofitting workspace boundaries and naming afterwards is disruptive and visible to users. The first business domain then goes end to end into production.

    • Workspace and capacity structure first. Tenant settings, capacity assignment and workspaces separating development, test and production, with Fabric deployment pipelines and Git integration between them.
    • Land once, shortcut rather than copy. Raw data into OneLake through Fabric Data Factory, using shortcuts to Azure Data Lake Storage or existing Delta tables instead of creating a second copy per tool.
    • Operational sources mirrored where it fits. Mirroring used to bring supported operational databases into OneLake as read-only replicas, which removes an extraction pipeline but not the modelling work behind it.
    • Transformations in notebooks or dataflows. Cleansing and conformance built in Fabric Data Engineering where code is clearer than a low-code flow, with the code held in Git and peer reviewed before promotion.
    • Serving layer chosen on the consumer. A Fabric Data Warehouse or a lakehouse SQL endpoint, decided on whether consumers write T-SQL or read through the semantic model, not on preference.
    • Semantic model with explicit measures. Power BI measures, hierarchies and storage mode set deliberately, with Direct Lake used where table size and query pattern genuinely suit it and import kept where they do not.
    • Streaming only where latency is paid for. Fabric Real-Time Intelligence with Azure Event Hubs for the measures that need sub-hour currency, and everything else left on batch.
  3. 03

    Security and governance

    Fabric inherits Microsoft Entra ID and the wider Microsoft compliance estate, which makes these controls straightforward to implement and equally easy to leave at their defaults. We configure them during the build. The broader cataloguing and classification programme is covered on our data governance and compliance page.

    • Workspace access through Entra groups. Workspace and item access granted to Entra ID security groups mapped to roles, with no standing access held by individual named accounts.
    • Row and object security proven by test. Restrictions implemented in the semantic model and tested with a representative account for each role before the report is released, including one that should see nothing.
    • Enforcement placed where it holds. Security applied at the layer that can actually be tested, since lake-level enforcement in Fabric has changed recently and inherited behaviour differs between Direct Lake modes.
    • Sensitivity labels applied by the pipeline. Labels set so classification persists into Power BI exports to Excel and PDF rather than stopping at the workspace boundary.
    • Tenant settings tightened deliberately. Publish to web, external sharing and unmanaged personal workspaces restricted, with every approved exception documented against a named approver.
    • Activity auditing routed to the security team. Fabric and Power BI activity logs forwarded to the destination the security team already reviews instead of a portal nobody opens.
  4. 04

    Adoption and enablement

    Power BI's real strength is that business users can answer their own questions. That only works when the model underneath is trustworthy and somebody has shown them how to use it properly.

    • Sessions split by role. Separate enablement for report consumers, self-service authors and the small group who will maintain the shared semantic model.
    • Endorsement used as a signal. The certified and promoted standard applied in the tenant so an endorsed dataset is visibly distinguishable from somebody's working draft.
    • A written measure dictionary. Each certified measure explained in plain language with its refresh schedule and its owner, published where users already look.
    • Retirement dates for what is replaced. Agreed dates for the spreadsheets and legacy reports being switched off, with reconciliation confirmed by the owner first.
    • A request path for new measures. A review step deciding whether a request extends the shared model or belongs in a personal report, so the model does not accumulate one-off logic.
    • Usage reviewed after launch. Workspace usage metrics checked on an agreed cycle, with content showing no viewers revisited with its owner or retired.
  5. 05

    Managed service continuation

    Capacity consumption, refresh reliability and semantic model performance all drift as usage grows and as the platform releases change. We watch them and act before users notice.

    • Refresh outcomes monitored. Pipeline and semantic model refreshes watched with alerting on failures, skipped runs and refreshes exceeding an agreed duration.
    • Capacity consumption attributed. Capacity units tracked by workspace and workload, with throttling events identified and resizing or workload separation recommended with numbers attached.
    • Slow models tuned on evidence. Partitioning, measure rewriting and moves between import and Direct Lake made where the metrics support it rather than on a hunch.
    • Platform releases tested first. Upstream schema changes and Fabric platform updates assessed in a non-production workspace before they reach the reports people rely on.
    • Permission and tenant drift reported. Workspace permissions, label coverage and tenant settings reviewed on an agreed cycle, with any drift from the approved baseline raised to a named owner.
    • Monthly report against the measures. Refresh reliability, data quality results, report usage and capacity cost reported against what was agreed at the start.

Reference architecture

A Fabric analytics platform, layer by layer.

How the pieces fit together on Microsoft Azure. Every engagement adapts this, and we will tell you which layers you already have.

  1. 01

    Sources

    Where the numbers already live, most of them not in Fabric yet.

    • SQL Server and Azure SQL Database
    • ERP and line-of-business APIs
    • Event streams and files
  2. 02

    Land

    One copy in OneLake, shortcut or mirrored rather than re-copied per tool.

    • OneLake
    • Fabric Data Factory
    • Shortcuts and mirroring
  3. 03

    Model

    Conformed, then modelled once as the layer every report must go through.

    • Fabric Data Engineering
    • Fabric Data Warehouse
    • Power BI semantic model
  4. 04

    Serve

    Reports, self-service and the few measures that genuinely need minutes.

    • Power BI reports and apps
    • Fabric Real-Time Intelligence
    • Excel on the model

Across every layer

  • Entra ID groups for workspace, row and object-level access
  • Deployment pipelines and Git from development to production
  • Capacity metrics with an agreed surge protection threshold
  • Sensitivity labels that persist into exports
Layer 03 is the one that gets skipped. Under delivery pressure teams run a pipeline into OneLake and point Power BI straight at it, and 18 months later there are 40 reports each carrying their own version of margin. The modelled layer is the cheapest thing to build first and the most expensive thing to retrofit.

Technology reference

The Microsoft Azure services we build with.

A reference architecture view of the platform services used in this domain, and what each one does in the design.

Lakehouse foundation

  • Microsoft FabricThe platform we standardise on for Microsoft-centred clients, so ingestion, engineering, warehousing and reporting share one capacity, one security model and one deployment pipeline.
  • OneLakeThe single storage layer every Fabric workload reads and writes, which lets us shortcut to data already held elsewhere instead of creating a copy per tool.
  • Azure Data Lake StorageWhere an existing raw landing zone already sits. We normally keep it and expose it into OneLake through a shortcut rather than migrating the files.

Ingestion and data engineering

  • Fabric Data FactoryPipeline orchestration and copy activity for source ingestion, including incremental loads, watermarking and retry behaviour for sources that fail intermittently.
  • Fabric Data EngineeringSpark notebooks and lakehouse tables for cleansing, conformance and the transformations that are clearer expressed in code than in a low-code flow.
  • Azure DatabricksRetained where heavy engineering or machine learning workloads already run, with curated Delta output surfaced into OneLake for the serving layer.

Operational and streaming sources

  • Azure SQL DatabaseThe transactional store behind line-of-business applications we ingest or mirror from, and occasionally the target where a workload needs relational behaviour rather than analytical scale.
  • Azure Cosmos DBSource for document and high-throughput operational data, brought into the lakehouse rather than queried directly by reports.
  • Azure Event HubsIngestion point for telemetry, device and application event streams before they reach real-time analytics.

Serving, modelling and reporting

  • Fabric Data WarehouseThe T-SQL serving layer for teams and third-party tools that expect a warehouse and connect over SQL.
  • Fabric Real-Time IntelligenceEvent stream processing and querying for the measures that must be current within minutes, such as production exceptions or dispatch delays.
  • Fabric Data ScienceNotebook-based modelling and experiment tracking for forecasting work that stays inside the same governed platform as the data it uses.
  • Power BIThe semantic model and reporting layer, where agreed measures are defined once, secured by role and consumed by self-service authors.

Governance the platform consumes

  • Microsoft PurviewWhere classification, lineage and labelling already exist we consume them rather than reinvent them. Standing that estate up is a separate engagement, covered on our data governance and compliance page.

Product names and icons are trademarks of Microsoft and Amazon Web Services, reproduced unmodified from their official architecture icon libraries to identify the technologies used in these architectures. Their presence does not indicate partnership, certification or endorsement by either vendor.

Azure questions

What people ask about doing this on Azure.

Cost, lock-in and the parts that go wrong, answered before you have to ask twice.

  • How much does Microsoft Fabric capacity cost to run each month?

    Capacity is the dominant line and it is priced by capacity unit, either reserved or pay-as-you-go, so the honest answer depends on which workloads you put on it. The traps are specific rather than mysterious: background operations such as refreshes and notebook runs are smoothed over up to 24 hours, so a spike you did not notice can throttle you the next morning, and every workload you consolidate onto one capacity competes with the reports users are looking at. We size against your observed refresh and query patterns, and we tell you where a smaller capacity plus scheduling discipline beats buying headroom.

  • Does Microsoft Fabric lock us in?

    Partly, and the split is worth understanding. OneLake stores Delta and Parquet over Azure storage, so your data is in an open format other engines can read, and that is the part you would keep. What you would not keep is everything above it: the semantic models, pipeline definitions, workspace and capacity configuration, and the Power BI reports themselves would all be rebuilt elsewhere. So the data is portable and the platform investment is not, which is true of the AWS side as well and is why the decision should follow your existing identity, applications and skills.

  • We are still on Power BI Premium P SKUs. What do we actually have to do?

    Move to Fabric capacity, and check your dates rather than assuming them. Microsoft stopped selling and renewing Power BI Premium per-capacity subscriptions from early 2025, with end-of-life dates that differ between Enterprise Agreement and non-Enterprise Agreement customers. The migration itself is mostly workspace reassignment and is not the hard part. The hard part is sizing, because Fabric capacity units and the old P SKU capacities do not map one to one once other workloads join the capacity, and that is where people either overpay or start throttling.

  • Should we use Direct Lake for all our Power BI models?

    No, and this is the most common piece of Fabric advice we end up reversing. Direct Lake is genuinely good for large tables read straight from Delta, and it removes the overnight refresh window. But Direct Lake on SQL endpoints quietly falls back to DirectQuery against views or where SQL-based row-level security applies, which users experience as a report that got slow for no reason. Direct Lake on OneLake does not fall back at all: cross the guardrails and the refresh fails. We test per model and keep import mode where it is the right answer.

  • Can we do this with our internal Power BI team instead of engaging you?

    If you have someone who already builds tested semantic models and understands capacity behaviour, quite possibly, and we will say so after the assessment rather than after the proposal. Fabric is not hard to start using, which is the problem: the expensive mistakes are structural. Workspace and capacity layout, whether business logic sits in the warehouse or the semantic model, and Direct Lake decisions per model are the ones that are painful to reverse once 30 reports depend on them. A short architecture engagement plus decision records is sometimes the whole answer.

  • Do we have to move off Azure Databricks to adopt Fabric?

    No, and often you should not. Where Databricks already carries the heavy engineering or machine learning work, the sensible pattern is to keep it, expose its curated Delta output into OneLake through shortcuts, and use Fabric and Power BI for serving and reporting. What we will be direct about is the cost of running both: two skill sets, two cost lines and a boundary somebody has to own. If Databricks is being kept because nobody wants to have the conversation, that is worth naming.

  • Will Fabric fix our data quality problems?

    No platform does that. Fabric makes it easier to test quality at each boundary and to see lineage when a figure is challenged, but the tests still have to be written and somebody still has to own the definition of a valid record. If a source system lets a user enter a delivery date in the past and nobody has decided which of two customer records is correct, that stays true after migration. We scope the quality rules explicitly during the build rather than implying the platform absorbs them.

Related industries

Where this work has the most leverage.

  • Manufacturing and Distribution

    Demand and inventory intelligence, production analytics and supplier automation built on data your planners already trust.

  • Professional Services

    Governed enterprise search, document intelligence and secure copilots that respect matter confidentiality and conflict boundaries.

  • Construction and Property

    Tender intelligence, addenda tracking and project reporting that keep estimators and contract administrators ahead of the documents instead of buried in them.

Free discovery workshop

Start with a data and analytics discovery workshop.

Bring one challenge. We will assess whether Microsoft Azure is the right platform for it before recommending anything.