Microsoft Fabric Architecture Explained: How OneLake, Workloads, and Capacity Actually Fit Together

Microsoft Fabric architecture, shown as glowing columns rising from one shared luminous base layer representing OneLake

Last updated: October 8, 2026

By ZapAI Team

TL;DR: Microsoft Fabric architecture comes down to three things working together. OneLake is the single storage layer. A set of workloads (Data Factory, Data Engineering, Data Warehouse, Real-Time Intelligence, Data Science, and Power BI) all read and write the same Delta tables in that lake. And a shared capacity model bills compute as one pool instead of per service. No workload keeps its own private copy of the data.

Microsoft Fabric architecture is OneLake as one shared storage layer, compute workloads that read and write the same Delta tables in it, and one capacity that every workload draws from.

If you’ve spent any time in a traditional Azure data estate, you know the pattern. A Synapse dedicated SQL pool holds one copy of your sales data, a Data Lake Storage account holds another, and a Power BI dataset holds a third. Three copies, three refresh schedules, three places for numbers to quietly drift apart. Microsoft Fabric is built to end that pattern, and its architecture makes sense once you see storage first, then compute, then the decisions layered on top.

That’s the order this guide follows. It’s written for the person deciding how to build on the platform, whether you’re migrating off Synapse, standing up a greenfield analytics stack, or trying to figure out why your Fabric capacity keeps throttling.

What Is Microsoft Fabric, Architecturally?

Fabric is a SaaS analytics platform that puts data engineering, warehousing, real-time analytics, data science, and BI into one environment on top of one data lake.

Skip the marketing description for a second. Per Microsoft’s Fabric overview, the platform is delivered as SaaS, so there’s no cluster or storage account for you to provision. It also leans on data mesh thinking, with domains and workspaces that let business units own their data instead of routing everything through one central data team. If you want the broader product tour first, start with what Microsoft Fabric is.

One thing worth clearing up early, because it trips people up in search results and in meetings. Microsoft Fabric and “data fabric” are not the same thing. Data fabric is an architectural pattern, not a product. It uses metadata to discover, connect, and manage data across many platforms without moving every dataset. Microsoft Fabric borrows some of that thinking, but it’s a specific commercial product with a specific storage layer underneath. Mixing the two up leads to some confusing procurement conversations.

OneLake: The Layer Everything Else Sits On

OneLake is the single, tenant-wide data lake that every Fabric workload reads from and writes to. Almost every architectural decision downstream traces back to it.

According to Microsoft’s OneLake documentation, OneLake comes automatically with every Fabric tenant, with no infrastructure to manage, and it’s built on Azure Data Lake Storage. Microsoft’s own shorthand is useful here. Think of it as OneDrive for data. The same way every Microsoft 365 app opens the same OneDrive file, every Fabric workload reads the same OneLake table.

The storage hierarchy runs tenant, then domain, then workspace, then item, then tables and files. There’s only one OneLake per tenant, and every workspace shares that single namespace. That design is exactly why you don’t end up asking “which copy is correct,” the question that haunts traditional multi-tool estates.

Underneath, tables are stored in open formats, primarily Delta Lake. Delta is Parquet files plus a transaction log, which lets plain files behave like database tables with ACID guarantees. Open matters more than it sounds. Tools that already speak Parquet or Delta, Azure Databricks included, can read OneLake data without a translation layer, and OneLake also supports the existing ADLS Gen2 APIs and SDKs.

The Workload Layer: What Actually Runs on Top

OneLake is storage. The workloads are compute, and each one is a specialized lens on the same data rather than a separate silo.

  • Data Factory handles ingestion and orchestration, connecting to source systems and moving data into OneLake through pipelines and dataflows.
  • Data Engineering is the Spark and notebook environment where PySpark, Scala, and Spark SQL do the heavy transformation work. Our Fabric data engineering team spends most of its time here.
  • Data Warehouse is the T-SQL-first, fully transactional relational experience for teams who think in stored procedures and star schemas.
  • Real-Time Intelligence is built for streaming and event data, like IoT telemetry and logs, that needs to be queried the moment it lands.
  • Data Science covers the machine learning workflow from training through deployment.
  • Power BI sits on top as the reporting layer, and it’s the one most people already know.

The list isn’t the interesting part. What connects the workloads is. Because one copy of data in OneLake powers all of them, you don’t maintain separate copies for the lake, the warehouse, BI, and real-time analytics. A notebook can write a table that a Power BI report reads an hour later, with nothing physically moving in between.

Several light beams drawing from one central pool, representing Fabric workloads reading the same OneLake data

Lakehouse vs. Warehouse: The Decision You’ll Actually Face

Pick Lakehouse if your team works in Spark and handles messy or unstructured data. Pick Warehouse if your team works in T-SQL and serves structured reporting. Both store Delta tables in OneLake.

This is where most teams get stuck early, because Fabric gives you two distinct ways to store and query structured data, and the difference isn’t cosmetic. Microsoft’s lakehouse vs. warehouse decision guide boils it down to data type and team skill set.

LakehouseWarehouse
Data typesStructured, semi-structured, unstructuredStructured
Primary languagePySpark, Spark SQL, ScalaT-SQL
Typical usersData engineers, data scientistsWarehouse developers, SQL engineers
Writes via SQLRead-only SQL analytics endpointFull multi-table transactions
Storage underneathDelta tables in OneLakeDelta tables in OneLake

In practice, mature Fabric estates usually end up using both in the same project. A common pattern puts the bronze and silver layers in Lakehouses and the reporting layer in a Warehouse, which is tuned for high-concurrency SQL. Neither one wins outright. They’re built for different jobs in the same pipeline, and because both sit on Delta, you aren’t locked into one paradigm forever. We cover the build side in more depth on our Fabric lakehouse and Fabric data warehouse pages.

Direct Lake Mode: The Real Architectural Innovation

Direct Lake lets Power BI read Delta tables in OneLake directly, giving near-Import speed without copying data into the semantic model.

If OneLake is the foundation, Direct Lake is arguably the cleverest thing built on it, because it solves a problem Power BI users have lived with for over a decade. Traditionally you had two choices. Import mode copies data into an in-memory model for speed but goes stale between refreshes. DirectQuery stays fresh but pays a latency tax on every visual click.

Per Microsoft’s Direct Lake overview, the semantic model loads columns straight from the Delta files in OneLake into the same VertiPaq engine Import mode uses. A Direct Lake refresh copies only metadata, a step Microsoft calls framing, which can finish in seconds. With automatic updates turned on, changes in OneLake show up in the report without a scheduled refresh.

One thing to know before you commit a production report. Direct Lake comes in two flavors. Direct Lake on OneLake never falls back to DirectQuery. Direct Lake on a SQL analytics endpoint does fall back when it can’t read a Delta table directly, for example when the source is a SQL view or the warehouse uses SQL-based row-level security. If you’re on the SQL endpoint flavor, test realistic queries under load before you assume Import-level speed.

Analyst reviewing a live dashboard on a laptop, representing Power BI reading OneLake data through Direct Lake mode

Getting Data In: Shortcuts vs. Mirroring

Shortcuts point at data where it already lives, with no copy. Mirroring continuously replicates a database into OneLake, so Fabric holds a real copy.

A OneLake shortcut references data in another system, such as Azure Data Lake Storage, Amazon S3, Dataverse, or another OneLake location, and lets you query or join it with local data without an initial migration. Nothing physically moves. It’s a pointer, and it’s the right call when data already lives somewhere sensible and Fabric just needs to see it.

Mirroring takes the opposite approach. It connects to an operational database, such as Azure SQL Database, SQL Server, Snowflake, or Oracle, and keeps a replica in OneLake in near real time. Reach for it when analytics should run against a replicated copy instead of hammering a live operational system, or when the source isn’t something a shortcut can reach at all.

Neither approach is better across the board. They answer different questions about where your data should physically live, and most estates use both.

Medallion Architecture: How Data Gets Organized Once It’s In

Fabric’s recommended layout is bronze for raw data, silver for cleaned data, and gold for reporting-ready data, with quality rising at each layer.

Once data lands in OneLake, through a shortcut, mirroring, or a pipeline, it needs structure. Microsoft calls the medallion lakehouse architecture the recommended design approach for Fabric.

  • Bronze preserves data exactly as it arrived, so there’s always a source of truth to fall back on.
  • Silver is where deduplication, standardization, and error correction happen.
  • Gold is shaped for reports and dashboards. It’s the layer your BI team actually queries.

This maps neatly onto the Lakehouse vs. Warehouse decision. Bronze and silver usually live in Lakehouses where Spark does the heavy lifting, and gold often graduates into a Warehouse once it has to serve governed, high-concurrency BI traffic. Microsoft also suggests putting each layer’s lakehouse in its own workspace, which gives you cleaner access control and cost tracking per layer.

Bronze, silver, and gold discs side by side, representing the medallion architecture in a Fabric lakehouse

The Capacity Model: How Compute Is Priced and Shared

Every Fabric workload runs on one shared pool of capacity units. You buy an F SKU, and Spark jobs, SQL queries, and Power BI reports all draw from it.

This is the part that catches architecture teams off guard, because it’s a different billing philosophy from the rest of Azure. Fabric doesn’t charge per query or per pipeline run. You buy a capacity sized in capacity units (CUs), and Microsoft’s licensing table lists SKUs from F2 up to F8192, each tier doubling the one below it. Cost scales roughly in line with CUs, as the Azure Fabric pricing page shows.

The behavior that matters most day to day is what happens when you run hot. Fabric uses bursting and smoothing to absorb short spikes, then throttles if you stay over. An undersized capacity doesn’t show up as a bigger invoice. It shows up as slow reports and delayed pipeline runs.

There’s also a licensing cliff to know before you size anything. F64 is the smallest SKU where users with a free Fabric license can view Power BI content. Below F64, every report viewer needs a Pro or Premium Per User license. For an organization with a large internal Power BI audience, that one threshold can reshape the whole cost calculation, and our Fabric vs. Power BI Premium breakdown walks through the break-even math.

Here’s the operational reality worth internalizing. A heavy Spark transformation and a business user refreshing a dashboard draw from the exact same CU pool. That’s very different from separately billed, separately scaled Azure services, and it means capacity planning has to account for every workload sharing the estate, not just the one you’re building this quarter.

Rows of servers in a data center, representing the shared capacity units every Fabric workload draws from

Security and Governance Layers

Fabric security narrows in layers, from workspace roles to item permissions to row- and column-level security, with tenant governance running through the OneLake catalog and Microsoft Purview.

Fabric workspaces have four roles. Admins have full control. Members can also share and manage permissions. Contributors create and edit content. Viewers can only view. Below that sit item-level permissions for finer sharing, and below that, data-level security such as row-level and column-level rules that restrict what a user sees no matter how they reach the data.

Tenant-level governance has recently moved house. Microsoft retired the Purview Hub report at the end of January 2026 and moved its security, sensitivity label, and DLP insights to the Govern tab in the OneLake catalog. That tab is now the admin’s main view of label coverage, endorsement, and security health. Microsoft Purview still supplies the sensitivity labels and data loss prevention policies, and labels flow downstream as data moves between items. That matters more every month as data starts feeding Copilot and AI agents that need to respect the same permission boundaries a human user would.

Where Does Fabric Fit Against Synapse and Databricks?

Fabric is Synapse’s successor for new work. Against Databricks, Fabric suits Microsoft-centric, reporting-driven teams, while Databricks suits multi-cloud and heavy custom ML.

You don’t design a Fabric architecture in a vacuum. Most enterprise teams are weighing it against what they already run. Microsoft has concentrated new feature investment in Fabric rather than Synapse, positioning it as the platform that brings Synapse, Power BI, Azure Data Factory, and Azure Data Lake together on OneLake. Synapse is still supported, but a greenfield build in 2026 has little reason to start there. Our Fabric vs. Azure Synapse comparison covers the details.

Databricks is a different conversation. Fabric is built to cut moving parts by putting engineering, warehousing, reporting, real-time analytics, and ML in one managed platform. Databricks is built for flexibility and gives teams more control, especially across clouds. If your roadmap is Microsoft-centric and driven by reporting, Fabric is usually the more natural fit. If you need multi-cloud compute and deep custom ML work, Databricks still has the edge, and plenty of enterprises run both, with OneLake shortcuts pointing at the shared data.

Common Fabric Architecture Mistakes Worth Avoiding

The three we see most often are treating OneLake as a file dump, sizing capacity as an afterthought, and securing only one access path.

The earliest one is treating OneLake as simple storage rather than the foundation of the whole architecture. Teams keep adding files, datasets, and workloads without a clear workspace and domain structure, and duplicate datasets and unclear ownership follow close behind.

Capacity planning gets skipped almost as often. Nobody notices until month-end, when a big Spark job and a board report land on the same CU pool and the throttling starts.

Then there’s security applied inconsistently across access paths. Tight row-level security in a Power BI semantic model doesn’t help if the same data is wide open through direct Spark or SQL access. A boundary like that only holds for users who happen to come in through the front door.

If you’re moving an existing estate onto Fabric, our step-by-step Fabric migration guide shows how to sequence workloads so these problems get designed out, not discovered later.

Fabric Architecture Questions We Hear Most

What is OneLake in Microsoft Fabric?

OneLake is Fabric’s single, tenant-wide logical data lake, built on Azure Data Lake Storage. Every Fabric workload reads from and writes to it, so data doesn’t need to be duplicated between tools.

How is Fabric different from a traditional data warehouse architecture?

Traditional architectures usually keep separate copies of data for warehousing, BI, and analytics tools. Fabric keeps one copy in OneLake and lets every workload query that same copy directly.

Is Fabric priced by capacity or by user?

Mostly by capacity. You pay for an F SKU sized in capacity units, and that pool is shared across every workload. Per-user Power BI licenses still apply below F64. From F64 up, users with a free license can view Power BI content.

When should you use a Lakehouse instead of a Warehouse in Fabric?

Choose a Lakehouse when your team works in Spark and handles raw, semi-structured, or unstructured data. Choose a Warehouse when your team works in T-SQL and needs governed, structured relational reporting. Many projects use both.

Does Fabric replace Azure Synapse?

Microsoft has positioned Fabric as Synapse’s successor and directed new investment there. Existing Synapse workloads keep running and stay supported, but a new analytics platform built today has little reason to start on Synapse.

Where This Leaves You

Most of what feels like a “Fabric decision” is really two decisions. Where should each piece of data physically live, and how much shared capacity are you willing to size for? Get OneLake’s role straight, match each workload to the job it’s built for, and the architecture stops looking like a list of product names and starts looking like one system.

In our experience the tools are rarely the problem. Skipping the planning step before workspaces start multiplying is. If you’d like a second set of eyes on your workspace design or capacity sizing, our Microsoft Fabric consultants do exactly this work. Book a free consultation to talk it through.

Leave a Reply

Your email address will not be published. Required fields are marked *

Book a free consultation