Enterprise Data Integration and Analytics Architecture: One View of the Estate a Board Can See

By Michael Branson | August 26, 2026

Quick answer. An enterprise data integration architecture is one named design for the whole estate: where data originates, how it moves, who runs it, and where people read it. Governance and AI-readiness are through-lines across all four layers, not a fifth layer bolted on at the end.

The question usually arrives as an argument about a number. The finance director shows one figure in the board pack and the operations director shows a different one for the same week, and the meeting that follows is about which report is right. Neither report is the problem. Two reports built on two extracts of the same source disagree as soon as the extracts differ in timing, filter or definition, and no reconciliation done at the consumption layer settles what was decided upstream of it.

Underneath that argument sits an estate nobody designed. Connections were built one at a time, each correct on the day an engineer wrote it, each carrying its own copy of the data and its own assumption about what a field means. An analyst builds an extract because the warehouse lacks the column. A second analyst builds a second extract because the first one runs at the wrong hour. Years of locally reasonable decisions produce point-to-point spaghetti and every report a snowflake, and both are what an estate looks like from the inside when no architect holds the whole view.

Microsoft’s Guidance to set your organization’s data strategy describes the same condition in three short sentences: “Data is spread across systems and teams. Standards vary. Governance is inconsistent.” The third one is the one an architecture can act on. Spread and variation are consequences of how an organization grew and how it bought; inconsistent governance has several causes, and the one an architecture reaches is nobody having named where a rule gets set, which is a design decision rather than a migration.

So the architecture a board asks for is not a diagram of servers. It is an answer to a plainer question: what do we have, and who answers for each part of it. That answer is the one architecture the board can see, and it is what this page holds.

What This Page Decides, and What It Hands Off

This page decides one thing: what the estate looks like when the whole of it is drawn once, which layer each recurring argument belongs to, and which guide settles that argument. That covers the four layers, the two through-lines, the order the layers are worth fixing in, and what happens when two layers give a platform team conflicting instructions.

It does not choose your integration tooling. Which pattern moves data, which Azure component carries it, and what governance each component needs are integration-layer decisions, and they are set out in Microsoft Integration Architecture for Large Enterprises: A Reference Guide for Regulated Sectors. Read that guide for the components; read this one for where they sit relative to everything else.

What the estate agrees to before anything is connected is not settled here either. Which system is believed when two disagree, what one system promises another, who is accountable for that promise, how fresh the receiving side has to be, and which options compliance has already closed are the subject of the companion guide Data Integration Discovery: What to Define Before Connecting Enterprise Systems, which owns them in full. That page is published separately and is named here without a link because it was not live at the time of writing.

Nor does it describe the standing work once the estate runs. Who answers for a dataset, what a failed overnight load means for the analysts downstream, how a regulator’s question about a figure gets answered, and which checks run on a schedule belong to the companion guide Data Pipeline Operations: Who Owns the Number, and What Happens When the Run Fails. That page publishes separately too, and carries no link for the same reason. The division between the three is worth stating once, because it is the division that keeps this page a parent: discovery decides what the estate agrees to, this page decides what the estate IS, and operations decides what a steward does about it on a Tuesday.

Two more hand-offs, on purpose. The platform decision between Azure Data Factory and Microsoft Fabric, including the combined pattern, belongs to the companion guide Azure Data Factory vs Microsoft Fabric, which is not yet published and therefore appears by title alone; the architecture described below holds on either, so this page neither makes that choice nor waits on it. And the consumption layer, meaning the semantic model, row-level security and who certifies a dataset, is owned by Enterprise Reporting System Design for Microsoft Environments: Architecture Decisions That Determine Whether It Scales. This page names the consumption layer’s job and stops at its boundary.

One thing this page is deliberately not. It is not an assessment of what you already have. Where the estate is undocumented and the first need is a measured picture of how bad it is, that work runs before an architecture can be written, and Data Integration Risk Consulting for Regulated Enterprises sets out what such an assessment inventories and produces.

The Four Layers of a Microsoft Data Estate

Estate arguments stay unresolved while the thing being argued about is left vague. Splitting the estate into four layers gives each recurring argument a place to land, because the layer decides who settles it and where the answer is read. The four layers are sources, integration, operations, and consumption.

The table states, for each layer, the question that layer settles, the one decision that settles it, and the surface an auditor or a new starter opens to check the answer. The third column is the one estates skip, and skipping it is what turns an architecture into a picture.

Layer The question this layer settles The decision that settles it Where the answer is read today Who owns the depth
Sources Which system is believed about this field, and what does it promise the systems downstream A written system-of-record designation per entity and per field, with a named person against it The governance domain and data product entries in Microsoft Purview Unified Catalog, which carry the designation and the name against it; the source system’s own export of the entity evidences the field values in dispute and nothing about who is believed The Data Integration Discovery guide
Integration What moves the data, on which pattern, under which controls One pattern and one carrying component per class of movement, chosen once and written down The pipeline and connection inventory in the Azure Data Factory monitoring experience or the Fabric monitor, and the published-API inventory in Azure API Management, each covering only the movement those services carry The Microsoft Integration Architecture reference guide
Operations Who answers for a dataset once it runs, and what a failed run means downstream One named accountable owner and one named steward per dataset, each with the date the name was last confirmed The ownership register, held where the organization already keeps controlled records, which is where the two names and their confirmation dates are read; the data quality steward permission list in Purview Unified Catalog shows who holds the steward permission and not who accepted a defect The Data Pipeline Operations guide
Consumption Which numbers are certified, and which rows a given reader may see One certified model per subject area, with a named certifier accountable for it The endorsement and certification state in the Power BI service, which records which models carry certification; which rows a given reader sees is read from the security defined on the model itself The Enterprise Reporting System Design guide

Three properties of that table carry more weight than its contents.

The rightmost column is a routing instruction, not a reading list. Each layer’s depth belongs to one guide, and the reason a parent page names four owners instead of answering four subjects is that the answers are long and they change at different rates. A page that summarizes all four goes out of date whenever any one of them changes.

Column four names the surfaces an auditor can open before the meeting. A layer whose answer exists only in a slide is an intention, and the difference shows up the first time a regulator or an internal auditor asks to see the answer rather than hear it. Where the drawn architecture and the running estate disagree about what feeds a dataset, the drawing is the weaker claim, because a pipeline’s actual source list sits in the Data Factory or Fabric monitoring experience for the auditor to open, while a diagram records what an architect meant at the time.

The precedence rule is worth settling before a disagreement forces it. When two layers give a platform team conflicting instructions, the layer nearer the source governs what a number MEANS and the layer nearer the reader governs how it is PRESENTED, so a definition dispute travels upstream and a formatting dispute does not. Where a compliance constraint is one of the two instructions, the constraint closes the option at every layer at once: a residency commitment or a classification rule outranks a preference expressed at any of the four, because the exposure outlasts the preference. And where the constraint and the source designation genuinely conflict, meaning the system named as the record cannot hold the data where policy requires it to sit, that is not an architecture decision at all. It is a finding, and it goes back to that entity’s business owner with the option set already narrowed.

Two layers, sources and operations, are owned in full by the two companion guides named in the table, and their treatments are deliberately absent here. What the estate view owes them is the boundary: sources hands operations a named owner for the entity, and operations hands consumption a dataset with its own named owner answering for that dataset’s state.

Retrofitting these four layers behind an estate that has been running for a decade is different work from drawing them at the start, and the order the layers get attention in changes what the retrofit costs. Talk to a senior Microsoft systems integration architect

Governance and AI-Readiness: The Two Through-Lines

Governance and AI-readiness are not a fifth and sixth layer. They are properties every layer either has or lacks, so an estate can pass a governance review at the consumption layer and fail the same review at the sources layer in the same quarter.

Governance is the question of where a rule is SET, separately from where it is ENFORCED. Many estates enforce in several places and designate no place where the rule is set, so a classification decision gets re-made as each enforcement point is built, once in a catalogue, once again inside a pipeline’s filter, and a third time inside a report’s row-level rule, and the three drift apart. Microsoft’s data strategy guidance names the target state directly: “Security and governance policies are defined once and applied consistently, instead of being recreated and enforced differently across multiple tools.” The operative half of that sentence is “defined once”. A rule defined in three places has three owners, and whoever changes it last wins by accident.

The estate-level control surface for this is the catalogue rather than any one workload. The same guidance states that “Purview provides a centralized data catalog, data classification, and policy enforcement across your entire data estate”, and that “The data can be in OneLake, Azure, on-premises, third-party SaaS, or other cloud platforms”. The second sentence is a scope statement, and it is the reason the catalogue is the estate’s control surface: its reach extends past the analytics platform, so an organization whose source systems sit outside Azure still has one place to apply the catalogue’s classification and policy rules. What a catalogue does not do is decide what the rule should be. Inside what compliance has already closed, that decision belongs to the business owner of the entity, at the sources layer.

Operating detail for governing an analytics estate, meaning the ownership models, the control set and who sits on which body, is treated in full in Enterprise Analytics Operating Model: A Governed Power BI and Fabric Framework. The readiness gates that have to be cleared before Purview itself can carry a governance model are set out in Microsoft Purview Deployment Guide: A Readiness Playbook for Regulated Enterprises. The estate view assumes both and repeats neither.

AI-readiness is a property of the data, and its test does not mention AI. An AI-ready data estate is one where an assistant reaching the data inherits the same answers an analyst would get by asking a colleague, and there are three of those answers. Which copy is the record. Who is allowed to see which rows. What the field means. Where all three are written down and readable by a service the assistant is connected to and permitted to read, the assistant reads them; where any of the three lives only in an analyst’s memory, there is nothing for it to read, and one of the things it does instead is produce a confident answer built on whichever copy it reached first.

Three checks make that concrete, and none of them requires an AI project to run:

  • The record check. For one entity two systems both hold, open the catalogue and see whether the designation is recorded there instead of agreed in a meeting. A missing entry puts the gap at the sources layer.
  • The access check. Take one restricted field and read where its restriction is written down: on the entity in Purview Unified Catalog, or inside one report’s own filter definition in the Power BI service. Where only the report holds it, each new consumer re-implements the restriction, and an assistant is a new consumer nobody reviewed.
  • The meaning check. Open the catalogue entry for one certified dataset and read whether an analyst who has never used it could tell what a column means. Where the description repeats the column name, or where the dataset has no catalogue entry at all, the meaning is written nowhere a service can read.

The platform side of this is more capable than it was, and it is worth stating precisely instead of enthusiastically. Microsoft’s What is Microsoft Fabric? overview records that “Fabric provides centralized data discovery, access control, and governance capabilities, helping organizations manage data access, sharing, and compliance consistently across workloads”, and What is OneLake? adds that “Tenant-level policies automatically protect any data that lands in OneLake for security, compliance, and data management”. Both statements are bounded, and the bound is the useful part: the first is scoped to workloads, and the second is scoped to data that lands in OneLake. Data still sitting in a source system that nothing has ingested falls outside both, so an organization part-way through a consolidation has some data governed by tenant policy and some governed by whatever its source system does, and the architecture has to say which is which per subject area.

Where an executive team is being asked to fund AI on top of an estate whose three checks have never been run, the checks are the conversation worth having first. Talk to a senior Microsoft systems integration architect

Where Data Factory, Fabric and Power BI Sit in the Four Layers

The question arrives as a comparison and it is not one. Azure Data Factory, Microsoft Fabric and Power BI get discussed as though an architect picks one, and at estate level they occupy different positions, which is why an organization can run all three and still have no architecture.

Data Factory is a mover at the integration layer. It carries data between places on a schedule or a trigger, and it is one of the places the pattern chosen at the integration layer becomes a running thing. The name appears twice in the Microsoft estate, once as the standalone Azure service and once as a workload inside Fabric, and an architecture document that never says which one it means is buying an argument for later.

Power BI is the consumption layer’s surface, and it is also a Fabric workload. That dual position is the source of most of the confusion, because a Power BI decision about how a report’s model is built is a consumption-layer decision whether or not the tenant runs Fabric at all.

Fabric is not a third thing at the same level. It is a platform that carries the integration and consumption layers at once, including the store the integration layer lands data in, which is exactly why “Data Factory or Fabric” is a badly formed question at estate level and a well formed one at the tool level. Microsoft’s overview states that “All Fabric workloads operate over OneLake, a unified logical data lake built on Azure Data Lake Storage”, and that “OneLake enables shared access to data across workloads without requiring data movement or duplication”. The architectural consequence is narrow and real: one stored copy can be read by several engines without being moved again. It is not a claim that copies stop existing, and an organization that keeps making extracts will keep having extracts.

The store itself is documented as answering a problem an estate of accumulated extracts arrives at. Microsoft’s OneLake documentation records that “Before OneLake, organizations often created multiple lakes for different business groups, which led to extra overhead for managing multiple resources”, and that “This siloed approach made it difficult to collaborate across teams, slowed down data projects, and increased the risk of duplication”. Duplication at that level is what every report a snowflake is built on, described from the storage side rather than from the report. The stated design answer leaves ownership where it already was: “This model maintains data ownership and enables federated governance, while still allowing authorized users to discover and use data without friction”. A finance team does not stop owning finance data by putting it somewhere a governed reader can reach.

One naming collision is worth heading off, because it breaks conversations. Inside the storage there is a second and different layering. Microsoft’s guidance on how to Understand medallion architecture for Fabric with OneLake states that “The three medallion layers are: bronze (raw data), silver (enriched data), and gold (curated data)”, that “The goal of medallion architecture is to incrementally improve the structure and quality of data”, and that “It’s the recommended design approach for Fabric”, with the implementation instruction to “Keep each layer separated in its own lakehouse or warehouse in OneLake, with data moving between the layers as it’s transformed and refined”. Those three layers are a QUALITY progression sitting inside the integration layer, in the store that layer lands data in. They are not a smaller version of sources, integration, operations, and consumption. A design review where one participant means the estate’s four and another means the storage’s three produces agreement that does not survive the week.

Which platform carries the integration layer, and whether the answer is Azure Data Factory, Microsoft Fabric or both in a combined pattern, is a tool decision with its own criteria, and the companion guide Azure Data Factory vs Microsoft Fabric owns it; that guide is named without a link because it is not yet published. Where that decision is already made in Fabric’s favour and the open question is who builds it, Enterprise Microsoft Fabric Development & Integration Services describes that work.

From Point-to-Point to a Managed Estate: The Order That Holds

The instinct is to start with the platform, because the platform is the part a sponsor can buy. That order fails in a way worth naming: the engineers land a capable platform, point the existing extracts at it, and arrive at the same disagreements between the same numbers, now hosted somewhere newer. The estate was not the tooling. It was the set of unrecorded decisions the tooling carried.

Microsoft’s data strategy guidance is unusually blunt about the size of the move, and the sentence is worth quoting because it contradicts how these programmes get funded: “Unifying the data platform is an investment in capability, not a wholesale replacement of every system.” An architecture that requires every source system to be replaced first has no first step, and a programme with no first step becomes a document.

The order that holds starts away from the platform.

  • Pick one subject area, not the estate. Customer, or claims, or grant awards. One subject area is small enough that its business owner can be in the room and large enough that fixing it changes what an executive sees. An estate-wide programme starts with an inventory phase whose end no sponsor can define.
  • Name the record before anything moves. For the entity that subject area turns on, the system-of-record designation gets written down with a person against it. That designation is the discovery work the companion Data Integration Discovery guide owns in full, and it is the input this order depends on rather than a step this page describes.
  • Build the new path beside the old one, and leave the old one running. The new path reads the named record, lands in the governed store, meaning whichever store the estate has designated for the subject area and set its classification and retention rules on, and produces a model an accountable owner certifies. Nothing is switched off while it is being proven, which is what keeps the retrofit reversible.
  • Cut over on a certification, not on a date. The old point-to-point link retires when its consumers read the new path and the certified model carries the number, and both halves of that are checkable, though where they are read depends on how each was built: consumers are read from the workspace access list and from every other route the report is shared through, and the old link’s remaining traffic from the pipeline run list in the Fabric monitor or the Data Factory monitoring experience when the link runs there, and from whatever schedules it when it does not.

Two disagreements are predictable, so the rules for them belong in the plan rather than in the incident. Where a subject area’s business owner and the platform team disagree about which subject area goes first, and neither is held up by an unlisted old link or an unsettled compliance question, the one with a named record owner goes first even when it is the smaller prize, because the pattern gets argued on the first one and the second starts from the argument already had. And where the plan’s order and a compliance constraint disagree, the constraint wins and the order changes around it, because a subject area whose data cannot legally sit in the governed store is not a candidate for this sequence at all until that is settled.

The stopping condition for a subject area is a state, not a milestone. It is finished when the certified model is the one an executive quotes, the old link has no consumers, and the ownership register carries a name and a confirmation date for the dataset. What that register contains and how the confirmation cadence gets set is the standing work owned by the companion Data Pipeline Operations guide, and it begins for a subject area on the day that subject area cuts over.

What this order does not do is compress. Each subject area costs what it costs, and the honest version of the argument for doing it at all is that the second one costs less wherever it can reuse the first one’s pattern, store and governance decisions, and costs what a first costs wherever its classification, its sources or its latency rule that reuse out. Delivery is senior and US-based.

When This Is Not the Work You Need

Some readers should stop here, and saying so is cheaper than selling an architecture to an organization that needs a different conversation.

Where the platform team cannot list the pipelines and extracts currently running, the sequence runs the other way. An architecture drawn over an unknown estate is an architecture of the parts an engineer happened to remember, and the assessment work described earlier is what produces the list first. The estate view waits for it.

If the whole estate is one source, one warehouse and a handful of models read by the analysts who built them, four layers is more structure than the problem has. The layers earn their cost when datasets have consumers who did not build them and when a number leaves the analysts who produced it.

Where the real complaint is that two dashboards disagree and the pipelines are sound, the problem sits in the consumption layer and the reporting design guide named above settles it. Redrawing the estate is an expensive way to fix a modelling disagreement.

And where the decision genuinely on the table is which platform to buy, this page is upstream of that and deliberately does not answer it. The architecture holds on either platform, which is what makes the platform question answerable later rather than first.

What is left is the organization this page was written for: connections built one at a time over years, an executive asking why two systems report different numbers, a proposal on the table to put AI on top of it, and no drawing of the estate that the finance director trusts. i3solutions has been a Microsoft partner since 1997, and the cheapest next step is to take one subject area and write down its four layers, naming who answers for each. Talk to a senior Microsoft systems integration architect

Frequently Asked Questions

What does a modern Microsoft data integration and analytics architecture look like?

It looks like four layers with a named owner against each, plus two properties that cross all four. The four layers are sources, integration, operations, and consumption. Sources settles which system is believed about a given field and what that system promises downstream, written as a system-of-record designation per entity and per field with a person named beside it. Integration settles which pattern carries each class of movement and which component runs it, chosen once and recorded. Operations settles who is answerable for a dataset in production and what a failed run costs the analysts reading from it, held as one named owner accountable for the dataset and one named steward running its checks, both dated when the names were last confirmed. Consumption settles which numbers carry certification and which rows a given reader may see. The two properties crossing all four are governance and AI-readiness, and treating either one as a fifth layer is how an estate acquires a governance function that reaches the reporting end and nothing behind it. What separates an architecture from a diagram is one more column: for each layer, the surfaces an auditor opens to check the answer, among them the governance domain and data product entries in Microsoft Purview Unified Catalog, the pipeline list in the Azure Data Factory monitoring experience or the Fabric monitor, the ownership register, and the endorsement and certification state in the Power BI service.

How do ADF, Fabric and Power BI fit together in one architecture?

They are not alternatives, and running all three is entirely compatible with having no architecture. Azure Data Factory is a mover at the integration layer, carrying data between places on a schedule or a trigger. Its name is used twice inside the Microsoft estate, for the standalone Azure service and for a workload of the same name inside Microsoft Fabric, so an architecture document has to say which of the two it means. Power BI is the surface at the consumption layer and a Fabric workload as well, so a decision about how a report’s model is built lands at the consumption layer whether or not the tenant runs Fabric. Fabric is positioned differently again: it reaches across the integration and consumption layers together, including the store the integration layer lands data in, which makes a choice between Data Factory and Fabric a sound question at tool level and a poorly formed one at estate level. Microsoft’s Fabric overview records that its workloads all run over OneLake, one logical lake built on Azure Data Lake Storage, and that those workloads share access to what is stored there without moving or duplicating it. Keep the consequence as narrow as the source keeps it: a single stored copy becomes readable by several engines, which is not the same as extracts ceasing to be made.

How do we get from point-to-point integrations to a managed estate?

One subject area at a time, and the first step is not the platform. Choose a single area such as customer, claims or grant awards, narrow enough to put its business owner in one room and consequential enough that repairing it changes a figure an executive reads. Before anything is moved, write down the system-of-record designation for the entity that area turns on, with a person against it. Then stand the new path up next to the existing link and leave that link running, so the retrofit can be reversed while it is still being proven. Retire the old link on a certification and not on a calendar date: it goes once its consumers have moved to the new path and the number comes from the certified model, and each of those two halves is readable, the consumers in the workspace access list and in every other route the report is shared through, and the remaining traffic wherever the old link runs, which is the pipeline run list in the Fabric monitor or the Data Factory monitoring experience only when the link was built there. Microsoft’s data strategy guidance treats a move of this size as an investment in capability and not as a wholesale replacement of every system, which matters because a programme conditioned on replacing every source system first has nowhere to begin. Where the plan’s sequence and a compliance constraint pull against each other, the constraint decides and the sequence re-forms around it.

What makes a data architecture AI-ready?

An estate is AI-ready when an assistant reading its data arrives at the same three answers a colleague would give an analyst: which copy is the record, who may see which rows, and what a field means. Where those three are recorded somewhere the assistant is connected to and permitted to read, the assistant reads them. Where one of them sits only in an analyst’s memory, one of the things it does instead is answer confidently from whichever copy it reached first. Three checks test this with no AI project running. First, open Purview Unified Catalog and look for the system-of-record designation on one contested entity, rather than accepting that it was agreed in a meeting. Second, take one restricted field and find where its restriction is written down, on the entity in Purview Unified Catalog, or in the filter definition of a single report in the Power BI service, because a restriction held only by the report gets rebuilt by every later consumer and an assistant is a later consumer nobody reviewed. Third, open one certified dataset’s catalogue entry, or find that it has none, and judge whether a newcomer to it could tell from the description what a column holds. Microsoft’s OneLake documentation applies tenant-level policy automatically to data that has landed in OneLake, and the boundary of that sentence is the point: data sitting in a source system nothing has ingested yet is outside it.

Who owns which part of an enterprise data architecture?

Ownership divides along the same four layers, and that split is what makes the estate answerable. At the sources layer the business owner of the entity settles what a given field means and which system is believed about it, within whatever a regulator or a compliance rule has already fixed, and it is not a decision a platform team can take for them. At the integration layer the platform team owns the pattern and the components carrying the data. At the operations layer one named owner on the business side decides what a known defect costs and, where no compliance constraint has already closed the question, whether it can be lived with, while one named steward runs the checks and raises what they find. At the consumption layer a named certifier answers for the model an executive quotes. The precedence rule follows from that ordering: given two layers instructing a platform team differently, meaning is governed from the layer closer to the source and presentation from the layer closer to the reader, which sends a definition argument upstream and keeps a formatting argument where it started. A compliance constraint behaves differently again, closing the same option at all four layers at once, because exposure lasts longer than a preference.

Two systems report different numbers. Is that an architecture problem or a reporting problem?

It is a reporting problem only when both reports read the same certified model and still disagree, and that case gets settled where the difference actually sits, in the model or in what each report does with it. The other case, and the commoner shape of the complaint, is a disagreement older than either report: two extracts of one source were taken at different hours, filtered on different rules, and nobody had designated which system was believed about the field in dispute. That is a sources-layer problem dressed as a reporting one, and reconciling the two reports treats the symptom for a quarter. Telling the two apart starts with one question. Ask which model each report reads, then open the endorsement and certification state in the Power BI service for each model named. Two different models put the argument upstream of the consumption layer, with the business owner of the entity. One shared certified model puts it in how each report filters or aggregates, which the reporting architecture settles.

Related Reading

About the Author

Michael Branson co-founded i3solutions and brings executive, operational, and technical perspective to organizations running complex, secure, and mission-critical Microsoft estates. He works with enterprise teams on the architecture and governance decisions that determine whether a data investment holds its value.