Copyright i3solutions. All Rights Reserved.
Email aski3@i3solutions.com, Phone 703.652.8966
Privacy Policy | Sitemap
Data Integration Discovery: What to Define Before Connecting Enterprise Systems
By Michael Branson | September 9, 2026
Quick answer. The connection got built before anybody wrote down who owns the data or what its fields mean, so when two systems disagree about the same customer there is no document that settles it and the argument reopens every month. Data integration discovery settles five definitions before any platform is chosen: the system of record, the data contract, the owner, the movement requirement, and the constraint set. Each one is a written answer with a named person against it.
Two systems hold the same customer, somebody rekeys between them every morning, and finance and operations produce reports that disagree by enough to matter. The request arrives in plain language. Somebody proposes connecting the systems. The proposal names a tool, because a tool is a thing you can buy, and the meeting that follows is about the tool.
That meeting is early by one step. A connection between two systems encodes a set of decisions whether or not anybody made them deliberately: which system is believed when they differ, what one system promises the other, who is called when the promise breaks, how fresh the data has to be and how much of it moves, and which of those choices a regulator has already constrained. Build first and those decisions get made by whoever writes the mapping, on the day they write it, and they are then expensive to revisit because other work has been built on top of them.
Discovery is the step that makes those decisions on purpose. It is not a questionnaire and it is not an architecture study. It is a short, bounded piece of work that produces five written definitions per data entity crossing a system boundary, and it finishes when those definitions have names against them. That finishing test is checkable against a real artifact: per Microsoft Learn’s Governance domains in Unified Catalog, a governance domain gives domain owners who govern and maintain the domain and its assets, so an entity nobody can name an owner into has not been defined yet, whatever the mapping already does.
Five Definitions to Settle Before a Platform Is Chosen
These are not the estate-level architecture and governance questions. Pattern choice, tool choice, mapping documentation, integration-layer security and failure handling across an estate are answered in full on the live System Integration Best Practices for Microsoft Enterprises page, and the closing section of this page carries the single link to it. What follows sits one step earlier: the definitions that have to exist before any of those choices is worth making.
| Definition | The question it answers | Where the answer becomes visible | Who answers for it |
|---|---|---|---|
| The system of record | Which system is believed when two of them hold the same field and disagree | The written designation, held per entity and per field, and the source system’s own export of that entity for the fields in dispute | The business owner of the entity, with the owner of each contributing system |
| The data contract | What the sending system promises the receiving system about shape, meaning, validity, cadence, change terms and names | The contract document itself, and the catalogue entry that carries its owner and its terms of use, such as a data product in Microsoft Purview Unified Catalog | The owner of the sending system, countersigned by each consumer that exists when the contract is agreed |
| The owner | Who is accountable for the contract when it changes, breaks or is disputed | The named person on the contract, and the governance domain in Microsoft Purview Unified Catalog that carries the domain owner for that data | The named business owner of the data, and not a committee or a team mailbox, because a dispute is settled by a person somebody can call |
| The movement requirement | How stale the receiving system may be, and how much data moves on an ordinary day and on the worst day | The measured record counts and update counts taken from each source system’s own reporting or query surface, written beside the business tolerance they are compared against | The business process owner who feels the staleness, with whoever pays for capacity |
| The constraint set | Which choices are already closed by classification, residency or the access rules of the receiving system | The classification and scan output from the Microsoft Purview Data Map, and the published data location for each service holding the data | The compliance or security owner, with the entity’s business owner |
The rightmost column names a role, and the copy you keep replaces each role with the name of the person who holds it, because a definition owned by a committee is undefined on the day it is tested. The third column names something somebody can open and read, so a definition is checked instead of believed. And where two owners in a row disagree, the constraint set decides first: a residency or classification limit closes an option before any preference is weighed, and only then does the entity’s business owner settle the rest.
The System of Record: Deciding Which System Wins, Field by Field
“Which system is the source of truth” sounds like one question per entity and is several. The CRM creates the customer. Finance owns the credit terms attached to that customer. The service desk holds the support tier. Asking which system owns “the customer” produces an argument; asking which system owns each field produces an answer, because the fields have different histories and different people care about them.
A designation that survives contact with the business answers four questions about each disputed field, in this order. Which system CREATES the value, because creation is the cheapest claim to verify. Which system’s copy the business ACTS on, which need not be the system that created it. Which copy an auditor or a regulator would be shown, which is the question that settles the case when the earlier answers split. And what happens to the other copies once the winner is named: whether they become read-only, whether they keep a local field the winner does not carry, and whether anybody is allowed to edit them.
Where the answers point at different systems, the third question governs. The copy an organization would defend externally is the copy it has to be able to reconstruct, and a designation that names a different winner is a designation somebody will quietly work around. That rule is worth writing into the designation itself, so the next consumer inherits the ruling instead of reopening it.
The designation is written per entity, per field, and it names a person. A useful test of whether it is real: take a field that appears in both systems’ own exports today, name the winner, and ask the loser’s owner what they lose. When the answer is a report, a workflow or a local field nobody mentioned, the designation is not finished. When the answer is nothing, the designation is probably correct and was probably already true in practice.
If the argument about which system wins keeps reopening between your business and platform teams, the useful next step is a conversation with somebody outside the argument. The shape of that work is on the record: i3solutions unified identity and automated provisioning across systems for 125,000 users by treating the interfaces as owned, governed contracts. Schedule a Meeting
The Data Contract: What Two Systems Agree to Exchange
A data contract is the written promise the sending system makes to the receiving one. It exists so that a change on one side is a negotiation instead of an outage, and so that the person who has to answer for a bad value is found before anybody starts looking.
Six things belong in it, and the list is short on purpose, because a contract nobody maintains is worse than none. The SHAPE: the fields exchanged, their types, and which of them may be absent. The MEANING: what each field is, stated in business words, because “status” means something different in a CRM and a service desk. The VALIDITY: the rules a value has to satisfy to be accepted, and what the receiver does with a value that fails them. The CADENCE: when the exchange happens and what the receiver may assume between exchanges. The CHANGE TERMS: how much notice a change requires, who has to agree, and what happens to consumers that cannot move in time. And the NAMES: the person accountable on each side.
That list is a superset of what a workflow build needs, and there is an existing treatment of the narrower case. Where the question is which contracts a specific automated workflow requires across SharePoint, Teams, Dynamics 365 and an ERP, Integrating Automated Workflows with SharePoint, Teams, Dynamics 365, and ERP sets out the inputs a flow build needs before anybody builds it, and this page assumes rather than repeats them. The estate case differs in one respect that matters: an estate contract binds consumers who arrive after it is signed and never negotiated it, so its change terms carry more weight than its shape.
Microsoft’s catalogue has somewhere for part of this to live. Its documentation on Data products in Unified Catalog describes a data product as a business concept with a name, description, owners, and a list of associated data assets, and states that each data product has an access policy that determines how users request access, the terms of use for the data, and who should approve access to the data. That covers the naming, the ownership and the terms of use. It does not cover the shape, the validity rules or the change terms, so the contract stays a written artifact with an owner, and the catalogue entry is where a consumer finds it. A team that treats the catalogue entry as the whole contract has recorded who owns the data and not what the data promises.
The contract is written per system pair and per entity, not per project. Where a contract is written for one project and its ownership ends with the project, the second consumer negotiates from scratch with whoever is available.
The Owner: Who Answers for the Contract, Not for the System
A discovery output can look finished and still leave the contract unnamed. The system has an owner, and that owner can say what the system does. The contract needs somebody who can say what the data MEANS and agree to a change in it, which is a different accountability from the one the system owner holds.
Microsoft states the problem plainly. Its documentation on Governance domains in Unified Catalog says that data ownership is a central aspect of data governance, and that often IT teams store and maintain data assets even though business teams own and use the data, which creates a gap between how data should be discovered and maintained and the teams that actually use it. The same page describes a governance domain as a boundary that enables the common governance, ownership, and discovery of data products and business concepts, and gives that domain owners who govern and maintain the domain and its assets. Its stated goal is a domain owner who manages their own data products and concepts and sets the rules for their access, use, and distribution.
Naming an owner in a catalogue is a record of a decision, not the decision. Three tests separate a real owner from a placeholder. The person can approve a change to the contract without escalating, which means the accountability sits where the authority does. The person is called when a value is disputed, and knows they will be. And the person’s absence is covered: a named deputy, or a stated rule that the contract freezes until they return. Where a candidate passes two of the three, the first governs: without the authority to change the contract, the other two describe a witness, not an owner.
One assignment is worth resisting. A platform team that owns every contract because it owns the integrations has taken accountability for meanings it cannot rule on, and the practical result is that disputes route to the team least able to settle them. The platform team owns the pipe. The business owner owns what travels through it. Where an organization genuinely has no business owner for an entity, that absence is a finding the discovery exists to surface, and it belongs in the output as a named gap rather than being filled by whoever is nearest.
The Movement Requirement: How Fresh, How Much, How Variable
How fresh and how much are treated as implementation details and then quietly decide the architecture. They belong in discovery because they are business facts, and because a business fact somebody measured is a much better input than a preference somebody stated.
Freshness is one question asked well: how stale may the receiving system be before somebody makes a wrong decision. Where the stated answer carries no consequence beside it, what it records is a preference, and the work is getting the real one onto paper. A useful way to find the real one is to ask what happens if the data is an hour old, then a day old, then a week, and note where the person hesitates. The number that goes in the discovery output is a tolerance with a consequence attached, so a later design conversation has something to trade against.
Volume is the question that is skipped, and it is the one an instrument can answer. Two figures per entity are enough to start: how many records exist, and how many change on an ordinary day. Both are readable from the source system’s own reporting or query surface, and the difference between them is what an incremental design has to carry. A third figure matters where the source is lumpy: the largest single day the business can produce, which is a month end, a fiscal close, an acquisition or a bulk load. An integration sized for the ordinary day and met with the worst day is a familiar way to discover that a request budget was a design input.
These two facts are captured in discovery and handed to the architecture conversation, which they inform without settling. The pattern selection, the tooling and the reference architecture for the estate belong to a separate step, and the reference guide for that step is listed under Related Reading.
There is a reason to settle this before a build instead of during one: a movement requirement nobody measured is the input that surfaces late, and it is cheapest to measure while the design is still a diagram. One IT systems analysis for a federal housing agency identified about $1.5 million in savings and led to processing roughly 35 percent faster by finding the real constraints. Schedule a Meeting
The Constraint Set: What Compliance Decides Before Architecture Does
Certain options are closed before anybody weighs them, and finding that out during design review is expensive in a specific way: the work already done was done against the wrong option set. The constraint set is captured in discovery so the architecture conversation starts inside the allowed space.
Three constraints are the ones to capture. What the data IS, in classification terms, because a field carrying regulated content changes what may be copied and where it may land. Where the data may PHYSICALLY sit, because a residency commitment is a hard boundary and not a preference. And who may SEE it once it crosses, because a copy in a second system is a second access surface with its own permission model.
Classification, the first, is measurable rather than remembered. Microsoft’s Data classification in Data Map documentation describes classification as a way of categorizing data assets by assigning unique logical tags or classes to them, based on the business context of the data, and states that the Data Map provides an automated classification capability while you scan your data sources. In the introduction above its Uses of classification section, the same page states “You get more than 200+ built-in system classifications and the ability to create custom classifications for your data.” Accessed 2026-09-09. Assets can be classified automatically when they are ingested as part of a configured scan, or edited manually in the Microsoft Purview governance portal after they are scanned and ingested. So a discovery that asserts an entity carries no regulated content has a way to check the assertion for the sources a configured scan covered, and the scan output is what the discovery output records.
Residency, the second, is published rather than inferred. Microsoft’s Where your Microsoft 365 customer data is stored documentation explains the data residency commitments for Microsoft 365 services and where the data is stored, and notes that tenants in certain Local Region Geographies have access to Advanced Data Residency, which provides more data residency commitments for certain in-scope services. The estate-level catch is in that documentation’s own pointers: the residency position for Azure, and for Dynamics 365 and Power Platform, is documented separately from Microsoft 365. An integration crossing those boundaries has more than one residency answer, and the discovery output carries one line per service rather than one line per estate.
Access, the third, is a question the receiving system’s own permission model answers, and asking it early changes what “integrate” means. Where the receiving system is more widely readable than the sending one, the integration widens access to the data as a side effect, and somebody has to agree to that in writing before it happens rather than discovering it at an access review.
The Discovery Itself: Who Answers, What It Produces, When It Ends
Discovery is the container for the five definitions and not a sixth definition of its own. It is short work, and the reason it is short is that its job is to produce written answers, not to study the estate.
Four kinds of people have to be in the room, and the discovery stalls when one is missing. The business owner of each entity, who settles meaning and system-of-record disputes. The owner of each contributing system, who knows what the system actually holds as opposed to what it was designed to hold. Somebody from compliance or security, who closes options early instead of late. And one person who writes, because an answer nobody recorded is an answer that will be re-litigated.
The output is a small set of documents, one per entity crossing a boundary, each carrying the five definitions and a name against each. Alongside them sit two lists that do more work than their length suggests: the disputes that were NOT settled, with who has to settle them and by when, and the gaps found on the way, which are entities with no business owner, fields two systems both claim to create, and constraints nobody had written down. Those two lists are the honest output of a discovery, and a discovery that produces neither has probably documented the estate rather than interrogated it.
Discovery ends on a condition rather than on a date. It is finished when every entity in scope has five definitions with names against them, or an explicit deferral with an owner. Anything else is study, and study expands to fill whatever room it is given. The scope that keeps this bounded is the set of entities that actually cross a system boundary in the work being contemplated, which is bounded by that work and not by the estate.
Two things this work is not. It is not an architecture assessment: no pattern is selected, no tool is chosen, and no reference architecture is drawn. And it is not a data governance program: the five definitions are per-entity working artifacts, and a standing governance function is a separate commitment, decided on one condition: whether contracts are being reopened faster than their named owners can maintain them one entity at a time. Where they are, the standing function is warranted; where they are not, the per-entity artifacts are enough. i3solutions is entirely U.S.-based. Every i3solutions employee is U.S.-based, every i3solutions project is U.S.-based, and every technology i3solutions delivers is U.S.-based.
When This Is Not the Work You Need
Stop here if the open question is which integration pattern to adopt, or if the integrations are already built and nobody can say what they do. Both are different work, and saying so is cheaper than selling five definitions to somebody who needs a different conversation.
If the open question is which integration pattern to adopt, which Microsoft tool suits which workload, how field mappings and conflict hierarchies get documented, where security controls belong at the integration layer, or how retry and dead-lettering should work across an estate, those are estate architecture and governance questions and they are answered in full in System Integration Best Practices for Microsoft Enterprises. The five definitions set out here are the inputs those answers are built on, and this page assumes rather than repeats them.
Where the integrations already exist and the problem is that nobody can say what they do, discovery is the wrong shape of work. The prior step there is a dependency map and a risk-sequenced triage of what is already running, which is the subject of Microsoft System Integration for Enterprise IT: How Regulated Enterprises Connect Disparate Microsoft Platforms Into a Governed Architecture.
If the scope is a single automated workflow rather than an estate, the flow-level version of these questions is smaller and faster, and Workflow Automation Discovery Questions for Process Owners asks them at that scope; the companion Power Automate API Integration page covers the Power Automate specifics and is not yet published. If the integrations are built and the question is how to run them, the companion Data Pipeline Operations page is being written and is not live either.
And if you want the connections delivered instead of the definitions written, Microsoft System Integration & Data Management describes that engagement.
Two or more systems hold the same data, somebody rekeys between them, a proposal to connect them is on the table, and none of the five definitions is settled in writing. That is the organization this page was written for. i3solutions has been a Microsoft partner since 1997, and the cheapest next step is to take one entity that crosses a boundary today and write its five definitions down, with a name against each. Schedule a Meeting
Frequently Asked Questions
What should we define before integrating enterprise systems?
Five definitions per data entity that crosses a system boundary, each written down with a name against it. The system of record: which system is believed when two of them hold the same field and disagree. The data contract: what the sending system promises the receiving one about shape, meaning, validity, cadence, change terms and names. The owner: the person accountable for that contract, who can agree a change to it without escalating. The movement requirement: how stale the receiving system may be, and the record counts and update counts for an ordinary day and the worst day. The constraint set: what classification, residency and access rules have already closed. The wider estate questions, meaning pattern choice, tool choice, mapping and conflict-hierarchy documentation, and retry and dead-lettering design, are answered on the System Integration Best Practices for Microsoft Enterprises page.
How do we designate a system of record?
Per entity and per field, not per system, because the fields have different histories. For each disputed field, answer four questions in order: which system creates the value, which system’s copy the business acts on, which copy an auditor or regulator would be shown, and what becomes of the other copies once a winner is named. Where those answers point at different systems, the copy the organization would defend externally governs, because that is the copy it has to be able to reconstruct. Then test the designation by asking the losing system’s owner what they lose: an answer naming a report, a workflow or a local field means the designation is not finished. Write the result down with a person’s name against it, and record the rule that settled it, because the same dispute recurs whenever a new consumer appears.
What questions should precede a data integration project?
The questions behind the five definitions, asked once per entity crossing a boundary, in a bounded session rather than a study. Four kinds of people have to answer them: each entity’s business owner, who settles meaning and system-of-record disputes; each contributing system’s owner, who knows what it actually holds; somebody from compliance or security, who rules options out before the design starts; and one person who writes, because an unrecorded answer gets re-litigated. The output is one short document per entity carrying the five definitions with a named owner for each, plus two lists that matter more than they look: the disputes left unsettled, with who settles them and by when, and the gaps found on the way, meaning the entities that turn out to have no business owner, the fields two systems both claim to create, and the constraints that were never written down. The work ends on a condition, when every entity in scope has five definitions with names or an explicit deferral with an owner.
Who should own data contracts between systems?
A named person on the business side of the data, not the platform team that owns the integration and not a committee. The platform team owns the pipe; the business owner owns what travels through it, and routing meaning disputes to the team least able to settle them is the common failure. The same gap is documented: Microsoft Learn’s Governance domains in Unified Catalog documentation (learn.microsoft.com/en-us/purview/concept-governance-domain), in the introduction above its Overview and elements section, states “Often, IT teams store and maintain data assets even though business teams own and use the data.” Accessed 2026-09-09. Three tests separate a real owner from a placeholder: the authority to change the contract without escalating, the standing expectation of being called when a value is disputed, and cover for the person’s absence, either a named deputy or a rule that freezes the contract until they return. Where an entity turns out to have no business owner at all, the gap itself is what gets recorded, named rather than filled by whoever is nearest.
What belongs in a data contract between two enterprise systems?
Six things, kept short so the contract is maintainable. The shape: the fields exchanged, their types, and which may be absent. The meaning: what each field is in business words, because a term like status means different things in a CRM and a service desk. The validity: the rules a value must satisfy to be accepted, and the receiver’s handling of a value that fails them. The cadence: when the exchange happens and what may be assumed between exchanges. The change terms: how much notice a change requires, who must agree, and what becomes of consumers that cannot move in time. Last, the names: the person accountable on each side. A catalogue entry can carry the naming, the ownership and the terms of use, and the shape, validity and change terms stay in a written contract the catalogue entry points at. An estate contract also binds later consumers that had no part in negotiating it, and that puts the weight on the change terms ahead of the shape.
How do we capture latency and volume requirements before choosing a platform?
Measure the volume and interrogate the tolerance, then write both down beside each other. For volume, two figures per entity are enough to begin: how many records exist and how many change on an ordinary day, both readable from the source system’s own reporting or query surface. Add a third where the source is lumpy, meaning the largest single day the business can produce at a month end, a fiscal close, an acquisition or a bulk load. For latency, ask what staleness the receiving system can tolerate before somebody makes a wrong decision, then walk the answer out to an hour, a day and a week and note where the person hesitates, because a tolerance stated with no consequence attached to it is a preference. What goes in the discovery output is a tolerance with its consequence attached, so a later architecture conversation has something concrete to trade against. The anchor for that is the data contract’s own cadence line, defined on this page as when the exchange happens and what may be assumed between exchanges: a tolerance that can be written there as a bound with the decision it protects is a measured requirement, and one that cannot be written there is recorded as unmeasured rather than rounded to a number somebody liked. The volume figures and the latency tolerance are inputs to the pattern and tooling decision and do not settle it.
Related Reading
- Microsoft Integration Architecture for Large Enterprises: A Reference Guide for Regulated Sectors, the pattern and tooling layer the five definitions feed into once they exist
- Data Integration Risk Consulting for Regulated Enterprises, the assessment shape for an estate whose integrations are already built and already carrying audit risk
- The Hidden Costs of Poor Data Synchronization Across Microsoft Systems, the cost argument for the rekeying and reconciliation this work exists to end
- Hire US-Based Senior Microsoft Integration Developers, the staffing route when the definitions are settled and the capacity to build is what is missing
About the Author
Michael Branson co-founded i3solutions and brings executive, operational, and technical perspective to organizations running complex, secure, and mission-critical Microsoft estates. He works with enterprise teams on the integration and governance decisions that determine whether a platform investment holds its value.
