Data Readiness for AI: The Four Tests a Dataset Has to Pass

By Michael Branson

Quick answer. An assistant pointed at the wrong table answers fluently and wrongly, and the meeting cannot say which table or why. Before one AI use case is allowed to read one dataset, four data readiness tests run on it, so the meeting has a record to read instead of a guess: ownership, quality, access, and context. Each is read from a named surface, the ownership register for the owner’s yes and the roles pane on the item for the identity’s reach, and a failed ownership or access test stops the use case whatever the other two show.

The pilot works, and that is what makes it hard to argue about. An assistant has been pointed at the customer table, it answers in a few seconds, and the answers are wrong in a way the room cannot characterize. A regional director asks why two accounts from last quarter came back under the wrong owner. The project manager cannot say, because the question is not about the model.

Somebody says garbage in, garbage out. The meeting agrees, and the next three weeks get spent cleaning the columns the loudest person believed were dirty. The phrase is true and it is unusable: it names a class of problem without naming an owner, a test, a threshold, or the record that would show who agreed to any of it, and the four tests below are what supply them.

Underneath the phrase sit four separate questions, and an organization holding an ownership register and a quality scan has half-answered two of them and never written the others down. A name on the register is not the owner’s recorded yes for this table being read this way. A product-level quality score does not say which columns have to be right, and to what floor, before an answer built on them is relied on. Nobody has named the identity the model reads as, or checked its reach against the people the answers go to. And nothing on file tells a model what a value in a column named STAT_CD means.

This page turns those four into tests a named person can run against one dataset, says what a pass looks like for each, names the surface the answer is read from, whether that is your own register or the platform’s own role list, and gives the rule for what happens when two of them disagree.

What This Page Decides, and What Other Pages Own

Before a dataset is handed to an assistant, one decision belongs here: whether that specific structured dataset is admitted to that specific AI use case, and on whose authority. That covers the authorization to add a new consumer, the quality baseline written for that consumer, the identity the AI system reads as, and the metadata a model needs to interpret the columns it touches.

It does not build your data operating model. Who is accountable for a dataset day to day, how a steward’s checks are scheduled, what happens the morning after a failed load, and what the weekly quality and cost figures are belong to the companion guide Data Pipeline Operations: Who Owns the Number, and What Happens When the Run Fails. This page reads that operating model’s ownership register as an input and takes a decision the register does not contain.

Which system counts as the source is a second hand-off. System-of-record designation, the data contract between two systems, and the constraints compliance imposes before an architecture is chosen belong to the guide Data Integration Discovery: What to Define Before Connecting Enterprise Systems. A dataset whose source is still contested is not ready for an AI use case for reasons that have nothing to do with AI.

Microsoft 365 content is a different readiness problem with a different surface and a different owner. Whether a tenant can safely switch an assistant on is set out at Is Your Microsoft Environment Ready for Copilot? What Must Be True First. The remediation program for oversharing across files and permissions is The Copilot Data Governance Fix, and the SharePoint half of that work is covered in SharePoint Modernization and Copilot Data Readiness. Files and their permissions belong to those pages. Tables, columns and their consumers belong here, with the register that records who authorized each one.

The whole-estate picture, sources through integration and operations to consumption drawn as one architecture, is the parent guide Enterprise Data Integration and Analytics Architecture, and this page is one gate inside it. Whether a use case needs Graph connectors, agents or custom integration is settled in Copilot Integration Architecture; that page chooses the integration path, and this one decides whether the data at the end of it is fit to be read. The policy frame those pages sit under, who approved an AI to work this way and how you would show later what it did, is Enterprise AI Governance for Microsoft Environments.

Why “The Data Is Not Ready” Names Four Different Problems

Reporting taught everyone to treat data quality as one score on a dashboard, and when an AI use case arrives that habit is what makes readiness confusing. A dashboard and an assistant’s answer fail differently, and the difference is whether an analyst sees the number before anyone acts on it.

An analyst reading a dashboard carries context the dashboard does not. When a regional total comes back far above last month, the analyst knows the region and stops. A manager reading an assistant’s answer has nothing to stop on: the sentence is composed with the same fluency whether the number behind it is right or wrong, and the part of it that came from a stale row looks exactly like the part that came from a current one. The reader’s own check, which reporting has quietly depended on for two decades, is gone.

Readiness is four tests, not one score on a dashboard. Ownership, quality, access and context each fail in a distinct way, and each has a distinct fix.

Ownership fails on a missing authorization, and a name on the register does not supply one. A large estate can produce a name for a table. Fewer can produce a record of that person agreeing that the table may be read by this assistant, for this purpose, on terms they set. A name is who to email when it breaks. An authorization is what an auditor asks to see. The fix is the six-field authorization record in the ownership register, naming the approver, the purpose, and the condition on which the yes lapses, so a later reuse of the same table arrives as a new decision and not as an assumption.

Quality fails at the column a use case depends on, and a product-level score can hide that column. A dataset good enough for a monthly report can be unfit for an assistant, and the reverse happens too. The score is an input to the decision and is not the decision, because a score does not know what the use case depends on. The fix is a baseline at column level. It names the dimension that matters here, the floor, and the action a breach triggers.

Access turns on which identity does the reading, and permissions are where that gets checked. The question is not whether permissions exist. It is which identity the AI system reads as, and whether that identity’s reach matches the audience for the answers. An assistant reading through one wide service account hands every reader the reach of that account. The fix is to name the identity and read its reach from the roles pane on the item, not from a change ticket.

Context is translation work, and the documentation is what it leaves behind. A model has your column names and your values and no colleague to ask. Where a column is named STAT_CD and carries the values A, I and P, a person asks somebody down the hall. A model produces a fluent sentence about active and pending accounts, right or wrong, with nothing in it to show which. The fix is a definition a stranger can act on, written on the glossary term the catalog attaches to the data product, with the allowed values spelled out in the same record.

The Four Readiness Tests, and Where Each Answer Is Read

The four tests below are run per use case, against the datasets that use case reads, and the ownership and access tests are settled before work is spent on quality or context. Skipping the last column is what turns a readiness review into an opinion, because a verdict with no named surface behind it cannot be checked by anybody else, and an audit has nothing to read.

Test The question it settles What a pass looks like Where the answer is read
Ownership Who authorized this dataset being read by this use case, and on what terms One named accountable person has recorded a yes for this consumer and this purpose, with the condition on which the yes lapses Your own ownership register, joined to the owner and access-policy fields on the data product in Microsoft Purview Unified Catalog
Quality What this dataset has to be right about for this use case A written baseline naming the dimensions that matter here, at column level, with a floor and a stated consequence below it The data quality scan results for that data product in Purview Unified Catalog, read per column and not at the aggregate
Access Which identity the AI reads as, and how far that identity reaches The identity is named, and its reach has been read and dated from the platform’s own permission surface The OneLake security roles pane on the Fabric item; for a store outside Fabric, the admin surface that lists each role assignment, named before the test is run
Context Whether a model can interpret the columns without a person in the room Every column the use case touches has a definition a new joiner could act on, and the ambiguous ones are mapped to one concept The glossary terms and critical data elements attached to the data product in Purview Unified Catalog

Two of the four are gates and two are conditions, and the difference decides what a failure means. Ownership and access are gates: a missing authorization record, or an identity that reaches further than the audience, stops the use case, and work on the other two does not change that. Quality and context are conditions: they fail in degrees, and a partial failure narrows the use case instead of stopping it, which is why the baseline is written per use case and not once per dataset.

Settle the precedence before a disagreement forces it into the open. Where the ownership test passes and the access test does not, access governs what happens next, because an owner’s yes authorizes a purpose and does not widen what an identity may read. Where ownership is the test that fails, no access answer revives the use case, because a gate is a pass or fail criterion and a failed one is not offset by anything. If the context test fails while the quality test passes, context governs, because a column nobody can define cannot be scored against a meaningful rule, and a clean score on an undefined column is a measurement of nothing. A breached quality floor is not raised by a definition. When both failing tests are conditions, the narrower of the two answers settles the scope of the use case, and the accountable owner is told on the record which test produced the narrowing, so the next conversation is about one thing.

One objection to running four tests on every dataset is that it will not scale. Across a whole estate that is true, and no pass runs across a whole estate: each covers one use case and the tables it touches, which for a first assistant is a handful and not an estate. The scaling cost arrives later and lands on context: the glossary term and the record of allowed values written for each column, the mapping where one concept carries two column names, and the domain owner who can say what a value means. It lands there because definitions written for one use case are part of what the next use case inherits.

Ownership: Who Authorizes a New Consumer

Ask who owns the customer table before an assistant reads it, and an estate with a register can produce a name. Ask whether that person agreed the table may be read by an assistant, for a stated purpose, and the room goes quiet. Those are two different questions, and the second one is what an auditor asks to see.

Why the second is harder has little to do with paperwork. The group holding write access to a dataset and the group that can rule on what its values mean are different groups, and an AI use case adds a third party who can do neither. The ownership test decides which of those two says yes, and puts the fact that they did on the ownership register where a stranger can find it.

The platform gives that yes somewhere to live. Data products in Unified Catalog sets out the flow it expects: “With data products, a user finds the data product and requests access to the data product. After approval, they get access to all the associated data assets.” The second sentence carries the consequence: an approval granted on the product reaches every asset grouped inside it, so the yes that goes on the record is wider than the single table the use case named. The same page bounds the object too, since “A governance domain can house many data products but a data product is managed by a single governance domain and can be discovered across many domains”, which is where the approver’s authority comes from and where it stops. What no field on that object supplies is the purpose the yes was given for, and purpose is what makes a later reuse a new decision instead of an assumption.

The authorization record holds six fields per dataset and per use case: the dataset, the use case, the purpose in one sentence, the accountable person who approved it, the date, and the condition on which the approval lapses. It belongs in the register the operating model already keeps, beside the approver and access-policy fields the data product carries in Purview Unified Catalog, so an auditor reading the catalog and an auditor reading the register are shown the same two names. That last field is the one that gets left out and the one that does the work. An approval given for a pilot with, say, twelve users is not an approval for the same assistant in front of four hundred, and the lapse condition is what turns that into a scheduled conversation instead of a discovery.

Two ownership questions arrive in these reviews, and each one has an answer worth agreeing before it is asked. When an executive wants a dataset an accountable owner has declined, the owner’s decision stands and the escalation goes to whoever the owner reports to, in writing, because an override on the record is defensible later and an override applied quietly is not. When a dataset has an owner who has left, the use case waits: the ownership register names a person nobody can now ask about the purpose, and naming a replacement is faster than arguing about the risk. Who that replacement should be, and how the ownership register keeps the name current after a departure, belongs to the Data Pipeline Operations guide.

Quality: A Baseline Written for the Use Case, Not for the Dashboard

The platform will give you a number. Overview of data quality in Microsoft Purview Unified Catalog lists “Out of box rules to measure six industry standards data quality dimensions (completeness, consistency, conformity, accuracy, freshness, and uniqueness).” Six named dimensions is a useful starting vocabulary; the baseline this page asks for is the one defined below, a three-line record for each column, not the platform’s six-dimension score, because the same page’s scoring model is the thing that decides what you see: those rules “are applied at the column level and aggregated to provide scores at the levels of data assets, data products, and governance domains”.

Read the aggregation for what it does to a readiness review, because the score on the report is not the score on the column. A data product carrying, say, forty columns can hold a healthy product-level score while the one column your assistant reasons over is the weak one, because the aggregate is doing exactly what an aggregate does. The baseline for an AI use case is written against the per-column results of the Purview data quality scan, for the columns that use case reads, and the product-level score is context around it, not the verdict.

Which of the six dimensions matter changes with the consumer, and three of them are worth reading against a model instead of against a report. Uniqueness stops being cosmetic: a duplicated customer record shows up in a report as two rows a person can see, and reaches an assistant as two competing answers with no signal about which is current. Freshness stops being a caveat: a stale row is asserted with the same confidence as a fresh one, and no phrasing in the answer distinguishes them. Completeness stops being a gap: where a value is missing, a model can produce a smooth sentence that does not mention the absence, which is worse than a blank cell because the reader never learns there was one.

The baseline is three lines per column the use case reads: the dimension that matters, the floor under which the dataset is unfit for this use, and the action a breach triggers. The last line is the breach criterion, and it is the one that gets skipped. The platform is honest that it will not supply it. The same Purview page describes the notification path as alerts you “Configure … to notify data owners and data stewards if data quality threshold missed the expectation.” Owners and stewards, which is correct and is not the whole path: nothing in that alert reaches the assistant or the people reading its answers, so the baseline names who suspends or narrows the use case while the defect is open, and that line is the control the alert does not supply.

Whether the scan can run against a given store at all is an authentication question. Purview’s own stated limit, on Data Quality in Microsoft Purview Unified Catalog, is that “Currently, Microsoft Purview can only run data quality scans by using Managed Identity as an authentication option”, and a store Managed Identity cannot authenticate to gets the same three lines from a check the owning team already runs. The readiness test cares that the number exists, that a named person reads it, and that the report or scan which produced it is named in the baseline, so the fallback path carries a surface the way the Purview path does. Standing the Purview side of this up, including its own control ownership, is the subject of Microsoft Purview Technology Readiness Services, and the scheduled operating rhythm those checks settle into afterwards belongs to the pipeline operations guide, not here.

Access: Which Identity the Model Reads As

The choice between two shapes is an architecture decision that gets made by accident, before anybody asks which identity the AI reads as, or names it in an access review. In the first, the AI reads as the person asking, so every answer is bounded by what that person could have opened themselves. In the second, the AI reads through one service identity against a curated set, and the boundary moves from the platform into the code the team built around it. Which shape is in force is a control, not a preference. Fabric documents the same pair for the SQL analytics endpoint, at OneLake Security for SQL analytics endpoints: in “User identity mode” the endpoint “passes the signed-in user’s identity to OneLake, and read access is governed entirely by the security rules defined within OneLake”, and in “Delegated identity mode” it “connects to OneLake by using the identity of the workspace or item owner, and security is governed exclusively by SQL permissions defined inside the database.” Both are used in production. Where the answers reach more than one audience, use the first: its boundary is the platform’s own role list, read on the item, while the second shape’s boundary is permissions the team defined, which the auditor has to read as a second artifact.

In OneLake the default is deny. How OneLake security controls data access states that “OneLake security uses a deny-by-default model, so users start with no access to data unless a OneLake security role explicitly grants access”, and that “You can also define data access with row-level and column-level security on tables.” Deny by default, with narrowing available underneath it. That inverts the readiness question a security analyst should be asking: not who is blocked from this dataset, because a user in no role already sees nothing in that item, but who was added to a role, and by whom.

Create and manage OneLake security roles carries the trap a first pass walks into: “When you add a user to a OneLake security role, remove them from the DefaultReader role. Otherwise, they keep full access to the data.” An access review that inspects the narrow role you just created and stops there reads as a pass and is wrong, because the wide role is still attached. The check is both halves, and the evidence is the roles pane on the item rather than a change ticket describing what was intended.

Two further facts change what the test has to look at. Data security in OneLake says that OneLake security “enforces that security consistently across all compute engines in Fabric”, so a model reaching the data through a different engine than the analyst does inherits the same roles, which is the reassuring half. The second, from OneLake Security for SQL analytics endpoints, is that “Newly created SQL analytics endpoints start in delegated identity access mode by default”, and that is the half an access review verifies per endpoint, because a default somebody has to change deliberately is a default somebody has to be shown to have changed.

The identity governance that sits above this, conditional access for AI surfaces and consent for the applications that reach your data, is a separate discipline with its own page: Microsoft 365 Access Governance for AI Readiness. Where the store is Fabric and the work is standing up those roles and the items under them, Enterprise Microsoft Fabric Development & Integration Services for Unified, AI-Ready Analytics is the delivery side of it. What this test owes the use case is one line: the identity the model reads as, named, with its reach read from the platform and dated, and the access review that produced it named beside it.

Context: What a Model Needs to Read a Column Correctly

When a new joiner is given a table with no documentation, they ask the person next to them. A model in the same position produces an answer anyway. The fix is a definition on the glossary term, written for a stranger to act on and not for the team that built the table.

Microsoft’s catalog carries two objects for this, and the pairing is the useful part. Governance domains in Unified Catalog holds glossary terms, “Business terms that provide context and can carry access and data-handling policies that determine how data should be managed and discovered”, so where a term carries the policy, a definition and the handling rule that follows from it live on the same object instead of in two documents that disagree. Alongside them sit critical data elements, “Logical groupings of important columns” that Microsoft marks for stronger governance, and the example it documents is mapping the column names CustID and CID to a single Customer ID concept. Read that example closely, because it is the failure an assistant makes on its own: two column names, one concept, and a model with no way to know they are the same thing until somebody maps them on the record.

Both of those objects sit inside a governance domain, and the domain is where a definition gets agreed once instead of twice. When a term means two things to two departments, the dispute sits between those departments, and no column comment settles it. Write the definition on the glossary term where the people who dispute it can both see it, with the handling policy that term carries, and the argument happens once.

What the test asks for per column the use case reads is modest, and it is written on the glossary term and the critical data element the catalog attaches to the data product, so the evidence and the definition are the same object. What the value means in a sentence a stranger could act on. What the allowed values are, spelled out, so A, I and P stop being a guess. And what the column is not, which is the line that stops a model reasoning about the account status column as if it were the contract status column. Where a concept has more than one column name across the systems that feed the dataset, the mapping is written on the critical data element before the use case goes live, because a model asked to reconcile them will do so silently.

Definitions cannot be written by the data team alone: they are business answers, and a data analyst writing them unaided produces a document that is wrong in the places it matters. The domain owner supplies the meaning, the analyst supplies the column, and the pairing is what makes the definition durable. An audit reads the same object the model reads. Definitions written this way carry forward: where the next use case reads the same columns and the same governed concepts, it inherits them without rework.

When This Is Not the Work You Need

Three kinds of reader reach this point with earlier work still to do before any of these four tests can be run.

Where nobody has yet listed which datasets a use case would read, these tests have nothing to attach to. Running them against the tables somebody happened to remember produces a readiness verdict about a subset with no name, and no record of which datasets were in scope. An assessment of what the integration estate contains, and what is broken in it, is set out at Data Integration Risk Consulting for Regulated Enterprises, and that work runs first.

A second reader arrives with an AI use case that is still a wish and not a described consumer, and the four tests cannot be scoped for them. The authorization record, the column baseline, the identity and its reach, and the definitions are each written against a described consumer. A wish is not one. A test written against “AI” in general returns a document, and a document is what this page exists to avoid producing.

When two systems disagree about which one is right, no readiness verdict will help. That is a system-of-record argument, it predates the assistant, and the Data Integration Discovery guide owns it. Settle the source, then ask whether the dataset is ready.

What remains is the case this page exists for: a described AI use case, a small set of structured datasets it reads, and four questions nobody has yet been asked to answer on paper. i3solutions has been a Microsoft partner since 1997. Start with a scoping conversation: bring the use case, the tables it touches, and the register you keep; that is enough to start on this one use case, no estate-wide review required. Talk to a senior integration architect

Frequently Asked Questions

How do we prepare enterprise data for AI use cases?

When an assistant is about to read a business dataset, run four tests on each dataset that use case reads, per use case and not per estate: ownership, quality, access, and context. Ownership asks whether the accountable person for the dataset has put a yes in the ownership register for this consumer, for a stated purpose, and said what would withdraw it. Quality asks what the dataset has to be right about for this particular consumer, set out per column with a threshold and a named action when the threshold is missed, because an aggregated product score can look healthy while the single column a model reasons over is the weak one. Access asks which identity the AI reads as and how far that identity reaches, taken from the platform’s permission surface and not from a change ticket. Context asks whether a model can interpret the columns with nobody available to ask, which means one usable definition per column and a mapping wherever a single concept carries more than one column name. Ownership and access behave as gates and stop a use case when they fail; quality and context fail in degrees and narrow it instead. Settle the two gates first. An answer that has already reached the wrong reader cannot be withdrawn. Begin with the datasets one described use case reads, because none of the four criteria can be scoped until the consumer is described.

What data quality does governed AI require?

Governed AI requires a baseline written for the consumer, not a general score, because a product-level score can look healthy when the one column the model reads is the weak one. Microsoft Purview’s data quality capability ships out-of-box rules across six industry standard dimensions, completeness, consistency, conformity, accuracy, freshness and uniqueness, evaluated per column and rolled up into scores for data assets, data products and governance domains. That vocabulary is the starting point. The roll-up is the trap, so read the score for the columns the use case touches instead of the figure sitting on the product. Uniqueness, freshness and completeness are the three worth reading again once the consumer is a model. Duplicates land in a report as two visible rows and in an assistant’s answer as rivals it has no way to choose between. Staleness carries no signal at all, because an out-of-date value reads exactly like a current one. Missing values are the hardest of the three, since a blank cell in a report is visible and an absent value inside a generated sentence is not. The baseline runs to three lines for each column the use case reads: which dimension carries the weight here, the score at which the dataset stops being fit for this consumer, and the action a breach triggers. Purview’s own alerting notifies data owners and data stewards when a quality threshold misses expectation, so that third line has to name who suspends or narrows the AI use case while the defect is open.

Who should own data used by AI models?

One named person on the business side is accountable for the dataset, and if that person has not authorized this consumer the name on the register settles nothing. A register that answers who owns the customer table has not yet answered whether the owner agreed an assistant may read it, for a stated purpose, on terms they set. The confusion is structural. Custody and meaning sit in different parts of an organization, so whoever holds the keys to a table is not, by holding them, the person who can say what a value in it signifies. Microsoft Purview Unified Catalog gives the agreement a home. In Microsoft’s words, at Data Products in Unified Catalog, “Each data product has an access policy that determines how users request access, the terms of use for the data, and who should approve access to the data”, and an approval on the product reaches the assets grouped under it. The field no product supplies is purpose, and purpose is what turns a later reuse into a fresh decision. Six fields go on the authorization record in the ownership register, in this order: the dataset, the use case, why it is read, who approved, when, and what would withdraw the approval. Where an executive presses for a dataset the accountable person has refused, the refusal holds and the appeal travels upward in writing, because an override on the record can be justified afterwards and a quiet one cannot.

What is data context and why does AI need it?

When a model reads a column and cannot ask anybody what a value means, what is missing is data context: the business meaning attached to that column, written for somebody who has never seen the table, covering what a value means, which values are allowed, and what the column does not cover. A model works from names and values with nobody to ask. A status column such as STAT_CD, holding A, I and P, gets an answer composed around a guess, and the guess reads exactly like knowledge. Microsoft Purview Unified Catalog holds two objects for this, glossary terms and critical data elements. A glossary term can carry the access and data-handling policy alongside the business definition, and where it does, a meaning and its handling rule sit on one object instead of in two documents that drift apart. Critical data elements are logical groupings of important columns singled out for stronger governance, and Microsoft’s documented example maps the column names CustID and CID onto a single Customer ID concept, which is the reconciliation an assistant would otherwise perform silently. Both objects live inside a governance domain, which is the boundary a definition and its handling policy are agreed within, so a term meaning two things to two departments is settled once. A later use case reading the same columns under the same governed concepts inherits those definitions without rework.

Which identity should an AI system read enterprise data as?

If the answers reach more than one audience, the AI reads as the identity of the person asking, so each answer is limited to what that reader could have opened for themselves; in Fabric’s SQL analytics endpoint that is user identity mode, which passes the signed-in user’s identity to OneLake. Do not use a single service identity reading a curated set there: it moves the boundary out of the platform and into the code the team built around it, and when somebody asks how a figure reached a reader who should not have seen it, the answer sits in that code instead of on the roles pane. The reach is read from the platform and dated in both shapes, with the access review that produced it named beside it. A regulated estate takes the first shape as well, because the boundary an auditor is shown is then the platform’s role list and not the code the team wrote. In OneLake security a user who belongs to no data access role sees nothing in the item, table access narrows further with row or column level security, and security that has been set applies across all engines in Fabric, so a model arriving through a different engine inherits the same roles. Two checks belong in the same pass. Creating the narrow role is only half the change, because anybody left in the DefaultReader role keeps full access, and an access review that inspects the new role alone reads as a pass. An item created with a SQL analytics endpoint begins in delegated identity mode, so the access review confirms the change was made deliberately instead of assuming it.

How is data readiness for AI different from data pipeline operations?

Different question, different moment, different owner, and the difference decides who you ask when an answer comes back wrong. Data pipeline operations, the subject of the guide Data Pipeline Operations: Who Owns the Number, and What Happens When the Run Fails, is the running of a data estate once integration is finished. It answers who is on the hook for the nightly run and what the numbers are. That work is continuous. Data readiness for AI is a decision taken once per use case, and it asks whether one dataset is admitted to one AI use case, on whose authority, against what baseline, read by which identity, and with what meaning attached to its columns. The operating model is an input to that decision, because the ownership test reads the register that model keeps, plus the approver and access policy the data product already carries in Purview Unified Catalog, instead of building either, and a dataset with no operating model behind it fails for reasons the four tests did not create. The practical sequence is to stand the operating model up first where none exists, then run the ownership, quality, access and context tests per use case on top of it, starting with the six-field authorization record. Reversing that order produces a readiness verdict which goes stale the first time a pipeline load fails and nobody notices.

Related Reading

About the Author

Michael Branson co-founded i3solutions and brings executive, operational, and technical perspective to organizations working in complex, secure, and mission-critical environments. He works with enterprise teams on the questions that decide whether a data investment survives contact with an AI use case.