Workflow Automation Discovery Questions for Process Owners

September 4, 2026

Workflow Automation Discovery Questions: What a Process Owner Answers Before the Build

By Michael Branson | August 26, 2026

Quick answer. Workflow automation discovery questions are the five a process owner answers before anyone builds: process truth, exceptions, volume, failure tolerance, and ownership. Answered after the build instead, they arrive as change requests against an automation that already runs the wrong process.

The first request arrives as a favour. Somebody from IT asks you to write down how invoice approval works so it can be automated, and you send back the diagram from the process manual. Six weeks later there is a demo, and the first thing you notice is that it does nothing sensible with what happens at month end.

Nothing was wrong with the build. It matched the document. The document described the process as it was designed, the work runs on the process as it is performed, and nobody in that room was accountable for the distance between the two.

The five questions below close that distance. They are written for the person who owns the work, not for the person who will build the automation, because each of the five ends in a ruling the business side has to give. What a build does with the work is decided in the specification, not in the code, and the specification is yours.

What This Page Answers, and What the Neighbouring Pages Decide

One process has been picked. If that is not true yet, the question in front of you is a different one, and Workflow Automation Decision Framework for IT Leaders is where it is answered: it scores candidate workflows against each other so a backlog can be ordered. Running that scoring as a facilitated exercise across a queue of candidates is its own piece of work, covered in How to Prioritize Automation Processes for Maximum Project Efficiency. Both answer which process. This page starts one step later, with a process already chosen, and asks what has to be true about it before anybody builds anything.

The questions that belong to other people are named here so you can hand them over rather than answer them badly. Whether this should be built on Power Automate or as a custom integration is an architecture decision, treated in a forthcoming comparison, Power Automate vs Custom Integration. Which system holds the master copy of a record, and how two systems agree, is engineering discovery, covered in a companion guide, Power Automate API Integration: Questions Before Development, not yet published. How the automation is designed to behave when a step fails, including retries and recovery, is design work covered in another forthcoming guide, Workflow Error Handling and Recovery. Your job on that one is question four below: you say what the business can tolerate, and the design is built to meet it. How the build itself runs is the territory of Power Automate Development.

Microsoft’s own planning material puts a planning step ahead of the build. Its introduction to Planning a Power Automate project addresses “a business user, an IT pro, or a professional app developer” and sets a plan step before the make step, described as “Identify the who, what, when, and why” and then “Design your new automated process \”on paper,\” and consider various methods of automation.” The five questions here are what the who, what and when look like when the process is a real one with a month end and a difficult customer.

The Five Questions, and What Each One Produces

Discovery goes wrong when it produces a meeting instead of an artifact. Each question below has an output somebody can hold, and a place the answer is read from, because an answer nobody can check is an opinion with a date on it.

Question What it settles The artifact it produces Where the answer is read from
Process truth What the work is when it is performed, as distinct from how it is documented A numbered step list of the process as performed, with the person or system doing each step The people who perform the steps, walked through their last few real items, checked against the record the work already leaves: the request queue, the shared mailbox, the tracking list, or the event log a process mining capability reads
Exceptions Which departures from that step list are real, roughly how many times each occurs, and what is done about each An exception register, one row per departure, each row carrying its trigger, its handling, its rough frequency, and what getting it wrong costs The same performers, working from a bounded set of completed items in the system of record instead of from memory, plus the rejected and reworked items in whatever list already holds them
Volume How many items arrive, how the arrivals cluster, and how many are in flight at once A count per period with its clustering and how many sit in flight together The system the items already sit in: the list, the mailbox folder, the ticket queue export, the finance system’s own report for the period
Failure tolerance What a wrong or missing outcome costs the business, how long the process can run degraded, and who has to be told A tolerance statement per kind of wrong outcome, written in business terms and signed by the accountable lead The process owner with the accountable business lead, tested against the audit, contractual or regulatory obligation the process carries, read from the obligation itself
Ownership Who decides, who approves each decision, who vouches for the data, and who answers when the automation does something unwanted Named individuals with named roles, one per decision point, plus a named process owner, a named data owner, and a named standing owner The delegation of authority the organization already publishes, checked against the approval records in the systems the process writes to

Two rules keep the table from becoming five separate conversations. When two answers conflict, the one drawn from a system record outranks the one drawn from memory, because the record is what an auditor will read; where no record exists, the conflict is written into the register as an open item with a named person against it, and it is not resolved by whoever spoke last. When the exception register and the step list disagree about what the process is, the exception register wins, and whoever holds the step list rewrites it until it accounts for the register.

The five are ordered as they are because process truth comes first and everything after it is an answer about the process the step list describes. Volume counted against a step list that turns out to be wrong is a number about the wrong thing.

Question One: Process Truth, or What the Work Is When It Is Performed

A process old enough to have trained a new joiner exists in three versions. There is the documented one, in a manual written when the system was installed. There is the trained one, taught to each new joiner by the person who sat next to them. And there is the performed one, which is what a team does on a Tuesday when the usual approver is on leave and the customer is important. Automation built from the first version automates a documented process instead of a performed one.

The performed version lives in people’s heads, and the way to get it out is not a workshop whiteboard. Take a bounded window of items the process has already finished, say the most recent fifty in whatever list or queue holds them, and walk three or four of them end to end with the person who handled them. Ask what they did, in order, and where each step’s information came from. Write the answers as a numbered step list, with a column naming who or what performs each step. That column earns its place later: automation takes steps away from named people, and a step whose performer nobody can name is a step whose removal nobody will notice until it is missing.

Where the work already runs through a system, part of this is read instead of recalled. Microsoft’s process mining and task mining in Power Automate builds its picture from “event log files that you can get from your system of recording”, and displays “maps of your processes with data and metrics to recognize performance issues”, with the gate stated on the same page: you can model the processes “for which you have data readily available”. So it covers the part of the work that already leaves a trace in a system, and the desktop half of the same capability watches “recorded user actions” for tasks that happen on somebody’s screen. What it does not cover is the part that happens in a conversation, and in a process with a manual step that part is where the difficulty sits. Read what the systems hold, then ask the people about the gaps between the records.

Two failure modes are worth naming while the step list is being written. The first is the tidy-up: a performer describes what they are supposed to do, because a step list feels like an audit. Asking about specific completed items instead of about the process in general takes much of that away, since a real item has a real history. The second is the missing branch, where a step list has one path because the person interviewed handles one kind of item. Two performers with different caseloads, walked separately, surface that.

You have enough when a person who performs the work can read the step list back to you and mark where it is wrong. If they read it and say it looks fine, either it is fine or they have not read it as a description of their day, and the difference is worth a second pass.

Question Two: Exceptions, Which Are Usually the Process

The step list describes what happens to an item when nothing unusual occurs. The work that consumes a team’s attention is generally the rest: the invoice with no purchase order, the request from a subsidiary on different terms, the approval that has to skip a level because the approver left. Teams describe these as edge cases. In a process that has been running for years they are frequently the reason the process needs experienced people to run it.

This is the question the business side owns, and the question a build inherits unanswered. A developer meeting an unlisted exception has three options: guess, stop and ask, or let the item fail. Each one costs more than the sentence that would have prevented it.

An exception register is a plain table, one row per departure from the step list, and four fields make a row usable:

  • The trigger. What is true about the item that makes it depart from the step list, stated so somebody could sort a list by it. “Supplier is foreign” is a trigger. “It is complicated” is not.
  • The handling. What is actually done today, including who does it. Where the honest answer is that different people do different things, that is the finding, and the row records both.
  • The frequency. Roughly how many, counted from the same window of completed items in the list or queue that holds them, instead of estimated in a meeting. Rough is fine. The distinction that matters is between a handful a year and several a week.
  • The consequence of getting it wrong. One line, in business terms, which is also the input to question four.

Build the register from evidence the organization already has. The rejected items, the reworked items and the ones that sat unusually long are where exceptions accumulate, and wherever the systems retain them they are already visible: the rejections in the approval history, the reopened tickets in the queue’s own export, the items your finance or service system flags for manual review. Pull that set, group it, and each group is a candidate row.

Then rule each row, because a register with no rulings is a list of problems. Each exception gets one of three futures, decided by the owner and written into the row: it is automated, meaning the automation handles it as a designed path; it is routed, meaning the automation recognizes it and hands the item to a named person; or it is excluded, meaning items of that kind do not enter the automation at all and continue to be handled as they are now. When frequency and consequence disagree about which future a row gets, consequence decides it, because a rare exception with an expensive wrong outcome is exactly the one an automation should refuse to guess at. A row left blank is a decision by default, and the default is that the automation will do something with that item and nobody chose what.

An exception register that everyone agrees is complete after one meeting is a list of the exceptions one person remembers. If the register keeps growing every time somebody new reads it, that growth is itself the finding, and worth an outside read before the build is scoped. Talk to a senior workflow automation architect

Question Three: Volume, and the Shape of It

Volume attracts confident wrong answers, because a figure is easy to offer and a count is not. “About two hundred a month” is a memory of a busy period. The number that matters has three parts: how many items arrive in a period, how those arrivals cluster inside it, and how many are in flight at once.

Read it, do not recall it. Whatever holds the items already counts them, wherever it keeps one record per item: the list or library the requests land in, the mailbox folder, the ticket queue’s own export, the finance system’s report for the period. Pull twelve periods if the system holds them, so a seasonal pattern is visible instead of inferred. When the average and the peak point at different designs, design for the peak, because complaints arrive at the peak, and a design built for the average is a design that fails at month end.

Clustering is the part that gets left out, and the part that changes the answer. Four hundred items a month arriving steadily is a different process from four hundred arriving in the two days after the invoicing run, even though both produce the same monthly figure. The second one has a queue in it, and a queue needs somebody to say what order the items come out in and what happens to the ones still waiting when the period closes. That is a business ruling, and it belongs in your specification.

How many items an automation handles at the same time is a setting, not a fact of nature. Microsoft’s Limits of automated, scheduled, and instant flows records that concurrent runs are “Unlimited for flows with Concurrency Control turned off” and “1 to 100 when Concurrency Control is turned on (defaults to 25)”, and that the control “is set in the flow’s trigger settings and is off by default”. The reason a process owner should know that a dial exists is not to set it. It is that “how many at once” has a right answer only once somebody has said whether two people working the same item at the same time is acceptable in your process, and that is your ruling to give.

Volume also feeds the decision about whether this is worth doing, which is a separate document with its own arithmetic. If the counting you do here is going to be reused to justify the spend, Building a Workflow Automation Business Case in Regulated Enterprises is where that case gets assembled.

Question Four: Failure Tolerance, and What a Wrong Outcome Costs

Buyers phrase this one as “what happens when it fails”. Underneath it are two different questions, and only one of them is yours. The design of retries, recovery and notification is engineering work, covered in a forthcoming guide, Workflow Error Handling and Recovery. What that design has to meet is a business statement, and if you do not supply it the design will be built against an assumption somebody made quietly.

There are three shapes of wrong outcome, and a process owner has a different answer for each:

  • Nothing happened. The item sat, and nobody noticed until somebody chased it. What does a day of that cost, and at what point does it stop being an inconvenience and start being a breach of something you have promised?
  • The wrong thing happened. A payment went to the wrong supplier, a record went to the wrong case, an approval was recorded against a person who never saw it. What does the wrong outcome itself cost, and separately, what do the correction and the explaining cost?
  • The right thing happened twice. Two payments, two records, two notifications to a customer. This one gets dismissed as tidy-up work, and where the receiving system is a finance or regulated one it is the costliest of the three.

For each shape, write down three things: what it costs, who has to be told, and how long the process can run in that state before it stops being a local problem. That is the tolerance statement, and it is short. Write it in business terms and have the accountable lead sign it, because that is the sentence that will decide, later and without you, how much money the build spends on protection. The tolerance statement is not final until that lead has been named.

Two consequences follow directly from that statement, and they are worth knowing before you write it. A process with a low tolerance for a wrong outcome gets checks that slow it down, while one that can absorb the occasional wrong item gets a simpler and cheaper automation. And the notification question, meaning who gets told when something is wrong, has a business answer before it has a technical one: the person who is told has to be the person who can act, and naming somebody who cannot act produces an alert everybody learns to ignore.

When two stakeholders give different tolerances for the same failure, the one accountable for the obligation behind it decides, and the obligation is read from the contract, the regulation or the service commitment itself rather than from seniority in the room; where no such obligation is written down anywhere, the tolerance is set by whoever carries the budget for the consequence.

If your problem is the opposite one, meaning automations already running that are failing without telling anybody, the sequence runs the other way and the starting point is When Power Automate Flows Fail Silently: Finding the Risk, which covers finding what you have and classifying it before anything is redesigned.

Question Five: Ownership, and Who Signs Off

Ownership sounds like the soft question, and it stalls builds that no technical problem would have stopped. It has four parts, and they come apart in real organizations: who owns the process, who approves each decision inside it, who owns the data the process writes, and who answers when the automation does something nobody wanted.

Start with the approvals, because they are the part the automation renders literally. Write, per decision point, who approves, what happens when that person is unavailable, and whether one approval is enough. The last of those is a business ruling with a direct consequence in the build. Microsoft’s Get started with approvals documentation describes behaviours that differ on exactly this point: with “Everyone must approve”, “A response is needed from each approver before the flow run is completed”; with “First to respond”, “Approval or rejection by any approver completes the request”; and with sequential approval, “Approvals are requested one at a time, in a specific order.” Today, with a human routing the paperwork, an absent approver is handled by somebody walking down the corridor. Automated, it is handled by whichever of those behaviours was chosen, and the choice is yours to make rather than to inherit.

Microsoft’s documented limits on automated, scheduled, and instant flows give the waiting question a boundary worth knowing. A run’s duration is capped at 30 days, calculated from the run’s start time and including “flows with pending steps like approvals”, and after that “any pending steps time out”. An approval nobody answers is not an approval that waits forever, so the specification needs your ruling on what should happen to an unanswered request long before that point.

Then the quieter three. Process owner is the person accountable for the outcome of the work, and if two names are offered, that is the finding. Data owner is whoever can say a record is correct, which matters because an automation writes to a system under an identity somebody has to grant it, and a process that writes to finance records with borrowed authority is a governance question waiting to be asked. The standing owner is the person who takes the call when the automation behaves in a way nobody expected after go-live, and the honest version of this answer names a person, not a team.

There is a fast way to test the approvals. Look at the approval history in whatever system currently records the decisions, and compare the names actually approving with the names your delegation of authority says should approve. When the published delegation and the approval history disagree, the disagreement is the finding and correcting it is a governance decision, and both belong in the specification before a build encodes either one.

Where approvals are the heart of the process rather than one step in it, Approval Workflow Automation Services for Enterprise Teams covers what an approval-centred build has to handle.

Ownership arguments surface late and they are rarely about the automation. If two people in your discovery both believe they own the same decision, settling that is a conversation to have before the build encodes one of them. Talk to a senior workflow automation architect

Who Should Be in the Room

Discovery gets scheduled as a workshop because a workshop is easy to book. It produces a wall of sticky notes and a photograph of it. What produces the five answers is a small number of named people in two different kinds of sitting, working from the records each question sends them to.

The people who belong there, by what they hold rather than by title:

  • The process owner, meaning whoever is accountable for the result. If nobody in the room can say this is them, that is the first finding and the rest of the discovery is provisional.
  • One or two people who actually perform the steps, including the person who handles the awkward items. That person is the one attendee nobody can substitute, because they are the one who can say what the awkward items actually were, and the step list without them is the tidy version.
  • The owner of the data the process writes to, so that the question of who vouches for a record is settled by the person who answers for it rather than assumed.
  • The approver, or a delegate who can speak for the approval rules, since the sign-off rulings are decisions and not observations.
  • One person from the delivery side, there to hear the answers rather than to shape them. Their value is asking what a step actually means when it is described loosely, and their risk is turning the sitting into a design session, which sends everybody home with a solution and no specification.

The two kinds of sitting do different jobs and mixing them is what makes discovery expensive. The step list comes out of separate, quiet conversations with performers, one at a time, and so does the exception register, because a performer describing their real workarounds in front of their manager describes the official version. The rulings, meaning the exception dispositions, the tolerance statement and the ownership names, come out of one joint sitting with the owner and the approver, after the draft artifacts exist, where the job is to disagree with something concrete.

This is a different room from the one that picks which process to automate. That room is scoring candidates against each other and wants people who can compare unrelated processes. This room has one process in it and wants the people who touch it.

When the Answers Say Not Yet

Discovery that concludes “do not build this now” has done its job and saved the money, and it is worth saying that out loud before the meeting that has to say it. Three findings should stop a build rather than shape it.

The exception register keeps growing. Every new reader adds rows, and the frequencies in the register say the exceptions are a large share of the counted volume. That is a process which has not been agreed, and automating it encodes one person’s version. The work in front of you is to settle the process, and that work is not an IT project.

No name survives the ownership question. Where the process owner is a committee and the approval history contradicts the delegation of authority, an automation makes the ambiguity permanent and much harder to see, because the routing rules are then buried in a build instead of visible in an inbox.

The process is about to change for a reason that has nothing to do with automation. A system replacement, a reorganization, or a regulatory change already on a published plan will rewrite the step list, and a specification written against the current version is a specification with a known expiry.

There is a fourth case, which is the happier one. Discovery sometimes shows the process is fine and the problem is one step, in which case the answer is a much smaller build than the one that was proposed, and the discovery has just removed the bulk of its cost.

What is left after those four is the organization this was written for: one process that has been chosen, and five answers written down. A step list somebody who does the work has argued with, an exception register with rulings in it, a counted volume, a signed tolerance statement, and named people. That is a specification a build can be quoted against, and the ambiguity it removes is what would otherwise arrive as change requests in its first month. i3solutions has been a Microsoft partner since 1997. The useful next step is to write the five answers down for the process you have already picked and see which ones are still blank. Talk to a senior workflow automation architect

Frequently Asked Questions

What should a process owner define before automating a workflow?

Five things, and every one of them is a business answer rather than a technical one: process truth, exceptions, volume, failure tolerance, and ownership. Process truth is the step list, numbered, covering the work as people actually perform it and naming who or what does each task, drawn from walking finished items with the staff who handled them. Exceptions is the exception register, one row for every departure from that list, carrying the trigger, the handling today, the frequency and the consequence of getting it wrong. Volume is a counted arrival figure with its peak, its clustering and the number in flight together, taken from whatever system already holds the items. Failure tolerance is a short tolerance statement, signed by the accountable lead, covering what a wrong outcome costs, who must hear about it, and how long that state is tolerable before it stops being a delay. Ownership is named individuals: the person accountable for the work itself, the approver at each decision, the owner of the data the automation writes into, and the person who answers after go-live.

What discovery questions prevent failed automation projects?

The ones whose answers only the business holds, asked before the build starts. The specification decides what the build does with the work, and the defects that do the damage are a step list describing the documented process instead of the performed one, an exception register assembled from memory in a single meeting, a volume figure nobody counted, nothing written down about the cost of a wrong outcome, and an ownership question left with two candidate names in it. Every one of those produces the same symptom at the demo: an automation that handles the ordinary case and argues with the real one. The counter-questions are direct. Can the person who performs this work read the step list, then tell me where it is wrong? Which departures from it are real, how many of each, and what do we do about them? How many arrive in a period, and what is the peak? What does a wrong outcome cost, and who must hear about it? Who approves, and what happens when they are unavailable?

How do we document exceptions before automating?

Build a register from evidence instead of from recall. Pull a bounded window of items the process has already finished, along with the rejected items, the reworked items and the ones that sat unusually long, all of which are normally already visible in the approval history or the queue’s own export. Group them, and each group becomes a row. Give every row four fields: the trigger, stated so somebody could sort a list by it; the handling as it is done today, including who does it; the rough frequency, counted from that same window; and one line on what getting the item wrong costs the business. Then rule each row into one of three futures: automated, meaning the build handles it as a designed path; routed, meaning the build recognizes it and hands the item to a named person; or excluded, meaning items of that kind stay outside the automation. Where frequency and consequence point at different rulings, consequence decides, because a rare exception with an expensive wrong outcome is the one an automation should refuse to guess at. A row with no ruling is still a decision, taken by whoever writes the build.

Who should be in the room for automation discovery?

A small group, defined by what each person holds. The process owner, meaning whoever is accountable for the outcome. One or two people who perform the steps, including whoever handles the awkward items, since only that person can say what the difficult items really were. Whoever owns the data the automation writes into, so the question of who can say a record is correct gets a real answer. The approver or a delegate who can speak for the approval rules. And one person from the delivery side, present to hear the answers rather than to design a solution. Run it as two kinds of sitting: quiet one-to-one conversations with performers to build the step list, then the exception register against the finished items the systems already hold, because people describe their real workarounds differently in front of their manager, and after that one joint sitting with the owner and approver to take the rulings once draft artifacts exist. This is a different group from the one that scores which process to automate first, which needs people who can compare unrelated processes.

How do we get a real volume figure for a process?

Read it out of the system that already holds the items rather than asking people to estimate it. Whatever the items already sit in has counted them, so long as it holds one record per item: the request list or library, the shared mailbox, the ticket queue’s export, the period report out of the finance system. Pull as many periods as the system retains, ideally twelve, so a seasonal pattern shows up instead of being inferred from a busy month. Then record three numbers, not one: how many arrive per period, how those arrivals cluster inside it, and how many sit in flight together. Clustering is what gets dropped, and what changes the design, because four hundred items arriving steadily and four hundred arriving in two days after an invoicing run are different processes with the same monthly total. Where the average and the peak point at different designs, design for the peak, since the complaints arrive at the peak.

What does failure tolerance mean, and who decides it?

Failure tolerance is a business statement of what a wrong or missing outcome costs and how long the work can run in that state, and the process owner decides it with the accountable business lead. It is not a technical decision, and it is the input the technical decisions are made from: how much a build spends on checking, verification and alerting follows from it. Write it for three shapes of wrong outcome separately, because they cost different amounts. Nothing happened, so the item sat until somebody chased it. The wrong thing happened, so a payment, a record or an approval landed against the wrong party, which carries the cost of the error plus the cost of correcting and explaining it. And the right thing happened twice, which gets dismissed as tidy-up and is the costliest of the three once finance or regulated systems are on the receiving end. For each, record what it costs, who must be told, and how long is tolerable, then have the accountable lead sign it. Where two stakeholders give different tolerances for the same failure, the one accountable for the underlying contractual or regulatory obligation decides, read from the obligation itself.

Related Reading

About the Author

Michael Branson co-founded i3solutions and brings executive, operational, and technical perspective to organizations running complex, secure, and mission-critical Microsoft estates. He works with business and technology leaders on the decisions taken before a build starts, where the value of an automation is either protected or lost.

CONTACT US

Leave a Comment

Your feedback is valuable for us. Your email will not be published.

Please wait...