The reproducibility gap in Copilot: why "it worked last time" isn't an audit answer
Reproducible AI IT automation is now a buying filter. Why suggest-and-forget copilots can't show how they'd do it again, and an agentic IT engineer can.
Ask a copilot vendor how their product would handle last Tuesday's access change again, and you'll usually get one of two answers: a chat transcript, or "it worked last time." Neither is an audit answer. A transcript shows what the AI suggested. It doesn't show what a human actually did with the suggestion, which policy allowed it, or what the tenant looked like afterward.
That matters now because "show me how you'd do it again" has moved from governance theory to a line item on procurement checklists for any vendor whose AI touches production systems. Our last post argued that reproducibility is the new procurement question. This one goes a step further: the copilot category fails that question by design, and an agentic IT engineer passes it by design. That turns reproducibility from a principle you nod at into a buying filter you can actually use.
What a reproducibility question actually asks
When an auditor, a security lead, or a procurement reviewer asks whether an AI-driven IT change is reproducible, they're asking for five things tied together in a single record:
- What was found. The state of the environment when the request came in: group memberships, license status, policy conflicts, whatever was relevant.
- What was decided. The plan, in order, and why that plan over the alternatives.
- What permitted it. The specific policy that authorized each action, under whose identity.
- What was executed. The actual operations against the backend, not a description of them.
- What came back. The system's response, success or failure, with timestamps.
If all five live in one record, you can explain the change, defend it, and run the same request through the same path again. If any link is missing, you have a story, not evidence. That's the core of reproducible AI IT automation: the record is produced by the thing that did the work, at the moment it did the work.
Why copilots fail the filter by design
A copilot's job is to suggest. A human's job is to act. That split is the product, and it's also the problem.
Walk through a typical case. A user can't open a SharePoint site. An admin asks a copilot what's wrong, and the copilot suggests adding the user to a security group. The admin opens the Entra admin center, adds the user to a group (maybe the suggested one, maybe a similar one with broader access because it was quicker), and moves on.
Now look at what exists afterward. The copilot has a chat log of a suggestion. The Entra audit log has a group membership change under the admin's name. Nothing connects them. The copilot never learns whether its suggestion was followed. The audit log never learns why the change was made. There's no record of what the copilot checked before suggesting, and no policy that bound the action, because the copilot didn't take one.
Six months later, someone asks why that user has access to a finance site. The honest answer is "an admin followed a suggestion, probably." Ask the copilot the same question today and it may suggest something different, because it reasons from the prompt in front of it, not from what happened last time. That's the suggest-and-forget pattern, and no amount of model improvement fixes it. The gap isn't in the intelligence. It's in the architecture: the reasoning and the action happen in two different systems, and only one of them writes anything down.
This is the same dividing line we drew in the three questions that separate agentic IT from chatbot copilots. Execution was question two. Reproducibility is what execution buys you after the fact.
How an agentic IT engineer closes the gap
Dex is an autonomous IT engineer. It doesn't hand the work back to a human. It investigates the environment, plans the sequence of actions, and executes the change itself across Microsoft 365, Google Workspace, Okta, and other SaaS tools. Because one system does all three steps, one system can record all three steps.
Take the same SharePoint request. Dex checks the user's group memberships, the site's permission inheritance, Conditional Access, and license status. It finds the actual cause, plans the fix, and checks each action against its policy engine before running it. Every action must match an explicit, structured policy enforced in code, below the model. No matching policy means no action. Then it executes, using delegated permissions, and records the result.
Each action and each refusal writes an entry to Dex's Activity Log and to the native Microsoft 365 logs: who asked, which policy authorized it, what ran, and what the backend returned. Two independent records, one from Dex and one from Microsoft, that you can reconcile against each other. And because Dex keeps persistent memory of environment quirks and working solutions, plus reusable skills, the next similar request starts from the same approach instead of a blank prompt.
The language model can still phrase things differently from one run to the next. The part that has to be reproducible isn't the wording. It's the path: the same findings feed the same policy checks, and the same policy either permits the same action or blocks it. That guarantee lives in the policy engine, not in the model's mood. For a walkthrough of the investigate, plan, and execute loop, see how Dex works.
Reproducibility matters more as the work gets harder
It's tempting to treat reproducibility as a box to tick for password resets. That's backwards. A password reset is easy to explain from memory. The hard cases are the ones you'll be asked about.
Dex resolves L1 through L3. That covers routine access and MFA work, and it also covers the Tier 2 and Tier 3 troubleshooting and configuration that used to need a senior engineer: untangling a Conditional Access conflict, fixing mailbox delegation across a department, tracing why a device keeps falling out of compliance. Those changes touch more systems, carry more risk, and are far harder to reconstruct after the fact. A decision record is worth the most exactly where a human would struggle to remember what they did and why.
When something genuinely needs human judgment, Dex escalates with the full investigation attached. The engineer who picks it up starts from the record, not from a blank ticket, and the record stays intact either way.
Turning reproducibility into a buying filter
Governance principles don't win evaluations. Concrete tests do. Here's how to make reproducibility one, in the room, before the contract:
Ask for a past change, not a demo. Pick an action the tool took in a pilot tenant last week. Ask for the full record: findings, plan, policy, action, backend response. A copilot vendor will show you chat history. That's a failed test.
Ask them to run it again. Submit the same request and compare the path. If the outcome differs, ask why, and expect the answer to point to a change in the environment or the policy, not "the model decided differently."
Ask for a refusal. Have the tool attempt something no policy permits, then show the record of the refusal. A system that can't log what it declined can't prove what it allowed.
Ask who did it. Every action should resolve to a requester and a permission scope. If the answer is a shared service account with broad API access, the record can't tell you whose authority the change ran under.
Score vendors pass or fail on these four. The copilot category fails the first test and never reaches the others. Not because the products are weak, but because they were never built to own the action.
What this means for ITSM modernization
Most ITSM modernization plans layer AI onto the existing ticket flow: a copilot to draft replies, summarize tickets, and suggest fixes. That speeds up the human who still does the work. It doesn't change who does the work, and it doesn't change what gets recorded. The ticket says "resolved," the admin log shows a change, and the reasoning in between stays undocumented.
Modernizing with an agentic IT engineer changes both. The work gets done before the ticket exists, and the record of how it got done is a byproduct of doing it. That's a stronger answer to your auditors than any ticket system has given you, and it's the answer procurement is now asking for.
"It worked last time" was always a weak answer. Now it's a disqualifying one. Ask every AI vendor on your shortlist to show you how they'd do it again, and buy from the ones who can.
Frequently asked
- What is reproducible AI IT automation?
- Reproducible AI IT automation means that for any change an AI system made in your environment, you can show what it found, what it decided, which policy allowed it, what it executed, and what the backend returned, and you can run the same request again through the same path. The record is produced by the system that did the work, not reconstructed afterward from memory or chat history. If a vendor can't produce that record for a change made last month, the automation isn't reproducible.
- Why can't a copilot answer the reproducibility question?
- A copilot suggests and a human acts. The suggestion lives in a chat window, the action happens in an admin console, and nothing ties the two together. The copilot doesn't know whether its suggestion was followed, modified, or ignored, and the admin log doesn't know why the change was made. That split is the design of the category, not a missing feature, so the decision record a reproducibility question asks for doesn't exist anywhere.
- How does Dex make IT actions reproducible?
- Dex investigates, plans, and executes the change itself, so the whole path runs through one system. Every action must match an explicit policy in a code-level policy engine, and every action (and every refusal) is written to Dex's Activity Log and to the native Microsoft 365 logs. Each entry captures who asked, which policy authorized it, what was executed, and the outcome. Persistent memory and saved skills mean the next similar request starts from the same approach instead of from scratch.
- What should procurement ask an AI IT vendor about reproducibility?
- Ask for three things. First, pick a change the tool made last month and show the full decision record: findings, plan, policy, action, backend response. Second, show the same request run again and explain any difference in the outcome. Third, show a request the tool refused and the record of why. A vendor that answers with chat transcripts or 'it worked last time' has failed the filter.
- Does reproducibility only matter for simple L1 tickets?
- No, it matters more as the work gets harder. Dex resolves L1 through L3: password resets and access requests, and also the Tier 2 and Tier 3 troubleshooting and configuration work that used to need a senior engineer. A password reset is easy to explain after the fact. A Conditional Access change or a mailbox permission fix across a department is not, which is exactly where a decision record earns its keep.