You can't audit what you can't reproduce: why "show me how you'd do it again" is the new procurement question
The agentic AI governance audit trail questions for your RFP: reproducibility, audit trail, data flows, and human oversight, with verifiable answers.
Enterprise procurement teams have stopped accepting "it's proprietary, trust us" from AI vendors. ISO/IEC 42001 and the agentic-AI checklists security teams have built on top of it have pushed reproducibility and audit trails onto RFPs as named requirements, and buyers now fact-check every answer. The sharpest version of the new question is simple: show me how you'd do it again. If a vendor can't replay why its agent took an action, nobody can audit that action. This post gives IT leaders the four on-record questions to ask any agentic vendor about reproducibility, audit trail, data flows, and human oversight, and shows how Dex answers each one with claims you can verify yourself.
Why reproducibility became an RFP line item
Two years ago, the AI section of an IT procurement questionnaire asked whether the vendor used customer data for training. Today it asks how the vendor governs an agent that takes action on its own. ISO/IEC 42001, the AI management system standard, gave risk and compliance teams a shared vocabulary for that: documented controls, traceability, and evidence that the system behaves as intended. Procurement turned that vocabulary into checklist items, and "reproducibility" and "audit trail" are now the two that show up most.
The reason is practical. An auditor doesn't review intentions; they review evidence. If an autonomous agent removed a user from a security group last Tuesday, the auditor wants to know why it was allowed, and whether the same request today would produce the same decision. A system that can't answer the second question can't really answer the first.
We covered the logging half of this in why the audit trail is the product. Reproducibility is the newer half, and it's where vague vendor answers fall apart fastest.
Four agentic AI governance audit trail questions to put on the record
Ask these in writing, and ask for evidence alongside each answer. Vague answers are easy to spot once you see them next to specific ones.
Question 1: Reproducibility. "Can you replay a decision?"
The vague answer: "The model is usually consistent."
That answer concedes the point. Large language models don't produce identical output on every run, and any vendor who claims their agent's reasoning is fully deterministic is overselling. So the honest question isn't "will the AI write the same sentence twice?" It's "is the decision about what the agent was allowed to do reproducible?"
What a credible answer looks like: the permission decision lives in deterministic code, outside the model. In Dex, every action must match an explicit, structured policy before it executes. The Policy Engine evaluates six layers (Global, Tenant, Target Rules, Department, Action, Runtime), and the rule is absolute: no matching policy, no action. Because that check runs in the execution layer rather than in a prompt, the same request evaluated against the same policy set gets the same allow-or-block decision every time. Prompt injection can't talk its way past it, because there's no prompt to talk to.
That's the part an auditor can replay. The model may phrase its investigation differently on a second run. The policy decision won't change.
What to ask for as evidence: the policy that authorized a specific past action, and a demonstration that an out-of-policy request is blocked.
Question 2: Audit trail. "Who did what, and why?"
The vague answer: "We log everything."
Logging everything is not the same as producing an audit record. The questions that matter are what each entry contains, who writes it, and whether you can check it against a source the vendor doesn't control.
What a credible answer looks like: a per-action record written by the system that performed the action. Every Dex action, and every refusal, writes an entry capturing who requested it, which policy authorized it, what was executed, and the outcome. That entry lands in two places: Dex's own Activity Log and your native Microsoft 365 logs (Entra ID, Exchange Online, SharePoint). The second source matters most. Microsoft generates those logs independently of Dex, so you can reconcile one against the other instead of taking the vendor's word for it.
Actions also run under delegated permissions, as the requesting user or admin with their existing access, rather than through a shared, broadly scoped API key. Your identity model stays the authority on who could do what, which keeps the agentic AI governance audit trail tied to identities your auditors already review.
What to ask for as evidence: one real action's record, end to end, next to the matching Entra ID audit log entry.
Question 3: Data flows. "Where does tenant data go?"
The vague answer: "Enterprise-grade, fully compliant."
"Fully compliant" with what, attested by whom, as of when? This is the question where unverifiable claims do the most damage, because a buyer who catches one overstated certification will discount everything else in the proposal.
What a credible answer looks like: specific mechanisms plus an honest certification status. Here is Dex's, stated exactly:
- Hosting and encryption: hosted on AWS, encrypted in transit with TLS 1.2+ and at rest with AES-256.
- Isolation: per-org isolated databases and encryption keys, so one tenant's data and audit trail never share storage with another's.
- Retention: zero data retention for Microsoft 365 data. Dex reads only what a task needs, then discards it.
- Certification status: Dex's own SOC 2 Type II attestation is in process, not complete. Dex is built by the SysAid team, whose platform is ISO 27001, ISO 27017, and ISO 27018 certified and SOC 2 Type II compliant with annual third-party audits. Dex is built on that foundation to the same standards.
Notice what's not on that list: a claim that Dex holds certifications it doesn't yet hold, and a claim of ISO 42001 certification. A vendor that is precise about what's still in process is easier to believe about everything else. The full detail is on our security and compliance page.
What to ask for as evidence: current certification reports or a written status for each one, and a data flow diagram showing where tenant data is read, processed, and discarded.
Question 4: Human oversight. "Who signs off on risky actions?"
The vague answer: "Humans stay in the loop."
Which humans, for which actions, and can the agent skip them? "Human in the loop" without a named control is a slogan.
What a credible answer looks like: a specific gate the AI cannot override. In Dex, sensitive or irreversible actions require explicit human approval through the Approval Engine before they run, every time. In Dex Pro, the admin console shows every action before it executes, with an "Approve Always" option an admin can choose for repetitive bulk operations they've already reviewed. Two guardrails sit below every policy as hard stops: Dex never grants admin roles and never bypasses MFA. And when a request matches no policy at all, nothing executes; it escalates to a human with full context attached.
That oversight model holds across the whole range of work Dex does. Dex resolves L1 through L3 autonomously, from password resets and access requests to the Tier 2 and Tier 3 troubleshooting and configuration work that used to wait for a senior engineer. Deeper work doesn't get a looser control path. It runs through the same policy check, approval gates, and audit record.
What to ask for as evidence: the list of action types that require approval, and a live demonstration of an approval request and an out-of-policy refusal.
How to score the answers
Put each vendor's responses side by side and grade every answer on one test: can I verify this without trusting the vendor?
- Strong: names a specific mechanism, points to evidence you can inspect, and is precise about what's still in process.
- Weak: describes intent ("we prioritize safety"), cites unnamed "industry-leading" standards, or blurs the line between a parent company's certifications and the product's own.
- Disqualifying: guardrails that live in the prompt, logs written by the model narrating itself, or a certification claim that doesn't survive a request for the report.
Reproducibility deserves the most weight, because the other three depend on it. An audit trail records what happened. Reproducibility proves the system would make the same call again under the same rules. Together they turn "trust us" into evidence.
What to do next
If you're drafting an agentic AI section for an RFP this quarter, lift the four questions above verbatim and require evidence for each. Then run the same exercise on Dex: ask us to replay a decision, pull an action's audit record next to your Entra ID log, and walk through the approval gates. Our security and compliance page lists exactly what's certified and what's in process, and the audit trail deep dive covers the policy-to-action-to-log chain in detail.
The vendors worth buying from are the ones who welcome the question "show me how you'd do it again." The answer is either evidence or it isn't.
Frequently asked
- What is an agentic AI governance audit trail?
- An agentic AI governance audit trail is the per-action record an autonomous AI system produces when it changes something in your environment: who requested the action, which policy authorized it, what was executed against the backend, and the result. To hold up in an audit, it has to be written by the execution layer that performed the action (not narrated by the model afterward), it has to cover refusals as well as completed actions, and it should reconcile against an independent source such as your native Microsoft 365 logs.
- Can an AI agent's decisions actually be reproduced if the model isn't deterministic?
- The model's wording won't be identical from run to run, and any vendor claiming otherwise is overselling. What can be reproducible is the decision boundary. If every action has to match an explicit policy enforced in code, the same request evaluated against the same policy produces the same allow-or-block decision every time. That is the part an auditor needs to replay. Dex enforces this with a deterministic, code-level Policy Engine: no matching policy, no action.
- Is Dex SOC 2 or ISO 42001 certified?
- No, and we don't claim to be. Dex's own SOC 2 Type II attestation is in process. Dex is built by the SysAid team, whose platform is ISO 27001, ISO 27017, and ISO 27018 certified and SOC 2 Type II compliant with annual third-party audits, and Dex is built on that foundation to the same standards. Dex does not hold ISO 42001 certification. For specific compliance requirements, contact support@dex365.ai.
- What should an RFP ask an agentic AI vendor about human oversight?
- Ask which actions require explicit human approval, whether the AI can override that requirement, and what happens when a request matches no policy. A credible answer names the specific gate and shows it running in code. In Dex, sensitive or irreversible actions require explicit approval through the Approval Engine every time, the AI cannot override the Policy Engine, and requests with no matching policy escalate to a human with full context attached.
- Does Dex only handle simple L1 tickets?
- No. Dex resolves L1 through L3 autonomously: routine work like password resets, MFA recovery, and access requests, plus the Tier 2 and Tier 3 troubleshooting and configuration work that used to sit with senior engineers. Every one of those actions runs through the same policy check, approval gates, and audit trail. Only genuine architectural or judgment calls escalate to a human.