CMMC ANSWER ENGINE

Can I use AI-generated or auto-generated reports as CMMC evidence?

JWJil Wright, Lead CMMC Certified Assessor · Last verified 2026-05-14
Author's note: The DoW, NIST, and the Cyber AB have not published specific official guidance on AI-generated evidence. The position below is my personal opinion as a CMMC Lead Assessor & SME. Other Lead Assessors may take a different approach.

This question is increasingly important as contractors lean on both traditional automation and AI tools to produce evidence. The two categories look similar on the surface and are very different in how an assessor will treat them.

What I will do as your Lead Assessor when I see AI-looking evidence

This is the practical reality, not a hypothetical risk. If a piece of evidence looks like it was generated by an AI tool — an LLM-summarized configuration write-up, an AI-drafted policy, a paraphrase of system state — I will dig into it. Specifically, I will:

  • Verify the artifact reflects your actual environment. AI tools produce plausible-sounding content that may not match what your system is actually configured to do. I will compare the AI-generated content to live system state.
  • Question the responsible staff to confirm they understand what the artifact says. If the staff member who owns the practice cannot walk me through the AI-generated artifact and explain it in their own words, the artifact does not represent operational reality.
  • Confirm the OSA actually does what the artifact describes. An AI-drafted policy or procedure that the organization does not actually follow is not evidence the practice is implemented — it's just a document.
  • Probe the chain of custody. Where did the AI input come from? What was the prompt? What was edited after generation? If the answers are vague, the credibility of the artifact takes the hit.

None of this is unique to AI — assessors verify operational reality against documentation for any artifact. But AI-generated content amplifies the gap between what's claimed and what's operating, which is why I scrutinize it harder.

Category 1: Traditional automation (acceptable when handled correctly)

System-generated reports produced directly by the asset being assessed have long been accepted at CMMC assessments — often preferable to manual screenshots because they preserve more detail and remove the risk of selective capture. Examples:

  • Scripts or APIs (PowerShell, Bash, REST API queries) that export an inventory directly from the source system — for example, exporting an identity provider's user list with MFA status, or pulling device compliance posture to CSV.
  • EDR / Vulnerability scanner / SIEM exports — reports generated by the security tool with its own timestamps and identifiers.
  • Compliance dashboards — provided the underlying data and the dashboard's source-system connection are auditable.
  • Scheduled automated reports — weekly or monthly exports that demonstrate operating cadence as well as content.

Category 2: AI-generated content (treat with caution)

The Lead Assessor advice from Wrightbrained Security: do not use AI-generated evidence unless you understand it well enough to know whether it is good, and do not assume an assessor will accept it. AI tools can produce convincing-looking artifacts that don't accurately represent the system state, that omit context, or that paraphrase in ways that obscure what actually happened. The contractor — not the AI — will have to explain and defend the evidence under assessor questioning.

Common AI-generated content that does not belong in an evidence pack:

  • LLM summaries of system state. An LLM that says "Based on your description, here's what your MFA configuration looks like" is not evidence — it's a paraphrase.
  • AI-drafted policies that have not been adopted. An AI-drafted policy is fine as a starting point, but until it is reviewed, approved, and adopted by the OSA, it is not evidence the practice is implemented.
  • Evidence reconstructed after the fact. If you didn't run the report at the time, you didn't capture that point-in-time state — you can't AI-generate it later.
  • Auto-summarized log excerpts. If the assessor wants log evidence, the assessor wants the log — not an AI's interpretation of what the log shows.

What automation cannot substitute for

  • The live demonstration that NIST 800-171A specifies for the Test method. A scripted output proves what the system was at the time of the script; it doesn't replace exercising the control.
  • Configuration screenshots when the screenshot adds context. A script may export "MFA is enabled" as a boolean; a screenshot of the conditional access policy shows the rule structure.

How to defensibly use traditional automation

For the kinds of automation that are clearly acceptable (Category 1):

  • Run scripts on a fixed cadence (monthly or weekly) tied to the SSP-described control cadence.
  • Have the script timestamp itself in the output filename and inside the file.
  • Have the script identify the source system (tenant ID, hostname, environment).
  • Store outputs in the evidence folder structure under the correct domain.
  • Document the script itself as part of the evidence — which script, who maintains it, where it lives.
  • Verify the script's output before relying on it — a script that's been broken for two months produces zero evidence even if it runs daily.

The defense test (applies to all evidence)

Before submitting any piece of evidence — manual, automated, or AI-touched — ask: can the staff member responsible for the practice walk an assessor through this artifact, explain how it was produced, and answer questions about what it shows? If the answer is no, the artifact is not ready to be evidence. AI-generated artifacts fail this test more often than traditional automation, which is why the more conservative position is the right one.

Common errors

  • Submitting AI summaries as evidence. If the assessor asks for evidence and gets a paragraph that starts with "This system uses...", that's not evidence.
  • Trusting scripts without verifying their behavior. A script that's been broken for two months produces zero evidence even if it runs daily.
  • Modified script outputs. Editing a CSV export to remove a non-compliant row destroys the chain of custody — the same problem as editing a screenshot. Preserve raw outputs and address findings through remediation, not through editing the artifact.
  • Treating AI-drafted policies as adopted policies. A draft is a draft. Adoption requires review, approval, and the artifacts of that process.
  • Submitting evidence the responsible staff cannot defend. If the operator cannot explain the artifact under assessor questioning, the evidence has a credibility problem regardless of how it was produced.

Sources

  • Wrightbrained Security — CMMC Assessment Preparation Master Guide — Section 2.3: Evidence Quality Rules — Evidence integrity and chain of custody — link
  • NIST SP 800-171A — Test method; evidence integrity expectations — link
Rulebook version: NIST SP 800-171A; Wrightbrained APREP
// READY WHEN YOU ARE

Is your security posture keeping you up at night?

Thirty minutes, no slide deck. Tell us what you're up against and we'll tell you honestly whether we can help.