3:00AM SOC2 is useless so now what?
It’s 3:00 AM and a client CIO is absolutely raging mad. I can almost feel little bits of spit hitting me through the monitor as she screams into her laptop (they now require face-to-face communications). Vendors are submitting vibe-coded reports that are so uniformly bad that you can easily identify Claude vs. ChatGPT down to the model version.
I know why I am on this call. Why are there 8 other people here?

I was on a long engagement so the Delve scandal (More details here) did not really affect me. I read about it, chuckled and filed it away as something less than critical. I also only missed the ISO 27001 subreddit drama because I wasn’t paying attention to that niche subreddit. I am paying attention now.
The CIOs and CISOs aren’t asking about their SOC 2 reports, they are asking about the SOC 2 reports that were required of their many, many vendors. Vendors across many different industries and domains.
Who is responsible?
Delve wasn’t the Independent Auditor. That assurance depended on the CPA firm independently obtaining and testing sufficient evidence.
There is a legitimate concern that an approval based primarily on an unreliable SOC 2 report could be difficult to defend following an incident. The AICPA expects user entities to evaluate matters disclosed in the report, including its scope, carve-outs and applicability to the service used. However, that responsibility does not relieve the independent auditor of its responsibility to obtain sufficient evidence, conduct the examination independently and support its opinion.
Am I responsible?
It is not my role to independently re-audit the auditor or reproduce their testing. My engagement scope is limited to validating the report, in most cases as a second opinion to the in-house process. About 1/3 of the time I am not offered additional evidence beyond the SOC 2 report—things like network or workflow diagrams, threat models, etc.
My process has evolved over the years. Right now I start with a sanity check to verify that the report is complete and makes sense within the context of what is being tested (I just had a report for a SaaS app submitted with an in-house asset inventory).
I check the auditor’s opinion and make sure the report period is okay (a recurring problem that’s getting worse). I make sure the description of the system is comprehensible (to a 15 year old). I look for anything that indicates they didn’t take their CIA responsibility seriously.
I then mentally prepare myself for the arduous task of reading the control descriptions and testing results.
This is where the Delve allegations create serious damage to the assurance process.
When I evaluate a SOC 2 report, I rely on the auditor’s representation that the underlying control actually exists and that the described testing was genuinely performed. I am not merely asking whether a Conditional Access policy appears in a screenshot. I am relying on the auditor to determine whether the policy was implemented correctly, applied to the relevant population, operated throughout the reporting period, and was effective for its intended purpose.
That trust is foundational.
If a meeting, access review, training event, or test result was fabricated and then accepted as audit evidence, the problem is not simply that one control failed. The problem is that the integrity of the entire examination becomes questionable.
Is AI responsible?
Not really. At least not in this instance. I have worked with AI systems long enough to know that even the strongest frontier models can confidently invent employees, meetings, teams, departments, policies, and entire organizational structures when prompted—sometimes correctly and sometimes incorrectly.
The dangerous part is not that AI can generate convincing compliance documentation. The dangerous part is when generated documentation is mistaken for evidence that an activity actually occurred.
So now what
There is no consensus forming around a possible solution. Even formulating the exact problem is becoming difficult because SOC 2 is such a comfort blanket. Like the Gartner Magic Quadrant.
Budgets are tight; looking back over the existing reports (around 60 over the past 24 months in my case) is not palatable. “Feed it to AI to pull out and verify the reputations of all the auditors” was suggested multiple times. I am not listening to this at 3:00 AM.