Back to Insights
FRAMEWORKAUGUST 27, 202614 min read

The First 48 Hours: What a Defensible AI Incident Response Looks Like

Sigilith Research

Institutional AI governance & accountability

It arrives as a support escalation, a journalist's email, a regulator's letter, or a line on an internal dashboard that someone finally reads: your AI system has produced an outcome that is harmful, wrong, or unlawful. A claim denied to someone entitled to it. A screening system that has been quietly excluding a protected group. A model output that caused a clinical, financial, or physical harm. The specifics vary; the next 48 hours do not.

Those hours contain five workstreams, and they run concurrently, not in sequence: preserve everything before anyone touches anything, reconstruct what the system actually did, determine what you are legally required to report and by when, communicate without asserting facts you have not established, and remediate in a way you can later prove. Each one is examinable afterward, by regulators, by opposing counsel, and by your own board.

This framework walks through all five. But it carries a premise worth stating at the top, because it is the part most incident-response templates omit: every one of these workstreams consumes records. Records of inputs, versions, policies, reviewers, and time. If those records exist when the incident begins, the response is fast, accurate, and defensible. If they do not, no quality of crisis management in hour three can compensate, because the missing input is a record that had to be written months earlier. Readiness for an AI incident is not primarily a process property. It is a recordkeeping property. We will return to that at the end, with the evidence accumulated along the way.

1. What counts as an AI incident, and why yours is not the first

The EU AI Act supplies the definition most regimes will converge toward. Article 3(49) defines a serious incident as "an incident or malfunctioning of an AI system that directly or indirectly leads to" any of four outcomes: the death of a person or serious harm to a person's health; a serious and irreversible disruption of the management or operation of critical infrastructure; the infringement of obligations under Union law intended to protect fundamental rights; or serious harm to property or the environment.

Read point (c) twice. A fundamental-rights infringement (a discriminatory screening outcome, an unlawful automated denial) is a serious incident in the same statutory list as a death. Most organizations' mental model of an AI incident is an outage or a jailbreak. The regulatory model includes the quiet failure that harmed no server and one person.

Two facts about AI incidents should shape the response before it begins. First, they recur. The AI Incident Database, maintained by the Responsible AI Collaborative and modeled explicitly on the incident databases of aviation and computer security, is "dedicated to indexing the collective history of harms or near harms realized in the real world by the deployment of artificial intelligence systems"; its founding paper's title states the purpose plainly: Preventing Repeated Real World AI Failures by Cataloging Incidents. Whatever failed for you has, in family terms, almost certainly failed before, which means "we could not have foreseen this" is a harder sentence to say than it used to be. Second, they are accelerating: Stanford's AI Index reports 362 incidents recorded in the database in 2025, up from 233 in 2024.

The major governance frameworks already assume you have planned for this. NIST's AI Risk Management Framework places incident handling inside the MANAGE function: post-deployment monitoring plans are expected to include "incident response, recovery, and change management" (MANAGE 4.1), and "processes for tracking, responding to, and recovering from incidents and errors are followed and documented" (MANAGE 4.3). An organization improvising its response in hour one is already outside the framework it probably claims alignment with.

2. Hour zero: preserve everything, before anyone fixes anything

The engineering instinct after a failure is to fix it. The defensible response inverts the order: freeze first, because the fix destroys evidence, and destroyed evidence is a second incident layered on the first.

The legal mechanism is the litigation hold, and its trigger is earlier than most teams assume. The rule, stated in Zubulake v. UBS Warburg and applied ever since: once a party reasonably anticipates litigation, it must suspend its routine document retention and destruction policy and put a hold in place to preserve relevant records. A serious AI incident (a harmed person, a regulator's letter, a pattern of wrongful denials) will usually qualify as "reasonably anticipates" on day one. The consequences of failing sit in Federal Rule of Civil Procedure 37(e): where electronically stored information that should have been preserved is lost and cannot be restored, courts may impose curative measures, and on a finding that a party "acted with the intent to deprive another party of the information's use in the litigation," may presume the lost information was unfavorable, instruct the jury accordingly, or enter default judgment. Our analysis of Mobley v. Workday documents what the downstream looks like when contemporaneous records were never created at all: dozens of third-party subpoenas, multi-year reconstruction, and the pointed observation that evidence cannot be manufactured retroactively, and that the attempt is itself discoverable.

What does a hold cover for an AI system? More than the incident ticket. The decision outputs and the inputs as the system saw them; model artifacts and version identifiers; prompts, configurations, thresholds, and the policy revisions in force; evaluation and testing results; deployment and change history; the human review trail; and the surrounding communications. Note what this list collides with: operational log rotation, which deletes on schedule, indifferent to what it deletes, and ephemeral infrastructure, which destroys state by design. The hold has to reach those systems in hours, not days.

The EU AI Act adds a preservation rule of its own that engineering teams routinely trip over: under Article 73, a provider investigating a serious incident may not conduct an investigation that involves altering the AI system in a way that could affect the subsequent evaluation of the incident's causes before informing the competent authorities. Rolling back the model, patching the prompt, and quietly retraining are all natural instincts. Sequence them after notification, or document meticulously why the alteration was unavoidable.

3. Reconstruct the decision: what did the system actually do?

Everything downstream (the report, the communications, the remediation) depends on answering one question precisely: what happened? For an AI decision, that decomposes into the elements we have cataloged in detail elsewhere: the inputs as the system saw them at decision time, not as the source databases read today; the exact model and pipeline version that ran, not the deployment window it probably fell inside; the policy and thresholds in force at that moment, with a revision identifier; and the human actor, if any: what they were shown, what authority they had, what they did. One substitute does not work: an explanation generated after the fact is an account of model behavior, not proof of what happened, as we have argued separately.

For organizations that write a decision record at issue, reconstruction is retrieval: an afternoon. For organizations working from operational logs, it is a join across deploy histories, CI systems, configuration state, and change tickets, none designed to be joined, several already partially rotated. Plan for this honestly in the incident plan: if your current answer to "which build produced this output?" is an archaeology project, the project starts in hour one and its duration is on the critical path of every deadline in the next section.

One discipline matters here more than speed: separate what is established from what is suspected. The reconstruction produces a timeline with confidence levels, because the regulatory reports and public statements that follow will be judged against what you knew and when you knew it.

4. What are the AI incident reporting requirements?

This is where the clock stops being metaphorical. Multiple reporting regimes can attach to a single AI incident, they run concurrently, and the shortest one sets your tempo.

Figure 1Five workstreams against the clock

Preserve

Issue the litigation hold, suspend log rotation, snapshot inputs, outputs, model artifacts, configurations, and the human review trail, before anyone fixes anything.

The clock · Reasonable anticipation of litigation, which a serious incident usually establishes on day one. Rotation and ephemeral infrastructure destroy evidence on their own schedule, so the real deadline is measured in hours.

The bars overlap because the work does: these are concurrent workstreams, not phases. Brass pins mark the outer reporting bounds that land inside the window (DORA initial notification at 4 hours from classification and 24 hours from awareness; the AI Act two-day clock for critical-infrastructure incidents and widespread infringement). The AI Act 10 and 15 day clocks, and GDPR 72 hours, sit past the right edge, and none of them pauses while reconstruction catches up.

The EU AI Act's Article 73 is the dedicated regime. Providers of high-risk AI systems must report serious incidents to the market surveillance authorities of the Member State where the incident occurred. Three deadlines, graded by severity, all counted from awareness:

  • 15 days for serious incidents generally, with the report due "immediately after the provider has established a causal link between the AI system and the serious incident or the reasonable likelihood of such a link"; the 15 days are a ceiling, not an allowance.
  • 10 days in the event of a death.
  • Immediately, and not later than 2 days, for a widespread infringement or a serious incident involving serious and irreversible disruption of critical infrastructure.

Three design features of the regime deserve attention. The Act explicitly permits an incomplete initial report followed by a complete one: the drafters expected you to report before you fully understand. Deployers have their own duty under Article 26(5): on identifying a serious incident, they must immediately inform first the provider, then the importer or distributor and the market surveillance authorities. And after reporting, the provider must perform the necessary investigations, including a risk assessment and corrective action, in cooperation with the authorities. The European Commission published draft guidance and a reporting template for Article 73 in September 2025, confirming among other things that an indirect causal link suffices to trigger the duty. Penalties for non-compliance run up to EUR 15 million or 3 percent of worldwide annual turnover, whichever is higher.

On timing: the application dates for the Annex III high-risk regime are in flux under the Digital Omnibus (see our Mobley analysis on the deferral), but the reporting duty for general-purpose AI models with systemic risk, under Article 55, is already in application, and the Commission published its reporting template for those incidents in November 2025. Treat the exact dates as moving and the architecture of the duty as settled.

The sectoral stack runs alongside, not instead. Depending on what the incident touched:

  • Personal data: GDPR Article 33 requires notifying the supervisory authority without undue delay and, where feasible, within 72 hours of awareness of a personal data breach, and Article 33(5) requires documenting every breach, notified or not.
  • EU financial entities: under DORA and its incident-reporting technical standards, a major ICT-related incident requires an initial notification within 4 hours of classifying it as major and no later than 24 hours after awareness, an intermediate report within 72 hours of the initial notification, and a final report within one month. An AI system failure inside a bank can be an ICT incident in exactly this sense.
  • Medical devices: under FDA regulations at 21 CFR Part 803, manufacturers report deaths, serious injuries, and reportable malfunctions within 30 calendar days, compressed to 5 work days where the event requires remedial action to prevent an unreasonable risk of substantial harm to public health.

The Commission's draft guidance addresses the overlap directly: where a sector already has an equivalent reporting regime (NIS 2, for instance), Article 73 reporting is confined to fundamental-rights infringements, with the rest following the sectoral rules. The practical consequence for the first 48 hours: someone's explicit job, starting in hour one, is to determine which regimes attach, in which jurisdictions, with which clocks. That determination is itself time-stamped evidence of diligence.

5. Communicate without deciding facts you have not established

Every statement made in the first 48 hours (to affected people, to regulators, to the press, in internal channels) becomes part of the record of the incident, discoverable and quotable. The failure modes are symmetric. Say too much, and you have asserted a cause you have not established: "an isolated error" is a factual claim that will be tested against whatever the reconstruction later shows, and a confident early misstatement reads afterward as concealment. Say nothing, and you have failed duties of candor that several of the regimes above impose, and surrendered the narrative besides.

The discipline that threads this is borrowed from the reporting regimes themselves, which are built for staged truth: the AI Act's incomplete-initial-report mechanism and DORA's initial, intermediate, and final sequence both assume knowledge improves over days. Statements should carry the same structure: what is established (and how), what is being done (hold, reconstruction, reporting, containment), what is not yet known, and when more will be said. Nothing in that structure admits liability; all of it demonstrates control. The people drafting these statements should be working from the reconstruction's established-versus-suspected ledger, which is one more reason the ledger exists.

Internally, the equivalent discipline is a decision log for the response itself: who decided what, on what basis, at what time. GDPR's Article 33(5) states the general principle for one domain (document every breach, including those not notified); the same posture serves the whole incident. The response is a sequence of judgment calls made under uncertainty, and the organizations that fare best in later scrutiny are the ones that can show each call was reasonable on the information available at the time.

6. Remediate, and prove the fix is a fix

Remediation has two halves, and the second is the one that fails silently.

The first half is the fix: correct the defective threshold, roll back or retrain the model, adjust the policy, re-run the affected decisions. Sequenced, as noted, against Article 73's constraint on altering the system before authorities are informed.

The second half is scope and proof. A defect that produced one discovered harm has usually touched other decisions, and the first question a regulator will ask about the fix is: which ones? With policy and version identifiers sealed into each decision record, that is a query. Without them, it is the multi-week reconstruction exercise that consumes organizations in every remediation, and the honest answer to "who else was affected?" becomes "we cannot be certain," which is an expensive sentence in every room it gets said in. Then the fix itself must enter the record: what changed, when, authorized by whom, validated how, so that next year's audit can distinguish the remediated system from the one that failed. NIST's MANAGE function frames this as continual improvement integrated into system updates (MANAGE 4.2) and closes the loop with communication: incidents and errors communicated to relevant AI actors, "including affected communities" (MANAGE 4.3). Contributing what can be shared to public incident resources is part of the same logic; the AI Incident Database exists so that the next organization's incident is not a rediscovery of yours.

7. Readiness is a recordkeeping property

Look back across the five workstreams and notice what each one consumed.

The hold in section 2 consumed an inventory: knowing where decision records live, so the freeze is a list and an instruction rather than a guess. The reconstruction in section 3 consumed the decision record itself: inputs, version, policy, reviewer. The reports in section 4 consumed established facts on a deadline shorter than any reconstruction project: a 2-day or 4-hour clock is not survivable by archaeology. The communications in section 5 consumed the boundary between established and suspected. The remediation in section 6 consumed version and policy identifiers across the whole decision history.

Figure 2What each step consumes
  • 01

    Preserveconsumes an inventory of where decision evidence lives

    The hold is a guess. Every system that might hold a fragment gets frozen late, and rotation has already eaten the first days of the window.

    Weeks
  • 02

    Reconstructconsumes the decision record: inputs, version, policy, reviewer

    Archaeology. Deploy logs, CI history, configuration state, and change tickets, joined by hand while every other workstream waits on the answer.

    Weeks
  • 03

    Reportconsumes established facts, on a statutory clock

    The clock outruns the archaeology. The report is late, hedged, or wrong, and the amendment trail becomes part of the record of the incident.

    Weeks
  • 04

    Communicateconsumes the boundary between established and suspected

    Statements guess. The guesses harden into commitments, and the reconstruction later contradicts them in writing.

    Weeks
  • 05

    Remediateconsumes policy and version identifiers across past decisions

    Scoping is a project measured in weeks, and its conclusion is a qualified estimate the regulator is invited to test.

    Weeks

None of the records in the left-hand headers can be created during the incident; each exists because it was written at decision time, or it does not exist at all. Flip the state and the same five steps change from a response into a salvage operation, which is the sense in which readiness is a recordkeeping property.

None of these records can be created during the incident. That is the whole point of the framework, and it is why the most useful incident-response exercise is not a tabletop of the communications plan. It is the drill we have recommended before: pick one real decision from eighteen months ago and attempt to produce its full record, timed. If the answer takes a week, your incident response plan has a week-long hole in the middle of it, positioned exactly where the two-day reporting clock runs.

The infrastructure conclusion follows without much argument, and it is the reason Sigilith's platform exists: decisions checked against the policy in force and sealed as tamper-evident records at the moment of issue produce, as a side effect, precisely the object every workstream above consumes. But the framework stands on its own, whatever tooling you use. The first 48 hours after an AI failure are a test administered without notice, and the grade is mostly determined before the incident begins: by whether, on the day your system did the wrong thing, there was already a record of exactly what it did.

Sources

Regulation and guidance

Preservation and evidence

Frameworks and incident data

Related Sigilith analysis

Also Applicable To

Public Sector
Critical Infrastructure
Telecommunications
Sigilith

Evidence infrastructure for consequential AI decisions: records built to outlive the systems that made them.

Est. in the decision path

{ CORRESPONDENCE }

syed@sigilith.com

Vendor-risk questionnaires and security reviews are welcome with a first message.

LinkedIn

© 2026 Sigilith, Inc. · A Delaware corporation. All rights reserved.

Set in Instrument Serif · Inter · IBM Plex Mono