Back to Insights
CASE STUDYAUGUST 7, 202622 min read

1.1 Billion Decisions, No Record: What Mobley v. Workday Reveals About AI Evidence

Sigilith Research

Institutional AI governance & accountability

A note on what this analysis is and is not. Mobley v. Workday is active litigation. Workday disputes the allegations against it. No court has found that Workday's products discriminated against anyone. Every ruling described here concerns procedure (what claims may proceed, what documents must be produced), not the merits. This analysis is not about whether Workday did anything wrong. It is about something the case has incidentally exposed, which applies to every organisation deploying AI in decisions that matter: the records needed to explain an AI-influenced decision are often held by nobody in particular.

At a glance

CaseMobley v. Workday, Inc., No. 3:23-cv-00770 (N.D. Cal.)
FiledFebruary 2023
JudgeHon. Rita F. Lin (district); Hon. Laurel Beeler (magistrate, discovery)
Claim typeDisparate impact: age, race, disability, sex
CertifiedNationwide ADEA collective, May 2025
Exposure window24 September 2020 – present
Applications referenced in court filings~1.1 billion
Potential collective sizeWorkday's own estimate: "hundreds of millions"
Employers using Workday10,000+, including 65%+ of the Fortune 500
Third-party subpoenas served on customers~35, against a disclosed list of ~353
Opt-in windowClosed 7 March 2026; collective membership is now fixed
StatusDiscovery. No trial date. No finding of liability.

1. The short version

In February 2023, Derek Mobley sued Workday, alleging that its algorithmic applicant-screening tools produced discriminatory outcomes. Three and a half years later, the case has become the most consequential piece of AI litigation in the United States, not because of anything a court has decided about discrimination, but because of what the discovery process has revealed about whether anyone can reconstruct what an AI system actually did.

On 29 May 2026, Magistrate Judge Laurel Beeler ruled on the discovery disputes. The plaintiffs wanted two things above all: Workday's internal bias-testing data, and the applicant data held in Workday's systems on behalf of its customers.

They got neither.

The bias-testing data was ruled privileged, because Workday's lawyers had curated the underlying data and used the results to give legal advice. The applicant data was ruled outside Workday's control for discovery purposes: Workday hosts it, but the customers own it, and the court found no legal right to obtain it on demand.

The plaintiffs did not leave empty-handed. The court ordered Workday to produce its EEO-1 and OFCCP filings, holding them relevant to what Workday knew about demographic disparities in its own workforce data. That matters, and it is worth saying plainly in an analysis that is otherwise about what could not be produced. But EEO-1 filings are aggregate workforce reports. They describe populations. They cannot describe a decision.

So the plaintiffs' response was to serve roughly 35 third-party subpoenas on Workday's customers.

Pause on that. To determine what happened in a set of automated hiring decisions, litigants are now issuing subpoenas across dozens of separate corporations, because the record of a single decision (this applicant, this system version, this configuration, this output, this human review) does not exist as a single retrievable object anywhere.

It was never written.

2. How the case escalated

A brief chronology, because the shape of the escalation matters more than any single ruling.

February 2023: Complaint filed. At the time, widely treated as a novel but narrow employment matter.

12 July 2024: The first genuinely significant ruling. Judge Lin holds that an AI vendor can be directly liable for employment discrimination under an "agent" theory. The court reasons that Workday's software does not merely implement employer criteria mechanically; it participates in the decision by recommending who advances and who is rejected. An employer, the court holds, cannot escape liability by delegating a traditional function like hiring to a third party. And that third party can be independently liable.

This is the ruling that turned a lawsuit into a category. It established that in an AI-mediated decision, liability follows the decision-making function, not the corporate boundary.

May 2025: Preliminary certification of a nationwide collective under the Age Discrimination in Employment Act.

July 2025: The court orders Workday to disclose every employer that enabled HiredScore AI features. Workday had argued HiredScore, acquired after the original complaint, was a separate product on a separate platform and outside the collective's scope. The court disagrees and expands the collective to include applicants "scored, sorted, ranked, or screened" using HiredScore. Disclosure deadline: 20 August 2025.

2 December 2025: The court approves the plan for notifying the collective. Notice goes out to anyone who applied through the platform since 24 September 2020 and was 40 or older at the time. Court filings reference approximately 1.1 billion applications processed in the relevant period. Workday itself suggests the collective could run to "hundreds of millions" of people. The court notes, pointedly, that allegedly widespread discrimination is not a reason to deny notice.

6 March 2026: Workday argues the ADEA's disparate-impact protections don't cover applicants at all, only current employees, and that Loper Bright undercuts the precedent holding otherwise. Judge Lin rejects both arguments. Two claims are dismissed with leave to amend.

7 March 2026: The opt-in window closes. Collective membership is now fixed, which means the scale of the matter is no longer speculative to the parties, whatever remains unresolved publicly.

30 March 2026: Plaintiffs file an amended complaint curing the deficiencies.

29 May 2026: The discovery order. See section 3. This is the one that matters for everyone who isn't a party.

22 June 2026: The court grants in part and denies in part Workday's motion to dismiss. The California FEHA claims survive: the court holds that plaintiffs adequately alleged a California nexus because Workday designs, develops and operates the screening tools from its California headquarters. A proxy-discrimination disability claim survives too. The case now spans race, sex, age, and disability.

Running alongside it: a class action filed on 20 January 2026 in Contra Costa County Superior Court against Eightfold AI, alleging the company operated as a consumer reporting agency: compiling applicant profiles from social media, location, device and cookie data, scoring candidates on a nought-to-five scale, and letting employers discard the low scores before a human saw them, all without the disclosures the Fair Credit Reporting Act requires.

The two cases attack from opposite directions. Mobley says the vendor is an agent, liable for outcomes. Eightfold says the vendor is a reporting agency, bound by process and transparency duties. An AI vendor now faces both framings simultaneously. So, increasingly, does the enterprise that deployed the tool.

3. The finding that matters: nobody holds the record

Strip away the employment-law specifics and the May 2026 discovery order describes a structural condition.

The bias-testing data. Workday had tested its systems for disparate outcomes. That testing data was exactly what plaintiffs wanted. The court held it privileged: Workday's attorneys had curated the underlying data and used the results in providing legal advice.

The mechanism here is worth understanding, because it is not an accident and it is not unusual. When an organisation runs its most probing self-examination under legal privilege (a rational, routine, defensible choice), that examination becomes unavailable as evidence. The organisation knows what it found. It cannot easily show anyone. And the analysis that is producible tends to be the analysis that was never probing enough to need protecting.

The applicant data. Plaintiffs sought the underlying applicant records. The court declined to compel production, finding plaintiffs hadn't established that Workday legally controlled that data under Rule 34. The Master Subscription Agreement let Workday produce customer data under a court order; that, the court held, is not the same as a right to obtain it on demand. Possession and control are not the same thing, and the difference decides who can be made to produce what.

And then the part that should have been the headline. Several of the third-party employers plaintiffs had already subpoenaed took the position that plaintiffs should seek the data from Workday instead. The court had just held that Workday does not control it. The court's response was to encourage the parties to work it out between them.

That is not a gap. That is a loop. The employer points at the vendor; the vendor is held not to control it; the plaintiff is left holding a subpoena that every recipient can plausibly say belongs to someone else. Nobody has to be acting in bad faith for this to happen, and it is worth saying that nothing in the record suggests anyone is. The loop is a property of the architecture, not of the parties.

The consequence. Roughly 35 subpoenas to third-party employers. A 353-customer disclosure list. A multi-year discovery exercise to assemble something that could have been a single record written at the moment of each decision.

Now the part that should concern every executive reading this:

The plaintiffs are not the only party that cannot reconstruct these decisions. Neither, in all likelihood, can Workday. Neither can the employers.

This is not an allegation of concealment. It is a description of how enterprise AI is architected. The vendor holds the model but not the outcome. The employer holds the outcome but not the model version, the configuration history, or the scoring logic. The most searching analysis sits behind privilege. No party ever wrote down, at the time, the one thing that would have answered the question: this applicant, this system version, this configuration, this input, this output, this human review, this final action.

Figure 1Who holds each element of a reconstructable decision record
VendorEmployerBehind privilegeNobody, contemporaneously
  • Configuration is mutable by design and usually stored as current state, not as history. The value in force eighteen months ago is frequently unrecoverable from either side.

Custody assignments describe the typical enterprise architecture described in the rulings, not findings of fact about any party. Select any row to see why that element resists production. 3 of 8 elements have no contemporaneous holder at all.

Rohan Sharma, a US delegate to ISO/IEC JTC 1/SC 42 and a member of the OECD.AI expert group on risk and accountability, put the general principle precisely in Corporate Compliance Insights: the records needed to evaluate an AI-influenced decision "may be distributed among parties that do not share the same retention obligations, access rights or litigation strategy."

Distributed among parties. Not missing. Not destroyed. Distributed, which in practice is the same as absent, because no one party can assemble them and no party has the right to compel the others quickly.

4. The distinction almost everyone gets wrong

Organisations use "auditability" as though it names a single capability. It names two, and they are not substitutes.

A bias audit is a population-level statement. It compares selection rates across demographic groups, examines error rates, tests for statistically significant disparities. It answers: does this system produce skewed outcomes across a population? It is a genuinely valuable control and several jurisdictions now require it.

Decision reconstruction is an individual-level statement. It answers: what happened to this person, on this date, and why? That requires the eight elements in Figure 1 above: the deployed system, the configuration, the inputs, the derived features, the output, any change between validation and deployment, the human review and its authority, and the final action.

A system can pass the first and fail entirely at the second. A vendor can clear a population-level assessment and still be unable to recreate any particular decision. An employer can retain the application and the final disposition and still lack the model version, configuration history, and validation evidence that would explain it.

Regulators ask the first question. Litigants, applicants, and ombudsmen ask the second.

New York City's Local Law 144 illustrates the gap. It requires an independent bias audit, a published summary, candidate notification, and disclosure of data-retention policy. Real transparency obligations, and none of them guarantee that the technical records needed to reconstruct one individual's decision are allocated to anyone in particular.

California goes further on retention: employers and other covered entities must keep employment records including automated-decision-system data for at least four years, and providers of automated-decision systems must retain relevant records for at least four years after the system was last used.

But retention is not production. A record can exist and still be unavailable, because the employer has no contractual right to obtain it promptly, or lacks the ability to interpret it, or because the party holding it has its own litigation posture. Mobley is the demonstration.

5. The retention trap

The Mobley exposure window opens on 24 September 2020.

Read that as an operational fact rather than a legal one. Decisions made in late 2020 are now evidence. Systems that have been retired, retrained, reconfigured, and replaced several times over are now the subject of forensic inquiry.

An organisation might have commissioned a rigorous bias audit in 2025 and still have no contemporaneous record of how its screening behaved in 2021, 2022, or 2023, the years actually under examination. A 2025 audit is not evidence about 2021. It is evidence about 2025.

One employment law commentator's framing of this is blunt and worth repeating: in litigation, that is not a compliance gap, it's negligence.

Whether or not that characterisation would survive a court, the operational point stands. Evidence has to be created at the time. It cannot be manufactured retroactively, and the attempt to do so is itself discoverable.

The exposure window on any AI decision made today is not the current audit cycle. It is the longest applicable limitation period, plus the duration of any inquiry: realistically five to eight years. Every AI-assisted decision your organisation issues this quarter carries an evidentiary obligation extending into the 2030s, against standards that do not yet exist.

6. The contract gap

There is a second failure mode, upstream of the technology.

Analysis of AI vendor agreements by the contract-review platform TermScout, reported through Stanford's CodeX centre in March 2025, found that 88% of AI vendors cap their own liability, often at the level of monthly subscription fees, while only 17% commit to complying with all applicable laws. The sample is described as thousands of IT services agreements; the methodology is a commercial contract-review dataset rather than a peer-reviewed study, and the figures should be read as directional.

What makes the dataset useful is not any single number. It is that the same platform scored general SaaS agreements on the same terms, which turns two isolated statistics into a comparison. And the comparison says something the isolated numbers do not.

Figure 2AI vendor contracts take more and promise less than the SaaS contracts they are modelled on
Vendor claims data-usage rights beyond service delivery+29 pts
AI
92%
SaaS
63%
Vendor commits to complying with all applicable laws19 pts
AI
17%
SaaS
36%
Contract includes compliance-related warranties25 pts
AI
17%
SaaS
42%
Contract caps the customer's liability6 pts
AI
38%
SaaS
44%

Source: TermScout contract dataset, reported via Stanford Law's CodeX centre, March 2025; sample described as thousands of IT services agreements. A commercial contract-review dataset rather than a peer-reviewed study: read the gaps as directional, not precise. Separately, 88% of AI vendors cap their own liability, against the 38% who cap the customer's.

AI contracts take more and promise less than the SaaS contracts they are modelled on. Vendors claim data-usage rights beyond what service delivery requires at 92%, against a 63% market average. They commit to legal compliance at 17%, against 36%. They offer compliance warranties at 17%, against 42%. And the liability caps run in one direction: 88% of AI vendors cap their own exposure, while only 38% cap the customer's, a lower rate than the 44% general SaaS figure.

Read that last pair slowly. The one protection these contracts are less likely to extend than an ordinary software agreement is the one that would limit what the customer can be held responsible for.

The practical consequence: an employer sued over discriminatory outcomes produced by a vendor's screening tool may find its contract provides audit rights, security commitments, and indemnification, but no mechanism to obtain the evidence needed to explain what happened. The indemnity pays a legal bill. It cannot recreate a decision.

Sharma's prescription is the right one: enterprise AI agreements need an evidence schedule that names, for each category of record, which party creates it, controls it, retains it, and must produce it. Seven areas at minimum:

  1. System identity and versioning: which model, which version, and how changes are recorded
  2. Decision-event records: linking applicant, requisition, system version, configuration, input, output, timestamp, and final action
  3. Validation and impact-assessment evidence, with an explicit distinction between routine compliance testing and analysis conducted for legal advice, because those receive different treatment in litigation
  4. Data control and legal holds: not "who owns the data" but possession, legal control, exportability, subpoena response, preservation
  5. Human review: not a checkbox. Reviewer identity, information presented, authority to override, reason for the decision
  6. Retention and termination: usable exports before migration or deletion
  7. Audit and remediation rights, including the right to suspend automated screening without breaching volume or exclusivity commitments

Point 3 deserves emphasis. Running all your testing under privilege protects it from disclosure and simultaneously renders it useless as a defence. You cannot rely on evidence you have successfully made unavailable. Organisations need two tiers: privileged legal analysis, and producible operational records generated in the ordinary course of business. Mobley shows what happens when only the first exists.

7. What it costs

Direct litigation exposure in Mobley is impossible to estimate responsibly: the collective size is contested, no liability has been found, and settlement values in disparate-impact collectives vary enormously.

The reconstruction costs, however, are the same in every matter of this shape and can be reasoned about. The model below is arithmetic, not a finding: it multiplies a per-recipient cost band by the number of organisations pulled in. Both inputs are adjustable, because the honest thing to do with an estimate is to show its working rather than assert a total.

Figure 3Third-party discovery cost, with its working shown

External legal + vendor

$2.6M$8.8M

Internal effort

210420 person-weeks

across engineering, HR and legal, at 6–12 weeks each

The same question, answered two ways

Record exists and verifiesHours
Record must be reconstructedMonths

Drawn to scale would make the first bar invisible: the gap is roughly three orders of magnitude, and nothing applied afterwards closes it.

Arithmetic, not a finding. Per-recipient costs assume external counsel and vendor fees to scope, negotiate, review and produce, in the $75k–250k band the matter's shape implies. Adjust the lower bound above and the band moves with it. Internal reconstruction effort is reported separately because it does not arrive as an invoice, and excluded from the dollar total for the same reason.

At the Mobley figure of roughly 35 recipients and a conservative $75k–250k per recipient in external legal and vendor costs, third-party discovery costs land in the $3–9M range, borne largely by companies that are not defendants, did not choose the litigation, and have no control over its outcome.

Internal reconstruction is the larger, quieter line. Each subpoenaed employer must determine what its own system did years ago: engineering time to locate archived configurations, HR time to match applicants to requisitions, legal time to assess. Six to twelve weeks of cross-functional effort per organisation is a realistic estimate, and it does not appear in the model above because it does not arrive as an invoice.

The asymmetry is the point. When a record exists and is verifiable, responding to an inquiry is a retrieval operation: hours. When it does not, it is a reconstruction project: months, with an uncertain output that may itself become contested. The cost multiple between those two states is roughly three orders of magnitude, and no amount of tooling applied afterwards closes it, because the missing input is a record that was never written.

The strategic cost. The most expensive consequence isn't the legal bill. It is that organisations watching this stop deploying AI in decisions where it would create value, because they cannot answer the question their general counsel now asks first: if this is challenged in 2031, what will we produce? The inability to answer that question is now a material brake on enterprise AI adoption: a cost that never appears on any invoice.

8. The regulatory calendar just moved. The evidentiary one didn't.

This is the part of the analysis that has changed most since the Mobley discovery order, and it is the part most likely to be getting read backwards inside large organisations right now.

Through 2024 and 2025, the reasonable planning assumption was that a wave of AI regulation would arrive on a known schedule and would tell you what to keep. Two of the most demanding regimes were Colorado's AI Act and the EU AI Act's high-risk obligations. Both have since retreated:

  • Colorado. SB 24-205 was postponed once, to 30 June 2026. Then on 14 May 2026 the governor signed SB 189, which replaced it outright, pushing the effective date to 1 January 2027 and removing the duty of care against algorithmic discrimination, the deployer obligation to run a risk-management programme, the impact-assessment requirement, and the attorney-general reporting duty. What remains is a narrower disclosure and transparency regime.

  • The EU. Under the Digital Omnibus on AI, provisionally agreed on 7 May 2026, the Annex III high-risk obligations are deferred by sixteen months, from 2 August 2026 to 2 December 2027. Annex III is the annex that covers employment and worker management. Article 12, the automatic event-logging requirement, travels with it. This still requires formal adoption, so treat the date as provisional rather than settled.

Now put that beside what happened in the same six months in the Northern District of California: the discovery order of 29 May 2026, and the 22 June ruling expanding the case to race, sex, age and disability.

Figure 4The two calendars are moving in opposite directions

Statutory obligationlater and lighter

Evidentiary exposureearlier and heavier

2020202420282032

Aug 2026 → Dec 2027EU AI Act · Annex III high-risk

Annex III covers employment and worker management, and carries the Article 12 automatic event-logging duty. Under the Digital Omnibus on AI, provisionally agreed 7 May 2026, those obligations are deferred from 2 August 2026 to 2 December 2027. Formal adoption is still outstanding, so treat the date as provisional rather than settled.

Statutory dates as at August 2026; the EU deferral is provisionally agreed and awaits formal adoption. Select any row for the detail. The exposure track is set not by legislatures but by limitation periods and the discovery rules, which is why it does not move when a commencement date does.

The two calendars are moving in opposite directions. Statutory record-keeping obligations got later and lighter. Evidentiary exposure got earlier and heavier. And the second calendar is not set by legislatures at all; it is set by limitation periods and the discovery rules, which is why a case filed in 2023 is currently examining decisions made in September 2020, years before any of these statutes existed or was even drafted.

The practical error this invites is specific and expensive: treating the delay as breathing room. An organisation that had been pacing its evidence programme to the EU's August 2026 date has just been handed sixteen more months of obligation relief and precisely zero months of exposure relief. Every decision issued during that window is still discoverable, still sits inside a limitation period running into the 2030s, and still cannot be reconstructed afterwards if the record was never written. The reprieve is real. It applies to the wrong calendar.

The rest of the landscape has not moved:

  • NYC Local Law 144: bias audit, published summary, candidate notice
  • California ADS regulations: four-year retention for employers and providers
  • Illinois: AI-use disclosure in employment decisions
  • ISO/IEC 42001 and 42005: AI management systems and impact assessment
  • NIST AI Risk Management Framework: currently under revision

None of these is a litigation safe harbour. ISO/IEC 42001 certifies that you operate an AI management system. It does not certify that any particular decision was lawful. Standards structure controls; they do not substitute for evidence.

The direction of travel is unambiguous even where the dates have slipped. Any decision denied or modified by an AI-assisted process increasingly needs a decision trace producible to a regulator or in litigation: not a denial code, but a record of what the system saw, what rules applied, and what reasoning followed.

9. What an evidence layer changes

The counterfactual is worth working through concretely, because the fix is less exotic than the problem suggests.

Suppose that at the moment each screening decision was issued, a single sealed record had been written containing:

  • The exact output: score, ranking, recommendation, as delivered
  • The system version: model or rules engine identifier, frozen at decision time
  • The configuration: thresholds and criteria in force for that customer, that requisition
  • The inputs: the applicant data actually presented to the system
  • The policy: the screening rules that authorised the outcome, with a revision number
  • The human review: reviewer identity, what they saw, their authority, their action
  • The timestamp: sealed at issue, not asserted afterwards
  • A cryptographic seal binding all of the above, verifiable by anyone

Six things change.

One: the question becomes retrievable rather than reconstructable. What happened to applicant X on 14 March 2022? is answered by fetching a file, not by a six-week forensic exercise across three companies.

Two: the loop dissolves. The stand-off in section 3 (employers pointing at the vendor, the vendor held not to control the data) exists because the record lives in exactly one party's system and its custody is therefore contestable. A sealed record can be held by both parties, independently verifiable by each, with neither dependent on the other's cooperation or litigation posture. There is nothing left to point at each other about.

Three: the privilege trap disappears. A record generated automatically in the ordinary course of business, before any dispute exists, is not attorney work product. It is a business record. It is producible, and therefore usable as a defence. Organisations stop having to choose between protecting their analysis and being able to rely on it.

Four: retention becomes verifiable rather than asserted. California's four-year requirement is satisfied by a record that can be proven unaltered across four years, not by a database row that must be taken on trust.

Five: reconstruction becomes possible across model retirement. The system that made a 2021 decision no longer exists. Its identity, version, and configuration, sealed at the time, still do.

Six, and this is the one executives care about most: remediation scoping becomes tractable. When a policy or threshold is found defective, the question is immediately how many past decisions are affected and who are they? With sealed records containing the policy revision and configuration for each decision, that is a query. Without them, it is the six-week exercise that is currently consuming dozens of companies in this matter.

What an evidence layer does not do, and any vendor claiming otherwise should be treated with suspicion: it does not make a discriminatory system lawful. It does not substitute for bias testing. It does not prevent disparate impact. A perfect record of an unlawful decision is a perfect record of an unlawful decision.

What it does is convert an unanswerable question into an answerable one, and give organisations that are operating lawfully the ability to demonstrate it, which today most of them cannot.

10. Seven things to do now

For any organisation using AI in decisions affecting customers, employees, applicants, or claimants:

  1. Map your exposure window. Identify the earliest AI-assisted decision still within a limitation period. That date, not your last audit, is where your evidentiary obligation begins.

  2. Test one reconstruction. Pick a single decision from eighteen months ago and try to produce the full record: system version, configuration, input, output, human review. Time the exercise. The result will be more persuasive to your board than any assessment.

  3. Separate audit from reconstruction in your controls, your budget, and your vocabulary. Confusing them is the single most common error in this area.

  4. Read your AI vendor contracts for evidence rights, not just liability caps. Ask: if we are subpoenaed about a decision this tool made, what are we contractually entitled to receive, and how quickly?

  5. Audit your privilege posture. If all your probing analysis is privileged, you have protected it into uselessness. Build a producible operational tier.

  6. Require a decision-event record, not a log line, for every consequential AI-assisted decision. Applicant, system version, configuration, input, output, reviewer, action, timestamp.

  7. Write the record at the moment of the decision. This is the only step that cannot be retrofitted. Everything else on this list can be started next quarter. This one is lost forever for every decision you issue before you do it.

Point 2 is the one we would start with, and it is why we built a two-minute version of it: you approve an ordinary decision, move eighteen months forward, and watch which parts of your own record survive. It takes rather less time than the real exercise, and it tends to end the internal debate about whether this is a real problem.

Sources

Also Applicable To

Public Sector
Critical Infrastructure
Telecommunications
Sigilith

Evidence infrastructure for consequential AI decisions: records built to outlive the systems that made them.

Est. in the decision path

{ CORRESPONDENCE }

syed@sigilith.com

Vendor-risk questionnaires and security reviews are welcome with a first message.

LinkedIn

© 2026 Sigilith, Inc. · A Delaware corporation. All rights reserved.

Set in Instrument Serif · Inter · IBM Plex Mono