The Arup Deepfake: A $25 Million Video Call, and the Two Controls That Were Missing
Sigilith Research
Institutional AI governance & accountability
A note on what is known and how. This case study reconstructs the fraud from three kinds of source, kept distinct throughout: what Hong Kong police said publicly (a media briefing on 2 February 2024 and statements to reporters), what Arup itself confirmed (a brief statement in May 2024, after press reporting identified the firm), and details that exist only in press accounts of the police statements. The employee at the center of the case was the victim of a deception and stands accused of nothing. No arrests have been publicly announced, no recovery has been disclosed, and because the matter has not been tried, none of the details below have been tested as evidence in any court. Where accounts conflict on minor points, this analysis keeps to what the police and the firm actually said.
At a glance
| Target | Arup's Hong Kong office (confirmed by the firm, May 2024) |
| First contact | Mid-January 2024: a message purportedly from the firm's UK-based CFO about a confidential transaction |
| The call | A multi-participant video conference; police said every participant the employee saw was an AI-generated likeness of a real colleague |
| Transfers | 15, to five local bank accounts |
| Amount | HK$200 million, about US$25.6 million at the time |
| Reported to police | 29 January 2024 |
| Disclosed | Police media briefing, 2 February 2024, company unnamed |
| Classification | Obtaining property by deception; assigned to the cybercrime unit |
| Status | No publicly announced arrests; no disclosed recovery |
1. What happened in the Arup deepfake scam
In mid-January 2024, an employee in the finance department of the Hong Kong office of Arup, the London-headquartered engineering consultancy whose credits include the structural engineering of the Sydney Opera House, received a message that appeared to come from the firm's chief financial officer in the United Kingdom. A confidential transaction was required, it said. According to Hong Kong police, the employee's first instinct was the correct one: it read like phishing.
Then came the video call. The employee joined a group video conference with the CFO and several colleagues, people whose faces and voices matched people the employee knew of. Acting Senior Superintendent Baron Chan Shun-ching later told the city's public broadcaster RTHK: "(In the) multi-person video conference, it turns out that everyone [he saw] was fake." Every participant except the victim was an AI-generated likeness of a real person, assembled, police said, from publicly available video footage with generated voices added.
The doubt the message had raised did not survive the meeting. Over the days that followed, acting on instructions from the call and on follow-up contact, the employee made 15 transfers to five local bank accounts, totaling HK$200 million: about US$25.6 million at the time, rounded in international coverage to anywhere from US$25 million to US$26 million.
The deception ended the way it could have been prevented: with a verification. The employee followed up with the firm's UK headquarters, and the report reached Hong Kong police on 29 January 2024, roughly two weeks after the first message. On 2 February, police described the case at a media briefing without naming the company, classified it as obtaining property by deception, and called it the first case of its kind they had seen: a fraud staged inside a multi-participant deepfake video conference.
The company stayed unnamed for three months. In May 2024, after press reports identified the firm, Arup confirmed it was the target, saying it had notified Hong Kong police in early 2024, that "fake voices and images" were used, and that "our financial stability and business operations were not affected and none of our internal systems were compromised." Arup's global chief information officer, Rob Greig, added: "Like many other businesses around the globe, our operations are subject to regular attacks, including invoice fraud, phishing scams, WhatsApp voice spoofing, and deepfakes. What we have seen is that the number and sophistication of these attacks has been rising sharply in recent months."
An employee in the finance department of Arup's Hong Kong office receives a message purportedly from the firm's UK-based chief financial officer: a confidential transaction is required. Police later said the employee's first instinct was the right one; it read like phishing.
In the fraud's path · One person's suspicion. At this point, the fraud is losing.
Dates and details are drawn from Hong Kong police statements reported in February 2024 and from Arup's May 2024 confirmation, both cited in the Sources section. The line to read twice: the verification that ended the fraud on 29 January is the same act that, performed two weeks earlier, would have prevented it.
2. How the deepfake video call actually worked
The mechanics police described are less exotic than the headline suggests, and the gap between the two is the most instructive thing in the case.
This was not, on the police account, real-time puppetry. The impostor participants were built in advance from footage the executives had already published to the world, with AI-generated voices layered on. During the call, the employee was asked to give a self-introduction, but the other participants never genuinely interacted with the victim; they issued instructions, and the meeting ended abruptly. Contact then continued through instant messaging, email, and one-to-one video calls, where the transfer instructions accumulated.
Read as attack design, the choreography is careful. Everything that would have required convincing real-time synthesis was avoided. The one participant invited to speak was the real one; the self-introduction manufactured the feel of a genuine meeting while demanding nothing from the fakes. The abrupt ending removed the risk of unscripted conversation. The heavy persuasive work was done by the gestalt (a room apparently full of familiar, senior people), not by any single flawless forgery.
Chan drew the structural lesson at the briefing: "In the past, we would assume these scams would only involve two people in one-on-one situations, but we can see from this case that fraudsters are able to use AI technology in online meetings, so people must be vigilant even in meetings with lots of participants."
One more point keeps the case current. The pre-recorded shortcut was a 2024 convenience, not a 2024 limit. By December of that year, the FBI was warning that criminals use generated video for "real time video chats with alleged company executives." The Arup attackers avoided live interaction because they could; their successors increasingly do not have to.
3. Presence was assumed because faces appeared on a screen
Here is the uncomfortable reading: the fraud defeated no technical control, because no technical control was in its path.
The phishing message met the one defense actually present, a person's suspicion, and was losing to it. The video call existed to kill that defense, and it did. From that point forward the scheme ran on a chain of silent inferences: faces on a screen were taken as proof of presence; presence was taken as proof of identity; identity as proof of authority; and authority as license to move HK$200 million in fifteen installments. Every link in that chain was assumed. None was checked, because in most organizations there is nothing to check it with.
Remote identity verification learned this lesson the hard way years ago, and we have written about it in detail: a camera is a precise, tireless, completely credulous witness that reports whatever it is shown. A video conference client is the same witness at one further remove, a screen faithfully rendering whatever stream it is handed. The camera in the Arup case was never fooled. It displayed exactly what it was given. The people watching it supplied the rest, using instincts (recognition of a face, the demeanor of a colleague, the social weight of a full meeting room) that evolved for rooms and were carried, unexamined, into a medium that cannot support them.
That is why "train employees to spot deepfakes" cannot be the whole answer. The employee in this case was not careless; the employee was initially suspicious, and the suspicion was defeated by evidence that would have persuaded most people. When detection quality is an arms race between generative models and human eyes, betting the treasury on the eyes is a policy decision, and a poor one.
4. How fast is deepfake fraud against businesses growing?
The Arup case is the largest publicly confirmed loss to a deepfake video call to date, but it sits on a short, steep, verifiable ladder.
2019: one cloned voice. Fraudsters used AI-generated audio of a German executive's voice to persuade the chief executive of a UK energy subsidiary to wire €220,000 (about US$243,000), a case disclosed by the parent company's insurer and reported by the Wall Street Journal as among the first of its kind.
2020: a cloned voice at scale. Court filings described by Forbes recount a branch manager in the United Arab Emirates authorizing transfers of US$35 million after a call that cloned the voice of a company director, supported by forged emails about a pending acquisition.
January 2024: an entire meeting. The Arup case moved the technique from one voice on a phone to a room full of synthetic colleagues on video. At the same February briefing, Hong Kong police also described the adjacent industrial pattern: deepfake methods paired with eight stolen identity cards were used in 90 loan applications and 54 bank account registrations over three months of 2023, with facial recognition systems deceived in at least 20 instances.
2024: the near misses. In May, fraudsters impersonated WPP chief executive Mark Read, using a WhatsApp account bearing his photograph, a Teams meeting, a voice clone, and YouTube footage to solicit one of the group's agency leaders; the attempt failed. In July, a Ferrari executive received WhatsApp messages and then a call convincingly mimicking chief executive Benedetto Vigna's southern Italian accent, and ended the scam with a single question: what was the title of the book Vigna had recommended to him days earlier? The caller hung up.
Regulators now treat this as a category, not a curiosity. In November 2024, FinCEN issued a dedicated alert (FIN-2024-Alert004) reporting an increase since 2023 in suspicious activity reports describing suspected deepfake media, particularly generative-AI-produced identity documents used to defeat verification, and pointing to phishing-resistant multifactor authentication and live verification checks as mitigations. Weeks later the FBI's public service announcement cataloged generative AI across the fraud lifecycle, including real-time video impersonation of executives. Projections, for what they are worth (they are modeling, not measurement), run in the same direction: Deloitte's Center for Financial Services has estimated that generative AI could push US fraud losses to US$40 billion by 2027 in its aggressive scenario, from US$12.3 billion in 2023.
Two entries on that ladder failed, and it matters why. WPP's attacker was beaten by vigilance; Ferrari's was beaten by a freshness challenge, a question whose answer could not exist in any harvested footage. That is liveness verification performed by hand. It worked. It is also a control that lives entirely in one person's presence of mind, under manufactured urgency, against an opponent who chooses the moment. The lesson of the ladder is not that humans always lose; it is that whether they win is currently left to chance.
5. The two missing controls
Strip the case to its mechanism and the fraud needed exactly two assumptions to hold: that appearing in a meeting proves a person is present, and that a meeting's say-so authorizes a payment. Each assumption maps to a control that existing practice already knows how to demand.
- 01Assumed
Presence“Who is actually in this meeting?”
Whoever the screen shows. Every face and voice in the call was generated from footage the executives had already published, and nothing in the call could establish otherwise.
- 02Claimed
Authority“Who approved this payment?”
- 03Unguarded
Execution“What stands between instruction and transfer?”
One employee's judgment, already worked on by a phishing message and a persuasive hour of familiar faces. Fifteen transfers to five accounts followed.
- 04Reassembled
Record“What survives to reconstruct it?”
Message logs and recollection. Who asked, who approved, and on what basis must be reassembled afterward from whatever the fraud's own channels happened to retain.
Countermeasures are described at family level only: fresh challenges bound to capture time, out-of-band approval, policy in the payment path, sealed records. No gate makes deception impossible; each one raises the price of the same attack from harvested footage and a persuasive hour to defeating a purpose-built control the attacker cannot perform into.
5.1 Verified live presence
What the Ferrari executive improvised is, structurally, a challenge-response check: introduce something fresh, at this moment, that a recording made yesterday cannot answer. The FBI's advice to families (agree on a secret word; hang up and call back on a number you look up) is the same family of control in consumer form. The liveness detection field has spent a decade systematizing exactly this idea for remote onboarding, as one of several countermeasure families: challenges bound to the moment of capture, forensic attention to the medium carrying the face, and integrity checks on the session itself. We describe those families at the family level only, there and here; the ceiling on that disclosure exists because the details are useful to precisely the wrong readers.
The corporate translation is a policy, not a gadget: a video call is not an authentication channel, and no instruction that moves money is actionable on the strength of one. Presence, for consequential requests, gets verified through something built to verify it, whether that is an out-of-band callback through directory-sourced numbers, a standing shared secret, or, at the high end, machine-verified liveness of the kind regulated onboarding already uses. Recognition is not verification. A face on a screen is a claim, not a credential.
5.2 An authorization gate with a sealed record
The second assumption is quieter and did more damage. Fifteen transfers to five new beneficiary accounts, totaling HK$200 million, framed as secret and urgent, executed by a single employee: nothing in that path required an approval that the attacker could not stage. Dual control for high-value payments is far older than deepfakes. What the Arup case demonstrates is that video presence can no longer serve as the second channel, because the second channel must be one the requester cannot perform into.
A payment gate of the kind this case argues for has two halves. The preventive half is policy in the path: value thresholds, new-beneficiary friction, and a second approver reached through a known, independent channel, none of it waivable by a meeting, however senior the meeting looks. The evidentiary half is the record: each approval written at the moment of issue, capturing who authorized, what they were shown, which policy version applied, and when, sealed so that no one, including the organization itself, can quietly rewrite it. We have set out what such a record must contain elsewhere; the short version is that if the Arup fraud had to be reconstructed today, the reconstruction runs on chat logs and recollection, which is the general condition of decision evidence, not a special failing of one firm. Nor does reconstruction wait: the first 48 hours after an incident run several workstreams concurrently, each consuming records that had to exist before the incident began.
The honest hedge belongs here: no control catalog makes deception impossible, and it would be false comfort to claim these two gates guarantee this fraud fails. What they change is the price. As it ran, the attack needed harvested footage and a persuasive hour. Against a live presence check, it needs to defeat a purpose-built verification system in real time; against an authorization gate, it needs to independently compromise a second person through a channel it does not control; and against a sealed record, whatever it achieves is visible, attributable, and reconstructable afterward. Screens are cheap to fake. Each gate is there to make a screen insufficient.
6. How can a business prevent deepfake video call fraud?
The case compresses into six working rules, all of them family-level and none of them secret:
- Reclassify video calls as unauthenticated channels in every payment and credential workflow. A meeting can discuss a transfer; it cannot authorize one.
- Verify out of band, through data the requester did not supply. Hang up and call back on a number from the directory, not from the message. The FBI's guidance is blunt on this because it works.
- Use freshness challenges for urgent requests. A shared secret or a question about something not on the public record costs nothing and defeated the Ferrari attempt outright. Treat it as a stopgap, not a system: it depends on composure the attacker is working to remove.
- Put policy in the payment path. Thresholds, new-beneficiary holds, and dual control that no meeting, message, or apparent executive can waive. The control must be immune to persuasion precisely because the humans are the target.
- Make verification socially safe. This scheme weaponized hierarchy, secrecy, and urgency. An employee who challenges a CFO's request must know, in advance and in writing, that the challenge is the job, not an insult to it.
- For the approvals that matter most, move presence and authorization onto systems built for them, and insist that every high-value approval leaves a sealed, tamper-evident record of who approved what, on what evidence, when.
The division of labor in that last rule (prove presence at the moment it is claimed; gate and seal authority at the moment it is exercised) is the right one to keep. The presence half is the work of liveness verification, the field our explainer describes; the authorization half is where our own work sits: the Sigilith Platform checks consequential approvals against policy and seals each one as a tamper-evident record. The Arup case is what the absence of both looks like, written in fifteen wire transfers.
A finance employee in Hong Kong sat in a meeting, surrounded by familiar faces, and every one of them was a rendering. The screen did its job perfectly. The controls that were supposed to stand behind the screen did not exist. That, and not the sophistication of the forgery, is the finding that generalizes.
Sources
Case reporting
- Hong Kong Free Press, Multinational loses HK$200 million to deepfake video conference scam, Hong Kong police say, 5 February 2024
- CNN, Finance worker pays out $25 million after video call with deepfake "chief financial officer", 4 February 2024
- South China Morning Post, "Everyone looked real": multinational firm's Hong Kong office loses HK$200 million after scammers stage deepfake video meeting, 4 February 2024
- CNN, Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee, 16 May 2024
- South China Morning Post, UK multinational Arup confirmed as victim of HK$200 million deepfake scam that used digital version of CFO to dupe Hong Kong employee, 17 May 2024
Regulator and law-enforcement warnings
- FinCEN, Alert FIN-2024-Alert004: Fraud Schemes Involving Deepfake Media Targeting Financial Institutions, 13 November 2024
- FBI Internet Crime Complaint Center, PSA I-120324-PSA: Criminals Use Generative Artificial Intelligence to Facilitate Financial Fraud, 3 December 2024
- Deloitte Center for Financial Services, Generative AI is expected to magnify the risk of deepfakes and other fraud in banking, 2024
Related incidents
- Wall Street Journal, Fraudsters Used AI to Mimic CEO's Voice in Unusual Cybercrime Case, 30 August 2019
- Forbes, Fraudsters Cloned Company Director's Voice in $35 Million Heist, Police Find, 14 October 2021
- The Guardian, CEO of world's biggest ad firm targeted by deepfake scam, 10 May 2024
- Bloomberg, Ferrari Narrowly Dodges Deepfake Scam Simulating Deal-Hungry CEO, 26 July 2024
Related Sigilith analysis
Also Applicable To