Early in 2024 a finance worker in Hong Kong got a message about a confidential transaction. He thought it looked like phishing.

He was right to be suspicious. He did what any well-trained employee does with a suspicious request, which is escalate it into a conversation.

So he joined a video call. The chief financial officer was there. So were several colleagues he recognized, faces and voices he had seen and heard before. The request was confirmed in front of him by people he knew.

He approved transfers totaling 200 million Hong Kong dollars, about 25.6 million US dollars.

Every person on that call was synthetic, generated from video and audio of Arup executives that was already publicly available. He was the only human present.

Hong Kong police described the case in February 2024. Arup, the engineering firm behind the Sydney Opera House, confirmed it was the victim in May.

He did not fall for it. He checked.

This is the part that should worry you, and it is usually skipped in the retelling.

The employee applied the control. He doubted the email. He sought confirmation from people rather than acting on a written instruction. That is exactly what every fraud-awareness program in the world instructs employees to do.

The control failed because of what it was made of. Confirmation by a person you recognize is a control built on perception, and perception is now a forgeable artifact.

Notice that no policy was broken. If you ran this as an incident review, there is no employee to retrain and no procedure to tighten, because the procedure was followed. The finding is that the procedure was never a control in the first place.

Arup's global chief information officer, Rob Greig, put the trend plainly afterward, saying the number and sophistication of attacks had been rising sharply.

The controls you never wrote down

Most organizations have a documented approval matrix. Underneath it sits a much larger set of informal controls that nobody has ever written down, because until recently they did not need writing down.

  • You know her voice.

  • You saw his face on the call.

  • That is how the CEO writes, she never uses that greeting.

  • The request came from someone in the meeting.

Every one of those is an authentication decision made by a human sensory system, and every one of them is now reproducible from material your executives have published on purpose. A keynote recording is a voice model. A quarterly earnings video is a face. A LinkedIn post archive is a writing style.

The more visible your leadership, the better the training data. Nothing about that is fixable through communications policy.

There is a second-order problem worth naming. These informal controls are invisible precisely because they work. Nobody audits them, nobody tests them, and nobody knows how many decisions rest on them, because they were never designed. They accumulated.

The regulator has already written this down

If you need something official to put in front of a risk committee, it exists.

On 13 November 2024 the Financial Crimes Enforcement Network, part of the US Treasury, issued alert FIN-2024-Alert004 on fraud schemes involving deepfake media targeting financial institutions. Deepfake media, in their definition, is synthetic content generated with AI to create realistic but inauthentic video, images, audio and text.

Their observation is the useful part. Beginning in 2023 and continuing into 2024, FinCEN saw an increase in suspicious activity reporting from financial institutions describing suspected deepfake media in fraud schemes against those institutions and their customers.

They also asked filers to tag these reports with the key term FIN-2024-DEEPFAKEFRAUD. A regulator does not create a tracking label for a hypothetical. Labels exist so volume can be counted, and volume gets counted when somebody expects it to grow.

The arithmetic points one direction

Deloitte's Center for Financial Services put a number on the trajectory in May 2024. They predict generative AI could enable fraud losses in the United States to reach 40 billion dollars by 2027, from 12.3 billion in 2023. That is a compound annual growth rate of 32 percent.

Treat the projection as a projection. The direction is the part that does not depend on the model being right.

The economics are simple enough to reason about without a forecast. Producing a convincing synthetic executive used to require a studio and a specialist. It now requires public footage and commodity software. When the cost of an attack collapses and the payout stays in the millions, volume follows.

Consider what the Arup attackers had to build. Not malware. Not an exploit. A meeting. The technical work was assembling a plausible room of people from footage a communications team had published deliberately, and the operational work was picking one employee and one plausible reason for secrecy.

Authentication is the wrong frame

The instinct is to detect the fake. Better tooling, liveness checks, watermarks. FinCEN itself points to live verification and phishing-resistant multi-factor authentication, and both are worth having.

But detection is a race you are structurally positioned to lose. The generator improves continuously and it only has to win once, against a human who is on their eleventh call of the day.

The more durable move is to stop letting perception authorize anything.

That sounds abstract until you make it concrete. The question is not whether the person on the call was really the CFO. The question is whether a transfer of that size can complete on the authority of a call at all.

What actually stops this

The controls that survive a perfect deepfake are the ones that do not care who was on the call.

Verification out of band, on a channel the requester did not choose. A callback to the number in the vendor master file, not the number in the email, not the number offered during the meeting.

Separation of the request from the release, with the second person reached independently. Two people on the same call is one channel with two witnesses.

A hard rule for changes to payment details, with a waiting period and a confirmation to a previously known contact. Most large frauds of this shape involve a change of destination, and destination changes are the narrowest place to put a gate.

Limits that bind regardless of who approves. If a single approval can move eight figures, the control is the approver's judgment, and judgment is what was attacked.

None of that is AI. It is treasury discipline that many organizations relaxed during a decade when speed was the priority and impersonation was hard.

There is a cost and it should be stated honestly. Every one of those controls adds friction to a legitimate payment, and somebody senior will be inconvenienced by them within the first week. That is the trade. The alternative is a process whose final safeguard is whether a tired person recognized a face.

The test to run this week

Try to defraud yourself, on paper.

Pick your largest routine payment. Write down every step between a request appearing and money leaving, and mark each step according to whether it depends on somebody recognizing somebody, or on a check that would still work if every human involved were a convincing fake.

Then count the recognition steps. Those are your exposure, and they are probably not in your controls documentation, because nobody thought to write down that we know each other.

The last question is the one to take to your board. If a call tomorrow showed your chief executive, sounded like your chief executive, and was joined by three colleagues your team knows by name, what in your process would still say no?

Sources