Deepfake fraud is the use of AI-generated voice or video to impersonate a real person, usually to deceive an employee into transferring money, sharing data, or approving a request. It is social engineering with the hardest part automated. Where a fraudster once relied on a convincing email, they can now clone an executive's voice from a few seconds of public audio, or generate a moving video of a leader on a call, and use it to make a fraudulent request feel legitimate. The request itself is old. What is new is that the voice on the phone and the face on the video call can now be fabricated well enough to pass a quick human check.
For a regulated company, this is a direct operational and financial risk, not a novelty. Publicly reported cases have included finance staff transferring large sums after video calls with what appeared to be their own executives, but were entirely synthetic. A deepfake payment fraud sits inside the same risk categories your auditors already care about: payment controls, fraud prevention, and the operational resilience expectations under DORA and NIS2. The attack does not breach a firewall. It breaks the human assumption that a familiar voice or face proves identity, and it does so at a scale and quality that did not exist a few years ago.1
What deepfakes are, and why they supercharge fraud
A deepfake is synthetic media, audio, video, or both, generated by AI to convincingly imitate a real person. Voice cloning models reproduce a target's timbre, accent, and cadence from short samples. Video models reconstruct a face and can drive it in real time, so an impersonator can hold a live conversation while wearing someone else's appearance. The output is no longer the obviously glitchy footage of a few years ago. It is good enough to survive a glance on a video call and a hurried phone conversation.
What makes this dangerous for fraud is the collapse of two old constraints. Convincing impersonation used to require either physical presence or a skilled human mimic, both of which were rare and expensive. AI removes that limit. Voice cloning now needs only seconds of sample audio, which is readily available for any executive who has spoken on a podcast, a webinar, an earnings call, or a conference stage. The raw material is public, the tooling is cheap, and the impersonation scales. That turns the convincing email of a few years ago into a convincing phone call or video meeting, against the same targets and for the same goals.
How deepfakes are used against companies
Deepfake fraud targets the points where a person can move money or release data on trust. The most common patterns map directly onto fraud your finance and operations teams already know, with the impersonation upgraded.
- CEO and executive impersonation on video calls, where a synthetic version of a senior leader joins a meeting, often with other fabricated participants, to pressure a finance employee into an urgent confidential transfer.
- Voice cloning for payment fraud, where a cloned executive or supplier voice calls to authorize a wire, change bank details, or push through a payment that bypasses the normal approval path.
- Vishing at scale, where AI-generated voice extends classic phone-based social engineering, impersonating IT support, a bank, or a colleague to extract credentials, one-time codes, or approvals.
- Supplier and vendor impersonation, where a cloned voice or video of a known contact requests a change to payment instructions, a variant of business email compromise with a far more convincing channel.
- Recruitment and onboarding fraud, where synthetic candidates pass remote video interviews to gain insider access, an emerging risk for distributed teams.
The shared arc is reconnaissance, synthesis, pretext, and execution. The attacker gathers public audio, video, and organizational detail, generates the cloned voice or face, makes an urgent and confidential request through a trusted-seeming channel, then relies on time pressure to get the action completed before anyone verifies it independently. This is the same class of attack we cover in business email compromise and phishing and social engineering, with the impersonation moved from text to live voice and video.
Deepfake fraud does not break your technology. It breaks the assumption that a familiar voice or face proves who you are talking to. The defense has to live in process, not perception.
Why deepfakes defeat trust-based controls
Most organizations run informal controls built on personal recognition. A finance officer will release an unusual payment because the CFO called and the voice was unmistakable. A help desk will reset access because the caller sounded like the person on the account. These controls were never written down as controls. They are habits that worked because faking a voice or face convincingly was hard. Deepfakes remove that assumption, and the informal control collapses with it.
The failure is structural, not a lapse of attention. Recognition controls authenticate the wrong thing. They verify that something sounds or looks like a known person, when what you actually need to verify is that the request is genuine and authorized. A deepfake satisfies the recognition test perfectly while failing the authorization test completely. Any control that rests on a human deciding whether a voice or face is real has been quietly broken, and the attacker is counting on the organization not noticing until after the money has moved.
| Control | What it tests | Holds against a deepfake? |
|---|---|---|
| Familiar voice on a phone call | Whether the voice sounds like a known person | No, voice cloning is built to pass this |
| Familiar face on a video call | Whether the face looks like a known person | No, real-time video deepfakes pass this |
| Callback to a number on file | Whether you reach the real person on a known channel | Yes, the attacker does not control that channel |
| Pre-shared code word | Whether the requester knows a secret a deepfake cannot generate | Yes, unless the secret has leaked |
| Mandatory second approver | Whether an independent person also authorizes the action | Yes, no single impersonated request can complete it |
How to detect a deepfake, and the limits of detection
Detection has a role, but it is the weakest layer and you should treat it that way. In a live interaction there are sometimes cues: unnatural blinking or lip sync, lighting that does not match the room, audio with no breath or background noise, a refusal to perform a simple unscripted action such as turning the head or holding up a hand. Context matters more than any single cue. An unusual request, unusual urgency, an unusual channel, secrecy, and pressure to skip the normal process are the real warning signs, and they hold whether or not you can spot a visual artifact.
The structural limit is that generation and detection are locked in an arms race, and detection lags. Automated deepfake-detection tools improve, then the next generation of synthesis defeats them. Worse, detection puts the burden on a human in the moment, expecting an employee on a live call to spot a fake under pressure from someone who appears to be their boss. That is not a control you can rely on or audit. The durable defenses do not try to detect the fake at all. They verify the request through a path that does not depend on perception.
Layered defense against deepfake fraud
Because the attack targets trust and process, the strongest defenses are procedural and organizational. The principle is simple: any sensitive action must be confirmed through a channel the attacker does not control, and no single convincing request can authorize it alone. Build the controls so that even a perfect deepfake hits a step it cannot pass.
- Require out-of-band verification for sensitive actions, confirming any payment or data request through a separate known channel, such as a callback to a number already on file rather than one the caller supplies.
- Set strict, mandatory procedures for payments and changes to bank details, so a single request, however convincing, can never authorize them alone.
- Use agreed verification methods for high-risk requests, such as a pre-shared code word or a required second approver, that a deepfake cannot supply.
- Train employees specifically on deepfake fraud, so they expect convincing impersonation and treat the verification steps as mandatory, not optional courtesies.
- Limit unnecessary public exposure of executive voice and video where practical, since every public clip is raw material for cloning.
- Build a culture where verifying a request is never treated as an insult, removing the social pressure that the fraud relies on.
- 01Pause on the triggerTreat urgency, secrecy, an unusual channel, or a request to change payment details as a stop signal, not a reason to move faster.
- 02Verify out of bandConfirm the request through a separate, known channel: call back a number already on record, or message the person on a trusted platform you initiated.
- 03Apply the agreed checkUse the pre-arranged verification method for the request type, such as a code word, a callback procedure, or a mandatory second approver.
- 04Hold the payment controlDo not release funds or data until verification completes. No live call, voice, or video overrides the documented payment process.
- 05Report and captureReport the attempt to security regardless of outcome, and preserve any recording or detail so the pattern can be tracked and staff warned.
These controls overlap heavily with defenses against phishing and business email compromise, because deepfake fraud is the same class of attack with a more convincing impersonation. The most effective way to make them stick is to test them. We rehearse exactly these scenarios through the social engineering work in our security awareness training, so the verification steps are practiced under pressure rather than read once in a policy.
The regulatory angle: EU AI Act and GDPR
Two EU regimes touch deepfakes directly. The EU AI Act, Regulation (EU) 2024/1689, sets transparency obligations for AI-generated and manipulated content. Providers must mark synthetic audio, image, and video in a machine-readable way, and deployers who create deepfakes must disclose that the content is artificially generated, with limited exceptions.2 These rules raise the legal cost of producing and distributing deepfakes, but they do not stop a criminal who is already committing fraud, so transparency law is a backstop, not a defense for your finance team. We explain the framework in the EU AI Act explained and on the EU AI Act compliance page.
GDPR applies because a person's voice and face are personal data, and a cloned likeness processes that data without a lawful basis or consent. For your own organization, GDPR is also a reason to manage how much executive voice and video you publish and how you handle any recordings captured during an attack. If your firm builds or deploys AI systems of its own, the same governance discipline applies internally, which is the subject of our work on securing LLM applications and our broader AI security service.
How Raptoric helps
Raptoric treats deepfake fraud as what it is: social engineering with the impersonation automated. We help you build and rehearse the controls that hold regardless of how convincing the fake is, focusing on out-of-band verification, payment-approval procedures, and the human judgment that turns a suspicious call into a reported incident rather than a wire transfer. The fastest place to start is testing your people and processes through our security awareness training, and if you are also deploying AI internally, our AI security service covers the governance side. To scope either, book a scoping call and a senior engineer will walk through your exposure and the controls that fit your organization.
