Attack technique

Deepfake

AI-generated synthetic audio, video, or images that convincingly imitate a real person's voice or likeness.

Definition

A deepfake is synthetic media, typically audio, video, or images, generated by artificial intelligence to convincingly mimic a real person's voice, face, or mannerisms. Attackers use deepfakes to impersonate executives, colleagues, or public figures in fraud schemes and social engineering attacks.

Deepfake technology has advanced to the point where a short sample of a person's voice, sometimes taken from a public interview or social media video, is enough to generate a convincing synthetic voice call. Attackers use this to impersonate a CEO or finance director on a phone call, instructing an employee to urgently transfer funds or share confidential information, a tactic often layered on top of business email compromise. Video deepfakes have been used in fraudulent video calls where an impersonated executive appears to join a meeting and approve a transaction in real time.

Deepfakes matter because they attack one of the last forms of verification people traditionally trusted, hearing or seeing someone directly, and the technology to create them is now inexpensive and widely accessible. Finance and treasury teams are prime targets because a single convincing call can authorize a large wire transfer before anyone questions its legitimacy. As synthetic media becomes harder to distinguish from genuine recordings, reliance on voice or video alone as proof of identity becomes increasingly risky, including in the Indonesian corporate environment where phone-based approvals remain common.

Organizations should require independent, out of band verification for any high value or unusual request, such as calling back a known number rather than trusting the number or video call that initiated the request. Establishing a verbal code word or secondary approval step for wire transfers and sensitive data requests adds a safeguard that synthetic media cannot bypass. Employees should also be trained to recognize subtle deepfake artifacts, such as unnatural blinking, audio latency, or inconsistent lighting, while understanding that detection alone should never replace verification procedures.

At a glance

Severity
High
Prevalence
Emerging, growing quickly
Primary targets
Finance and treasury teams, executives, and their close contacts
Also known as
Synthetic media fraud

How it works

  1. 1

    Voice or likeness sampling: the attacker gathers audio or video samples of the target, often from public interviews or social media.

  2. 2

    Generation: AI models use those samples to create a convincing synthetic voice, video, or image.

  3. 3

    Impersonation: the attacker uses the deepfake to pose as an executive, colleague, or trusted contact in a call or video meeting.

  4. 4

    Pressure: the fabricated interaction instructs the victim to urgently transfer funds or share confidential information.

  5. 5

    Fraud completed: the victim, trusting the familiar voice or face, complies before verifying through an independent channel.

Warning signs

  • An urgent request for a wire transfer or sensitive data over a call or video meeting
  • Slight audio delay, unnatural blinking, or inconsistent lighting on a video call
  • A request to bypass normal approval steps because of urgency or secrecy
  • Voice or video that feels almost right but slightly off in tone or pacing
  • Pressure to act immediately without time for independent verification

How to defend

  • Require independent, out-of-band verification for any high-value or unusual request
  • Call back using a known number rather than trusting the number or call that initiated the request
  • Establish a verbal code word or secondary approval step for wire transfers and sensitive data
  • Train staff to recognize deepfake artifacts, while never relying on detection alone
  • Apply dual approval for financial transfers regardless of who appears to request them

Real-world example

A finance manager at an Indonesian conglomerate receives a video call from someone who looks and sounds exactly like the company's CFO, urgently requesting an emergency wire transfer to close a confidential deal. The manager pauses, calls the CFO's known number directly, and discovers the real CFO never made the request.

How Claro helps

Claro's awareness training covers deepfake-enabled fraud scenarios so employees understand why voice and video alone are no longer sufficient proof of identity for sensitive requests.

Frequently asked questions

How is a deepfake attack different from vishing?

Vishing relies on a live scammer using their own voice and a false pretext over the phone. A deepfake attack uses AI-generated audio or video to convincingly mimic a specific real person, making impersonation far harder to detect by ear or eye alone.

Why have deepfake scams become more dangerous recently?

The technology needed to generate a convincing voice or video clone has become inexpensive and widely accessible, and only a short public sample of someone's voice or face is now enough to produce a believable fake.

Can you detect a deepfake just by watching or listening carefully?

Sometimes subtle signs like unnatural blinking or audio latency are present, but detection is unreliable as the technology improves. Independent verification through a separate channel is a far more dependable defense.

Does a verbal code word actually stop deepfake fraud?

Yes, when used consistently. A prearranged code word or secondary approval step confirms identity in a way that synthetic audio or video cannot replicate, since the attacker does not know it in advance.

Reduce your human risk

Claro measures and lowers the risk these terms describe, in English and Bahasa Indonesia.

Request a demo