Real-Time Deepfakes Are Here: What Sub-40ms Face-Swaps Mean for Video-KYC and "the CEO on the Call"
For years, "deepfakes" meant a pre-rendered clip someone spent hours producing. That era is over. Real-time video models now transform a live webcam feed — swapping faces, changing backgrounds, altering appearance — at sub-40ms latency, 100 frames per second. The edit lands before the frame finishes drawing. For security leaders, this collapses the assumption underneath a lot of identity and trust controls: that a face on a live video call is who it appears to be.
What changed
A growing field of real-time video-to-video models now does to live video what filters did to selfies — except the output is photoreal and controllable by a text prompt. Point one at a webcam and it can put a different person's face on the stream, in real time, with no perceptible lag. The compute that used to take a render farm and hours now runs continuously, live, for a couple of cents per second.
The threat model shift is simple: "it was a live video call, so it must be real" is no longer a safe inference. Liveness — the thing video was supposed to prove over a photo — can now be faked live too.
Three controls this breaks
1. Video-based identity verification (KYC / onboarding)
Fintech, crypto, and increasingly regulated onboarding flows ask users to show their face on camera, sometimes turning their head or blinking to "prove liveness." Real-time face-swap defeats exactly that: the attacker performs the head-turn and blink with their own movements while the model paints a stolen identity's face on top. The motion is genuine; the identity is not. Liveness detection built on "is this a real, moving human?" answers yes — to the wrong question.
2. Executive impersonation over video (the evolved BEC)
Business email compromise moved to voice (vishing) years ago. The next step is video. A finance employee gets a Zoom or Teams call from someone who looks and sounds like the CFO, urgently authorizing a wire. In 2024 a Hong Kong finance worker was tricked into paying out ~$25M after a video call with deepfaked colleagues — and that used pre-built avatars. Real-time swaps make this cheaper, faster, and interactive: the fake CFO can now answer your follow-up questions live.
3. "Send me a video to prove it's you"
The fallback everyone reaches for when a channel feels off — "hop on a quick video" — is now a weaker signal, not a stronger one. Any out-of-band check that relies on seeing a face instead of verifying a secret is on borrowed time.
What actually still works
The defenses that survive real-time deepfakes have one thing in common: they don't trust the face. They trust something the attacker can't synthesize.
- Out-of-band verification of the request, not the person. For any high-value action (wire, credential reset, data export), confirm through a separate, pre-established channel — a callback to a known number, an approval in a system the caller doesn't control. Verify the transaction, not the video.
- Shared secrets and challenge phrases for sensitive internal calls. A code word the real CFO knows and a live face-swap doesn't.
- Cryptographic identity over visual identity. Passkeys, signed approvals, and hardware-backed authentication don't care what someone looks like on camera. Push identity-critical decisions onto rails that verify keys, not faces.
- Process friction on the exact actions fraud targets. Dual authorization and mandatory cool-off on large wires turn "the CFO told me to, live, on video" into "the CFO and a second approver confirmed through the system."
- Train for it. The Hong Kong loss happened because a real employee believed a real-looking call. Awareness that "a convincing live video is not proof" is now a required control, not a nice-to-have.
Make it stick with your team
Reading that a face can be swapped in real time is one thing; getting the people who approve money and access to internalise it is another. The lesson lands hardest when it's experiential — the same way our Sus or Legit? and Who's the Rat? exercises teach social-engineering awareness by putting you in the seat rather than lecturing. Run a tabletop where a "known" executive on a video call requests an urgent wire, and make verifying-through-a-second-channel the muscle memory.
The bottom line for CISOs
Real-time deepfakes don't require you to detect the fake — that's a losing arms race. They require you to stop using "I can see them" as an authentication factor and move identity-critical decisions onto channels that verify secrets and keys instead of faces. Update your wire-transfer and privileged-action playbooks now to assume any face on any call could be synthetic, and make sure the humans who approve money and access have been shown, not just told, how good this has already gotten.
The Hong Kong figure references widely reported 2024 coverage of a deepfake-enabled video-call fraud. Real-time model capabilities and pricing referenced reflect publicly documented specifications of current real-time video-to-video systems.
Ready to practise the decisions these articles describe?
Run a free War Room →