Artificial Intelligence
Deepfakes: How They’re Made, How to Spot Them, and Why It’s Getting Harder
The short version
- A deepfake is synthetic media trained on a real person. AI models learn a face or voice from source video and audio, then swap it into words or scenes the person never did.
- The old spotting tricks are fading. Blurry edges and odd blinking still appear, but the models are improving; “look closely” is no longer a reliable defence.
- Verify through a trusted channel. Call the person back on a known number, be suspicious of urgent money requests, and treat realistic video as proof of nothing on its own.
Realistic fake video and voice are now part of everyday scams and online life, and most advice about spotting them is already out of date. This article gives a plain mental model of what a deepfake is, how one is made, and the defences that still hold. The honest headline: the visual tells are shrinking, and the durable skill is verifying the source rather than staring at the pixels.
What a deepfake actually is
A deepfake is synthetic media, usually video, audio or images, generated by AI models that have learned the patterns of a real person’s face, voice or mannerisms. The models study enough source material of the person, then generate new frames or audio that match them, putting the person into scenes or words they never did.
The word blends “deep learning” and “fake”, and that is an accurate description of the method. It is not a filter or a clever edit. It is a model that has learned how a specific person looks and sounds well enough to produce convincing new versions of them. The technology has moved fast, which is why the honest answer to “what is a deepfake” has changed even in the past few years. The definition stays the same. The quality does not.
How they’re made
The general process is simple to describe and hard to reverse. A model is trained on lots of source material of the person, interviews, social media clips, voice recordings, and learns their patterns: how they move, how their face catches the light, the rhythm of their speech. It then generates new frames or audio matching those patterns.
The part that changed the risk picture is how little source material is now needed. Modern tools can produce a convincing result from a surprisingly small amount of public audio or video, which is why the risk is widespread rather than state-level. No special equipment is needed anymore. The same generative technology that powers chatbots and image tools, explained in plain terms in the site’s guide to large language models, is what makes this possible, and the same honest note applies to the output: it can look confident and be entirely fabricated.
What they’re used for (the good and the bad)
Not all deepfakes are harmful. There are legitimate uses: dubbing a film into another language, restoring a voice for someone who has lost theirs, accessibility tools and education. The technology is a tool, and the tool has honest uses.
The harmful uses are what drive the concern. Non-consensual intimate images, which use a real person’s face without consent, are among the most damaging. Fraud is a growing category, fake calls from a cloned voice, or a video call that appears to show a chief executive authorising a payment. Disinformation uses realistic fake footage of public figures to spread false claims. And impersonation scams use a cloned voice or face to convince someone they are talking to a person they trust.
The risk is not limited to public figures. An ordinary person’s voice or face can be cloned from a small amount of public audio or video, a social media clip or a voicemail greeting is enough, and that clone can be used in a scam call to a family member or a fake video sent to a workplace. That is why the defences below matter to everyone, not only to people who appear on television.
The common thread is trust. Every harmful use depends on the victim believing the media is real, which is why the defences below focus on verification rather than pixel-peeping.
How to spot one, honestly
The honest warning comes first: the visual tells that used to give deepfakes away are becoming less reliable. The signs still exist, subtle artefacts around the face, blurring at the edges, odd blinking or mouth movement, lighting and skin that do not quite match, audio that does not sync or has an unnatural tone. But the models are improving, and a casual viewer can no longer assume the artefact will be visible. “Look closely” is not enough on its own.
The checks that still work are behavioural. Where did this video or call come from? Who sent it? Does the content match the stated source? A surprising video forwarded by an unknown account, or a call claiming to be from a bank, deserves the same scepticism as any other unsolicited approach. This is the same discipline the site’s guide to checking AI search answers applies to written answers: verify the source, do not trust the surface.
Why it’s getting harder
The reason detection is getting harder is an arms race. Generation models improve, detection tools improve, and then generation improves again. Each round closes the gap a casual viewer can see. The tells that were reliable two years ago are unreliable now, and they will be more unreliable in two more.
The consequence is a shift in where the burden sits. Spotting the artefact is a losing game, because the artefact is disappearing. Verifying the source is a stable game, because it does not depend on the quality of the fake. This is why the practical advice has moved from “look for the blur” to “check where it came from”, and why that advice will stay true even as the fakes get better.
The defences that still work
The practical list is short and human. Verify surprising requests through a known channel: if a call claims to be from someone you know, hang up and call them back on a number you already have. Treat urgency and secrecy as red flags, because both are the standard tools of a scam. Protect your own voice and face online where you can, since a small amount of public audio or video is all a cloning tool needs.
And know the reporting paths. In Australia, the eSafety Commissioner handles image-based abuse, including non-consensual deepfake images, and the reporting process is the first step to getting content taken down. For readers building a wider picture of how AI systems work, the site’s explainer on what AI agents can and cannot do and its look at whether machines can be creative cover the neighbouring questions about what this technology actually is.
Don’t believe your eyes
The technology will keep improving, and the artefacts will keep shrinking. The durable skill is not spotting the fake. It is verifying the source, and when a video or a call asks for something important, checking it like you would check any claim.
A realistic video is no longer proof that someone said something. That is a strange thing to have to accept, and it is the reality the technology has created. The defence is not suspicion of everything. It is a simple habit: for the messages that matter, the urgent ones and the surprising ones, go back to a source you trust and ask. That habit will still be working long after the blur is gone.
Sources: eSafety Commissioner (Australia), deepfakes and image-based abuse · CSIRO National Artificial Intelligence Centre, generative AI explainers · Australian Federal Police, online fraud and impersonation warnings
