Who's Afraid of the Big Bad Wolf? Voice Cloning and Modern Deception
Anyone familiar with the fairy tale of the Big Bad Wolf knows that he was able to deceive his victims with clever tricks. Today, artificial intelligence (AI) can deceive us in a similar way.
Social Engineering is at least as old as the fairy tales of the Brothers Grimm. In the story "The Wolf and the Seven Young Goats", the wolf uses various tricks to pose as the goats' mother. In the end, he manages to outsmart the goats, and they open the door for him, which has devastating consequences. This fairy tale teaches that one should not trust seemingly familiar characteristics such as appearance and voice unconditionally.
AI: The Wolf in Sheep's Clothing
AI can make work easier in many areas, but it also carries risks: on the one hand, one should never trust an AI unconditionally, as it is trained on large amounts of data from a wide variety of sources to recognize patterns and make decisions. If these data are incomplete, biased, inaccurate or faulty, the AI can draw wrong conclusions. One then receives inaccurate or incorrect answers that may sound plausible. On the other hand, new technologies enable more sophisticated attacks. Especially in the area of social engineering, attackers can now proceed more easily, faster and more specifically to deceive us, just like the Big Bad Wolf in the fairy tale. In the future, it will be more difficult to recognize phishing attacks. In the aforementioned fairy tale, the wolf pretends to be a known figure with clever tricks to win the trust of his victims. Today we speak of "Deepfakes" such as voice cloning as well as image and video manipulation. Just like the goats in the fairy tale, we should be careful about whom we trust and to whom we open the door!
We gave a presentation on this topic at the Night of Digitalization . Below you will find the video of the presentation as a shortened version.
How Artificial Intelligence Works
Put simply, an AI is based on statistical methods, probabilities and pattern recognition. Through training with large amounts of data, it learns which patterns occur in the data and how language is used. An AI language model is trained with extensive text data that encompasses a large part of documented human knowledge. This allows it to make predictions about which words are meaningful in a particular context. The difference between a chicken and a cow is recognized by the AI because it has processed numerous descriptions of both animals in the training data.

The text generated by an AI is new every time and can vary. The results are not deterministic, which means that an AI can give different answers to the same question – in contrast to a classic computer program that works with fixed instructions such as "if", "then" and "else".
New Attack Methods in Phishing and Vishing
Phishing is a popular social engineering method and can take place over various channels such as SMS, email and various messengers. With special AIs, personalized phishing attacks can be carried out automatically. These AIs, for example, scan the websites of the companies to be attacked and automatically create phishing messages using the collected information and send them to selected recipients.
In addition to written phishing attacks, there are also attacks by telephone, so-called vishing. Attackers, for example, pose as employees of a large software company in order to obtain sensitive information from their victims. Traditionally, this method required high personnel costs with a call center character. By combining text-generative and text-to-speech AI, many victims can be reached simultaneously with minimal effort.
Through adapted voice cloning and targeted vishing, methods such as the well-known "grandchild trick" become even more dangerous. The nature of the attacks is not new. However, they become much more successful through the use of voice cloning.
AI Technology for Voice Cloning
In connection with voice cloning, a speech AI model can be trained to mimic the voice of a specific person by using voice samples and learning how they speak. There are already open-source AIs with user-friendly interfaces that can be used by many people without extensive prior knowledge. Thus, attacks in this area can increase.
There are two main types of AI for voice cloning:
- Text-to-Speech (TTS)
In this method, the text that the cloned voice should read is specified. TTS technology simply reads a certain text aloud and thus generates an artificial speech output. - Inference
In this method, the trained model does not read a text file, but manipulates the voice characteristics in an existing audio recording. With real-time models, texts can be output directly while speaking with the cloned voice.

How Attackers Get Voice Samples
There are various ways in which attackers can obtain voice samples. The two most likely scenarios are social engineering and publicly available voice recordings:
- Social Engineering
Attackers can engage their target person in a conversation and record them secretly. In this way, they obtain voice samples that they can later use for cloning the voice. - Publicly Available Voice Recordings
Public recordings of speeches or social media profiles with video content on the internet can serve as a source for voice samples. This can be particularly risky for people who are in the public eye or are active on social media.
Legitimate Use Cases for Voice Cloning
In addition to the misuse of voice cloning for criminal purposes, there are also some areas in which it can be used sensibly and legally. For example here:
- Education and Learning
Through the cloning of speech, multilingual training materials can be created more easily. This could enable the availability of teaching materials in different languages and better prepare students with different backgrounds for learning. - Accessibility
Voice cloning can help people who are no longer able to use their voice due to physical impairments or accidents. With the help of a trained text-to-speech model, they can continue to communicate and speak with their own voice. - Entertainment
Voice cloning can be used to recreate the speaking roles of actors in films or video games. - Historical Presentation
Voice cloning can bring historical figures and events to life. By cloning known voices, people could better understand these figures and their life stories, as the presentation appears more authentic.
Social Engineering – Motives and Techniques
Attackers specifically exploit the psychological weaknesses of people and use various methods to achieve their goals. Their motives include financial enrichment through extortion or fraud, political and social manipulation to generate unrest in the population, identity or data theft, damage to reputation or disinformation.
Through cleverly devised tricks, attackers use one or more tactics. These are often based on the principles of persuasion psychology, as described by the well-known US psychologist Robert B. Cialdini in his book "Influence: The Psychology of Persuasion" (1984):
- Time Pressure
- Authority
- Reciprocity
- Social Proof
- Commitment and Consistency
- Liking
In 2016, Cialdini added another technique to the six: The principle of unity and shared identities.
Fraud has always worked by exploiting the human psyche. Over the years, however, the requirements have changed. Attackers constantly adapt their methods to be as successful as possible. With AIs, they have received a practical toolbox that makes it even easier for them to carry out attacks.
Detecting Fraud and Protecting Yourself
To protect yourself from fraud attempts, it is important to be vigilant and skeptical of unusual incidents. We have listed some strategies here on how you can protect yourself from social engineering:
- Ensuring Identity
- Be skeptical if unusual communication channels were chosen.
- Critically question unusual requests, even if they come from supposedly known persons.
- In case of doubt, choose a known and secure communication channel for query or verification.
- Only disclose the most necessary personal information on the internet.
- If possible, ensure that your voice is not freely available on the internet.
- Avoid passing on confidential information via insecure channels.
- Define and use secure communication channels, for example for banking transactions or other sensitive matters.
- Critical Consideration of Media Content
- Do not blindly trust images, videos and audio recordings from unverified sources. Content in social media in particular can be manipulated or faked.
- Check the source and look for further confirmations if something seems suspicious.
- Regularly inform yourself about new fraud tricks and techniques.
- Stay vigilant and look for signs of phishing, vishing or social engineering attacks.
Do you want to protect your company better against phishing, vishing and social engineering? We defend you against these dangers.
On 09.09.2024 in the category Communication Security published.
