Skip to main content

‘I am so excited about AI, I feel like shouting from the rooftops. But I remain critical.’

23 September 2024
Illustration in Scribit.Pro of two people drinking coffee.

An interview with Ferry Molenaar, who, together with Scribit.Pro, created an AI-generated audio description voice of himself.

Research: AI-generated voice

Following in the footsteps of the supermarket Aldi, Scribit.Pro also developed a voice generated by AI (artificial intelligence). This spring, the retailer replaced regular voice actor Diederik Ebbinge with a voice created using artificial intelligence. Ten Aldi employees were models for this voice. We at Scribit.Pro took up the challenge and developed an artificial voice based on sound recordings of our regular business partner Ferry Molenaar - also a voice actor. The aim: to investigate whether AI can improve the production of accessible video. Can Scribit.Pro develop an artificial voice-over that provides video content with a lifelike, natural-sounding audio description?

Voice actor Ferry Molenaar

Time to talk to Ferry. As a sound professional and enthusiastic user of media and technology, what are his views on artificial intelligence? And does his visual impairment play any role? What role does he think AI can play in video accessibility?

Ferry Molenaar (34) is involved in recording and publishing sound. He creates podcasts and audio tours and is available as a voice-over. Ferry is blind; due to the eye condition Aniridia, he lost his sight when he was 23. Scribit.Pro and Ferry work together regularly. He is an enthusiastic user of audio description. As a blind creative entrepreneur, he also regularly applies his expertise.

Our partnership has already led to many great things, such as De Beeldspieker, the most accessible podcast in the Netherlands. In this podcast, Ferry visits Dutch cultural institutions and he and his listeners get an auditory experience of a collection or exhibition.

For the latest collaboration, Scribit.Pro and Ferry investigated whether it was possible to create an AI-generated version of his voice that could serve as audio description.

Interview with Ferry Molenaar

You are a frequent user of audio description. To what extent does an audio description influence how you experience the product in question?

‘Audio description has a huge influence on my experience. I am a premium member of YouTube, I watch so much. In general, I watch videos where a lot of talking is done because I can easily follow them. For example, I am a fan of Mattias Krantz, who destroys pianos - in the most creative ways. The vlogger replaces the hammers of a piano with real ones, or makes a piano waterproof, fills it with water, and then sees how the instrument sounds. He tunes an entire piano to the same note and then calls a teacher for piano lessons, but all the keys sound the same. Hilarious. But I would love to see the face of that piano teacher. I already enjoy this content a lot, but if an image description could mention such details, it would be even more fun. That is the gap that audio description fills for me.

Language use, voice sound, and intonation of an audio description are also important for an experience. I believe that the image description should match the product. If a video becomes exciting or sarcastic, the audio description can reflect that, but in a fitting manner. You add something to the product, and you can make or break the story through it. I find this especially important for online videos because a YouTube video can go viral over a small detail. When I miss that detail, there is little enjoyment for me in such a video.

‘With the synthetic Xander, I understand everything with just half a word.’

Up until a year ago, I was very satisfied with synthetic voices. These are the voices I hear daily in reading software and the voice-over function on my phone, but they are also the voices that Scribit.Pro uses for audio description. The advantage: give such a voice the same sentence a hundred times and it will sound the same a hundred times. Another advantage is that I now know these voices so well that I can make them speak faster and faster. It becomes barely understandable, but as a blind user, you still know what is being said. For example, with Xander (a commonly used synthetic voice), half a word is enough for me, haha. No doubt there are users who have the same experience with the voices from Scribit.Pro. However, I fear, in the case of Mattias Krantz's videos, that a standard synthetic audio description would ruin the entire video.’

Time for a human version of artificial voices, then. For which, ironically, we have to enlist the help of artificial intelligence. How do you feel about AI? To what extent is that attitude shaped by your work as a sound professional and how much is your visual impairment a factor?

‘I would love to shout from the rooftops, that's how excited I am about AI. I remain very cautious as well. I'm incredibly enthusiastic and the developments are already very advanced, but I don’t want to be swept away by it. It can help us, but no more than that. It's not going to take over the world. Artificial intelligence is really just predictive thinking. And right now, so many things are still going wrong. That's why I want to keep control. It's important that we as humans stand between AI and reality. As long as we stay in the director's chair and can tell AI systems that they're doing something wrong or need to do it over, things will be ok.

At the moment, there are about seven points where artificial intelligence takes over or improves my daily tasks. Think of writing newsletters or LinkedIn posts; chores I don't particularly enjoy and I'm not very good at. Brainstorming with ChatGPT is also incredibly fun, I always have it open. AI can also elevate the worst sound recording to studio quality. Or suggest a title for a new podcast that I would never have thought of myself. In short, it complements my shortcomings. And I always remain the final editor who decides whether I agree with AI.


I always have ChatGPT open.’


In response to the question of how much my visual impairment is a factor, I would say that this impairment is becoming less and less of a factor. My blindness is guiding AI, enabling me to do things I couldn't do before. More often, the thought is: let's just do this. Even though I am blind. AI recently helped me replace the nose wheel on my robot vacuum. It used to take me an intensive hour of googling. Now ChatGPT guided me through it. I had to repeat three times that I am blind and can't do anything with visual instructions, but eventually, I'm fussing with that thing myself and I replaced the nose wheel. And I even had fun doing it.’

Can you tell us a bit more about the AI experiment you did with Scribit.Pro?

‘Currently, Scribit.Pro uses synthetic voices for audio description of video content. The idea of the experiment was to see if an AI-generated voice based on my own voice would result in a more pleasant sounding description. We collected voice materials and input them into ElevenLabs, an AI audio generator. That AI model created a voice out of all that data. My voice, but also not. For three days, I wandered around my house in a daze, hearing myself repeatedly say: ‘It's really creepy.’ Until at one point, I also input that phrase into ElevenLabs. Then it really got creepy, haha. Two and a half years ago, I had a synthetic voice made of my own voice, but it didn’t resemble me at all. I sounded like I had swallowed a towel. But now we’ve been able to create a voice where I really hear myself. It doesn’t sound like me, it is actually me. However, it’s not so perfect that it could replace me in my work as a voice-over. That’s a huge relief. But for Scribit.Pro, it can definitely be a serious alternative. It results in a more natural and pleasant sounding audio description, at prices that can still be kept low.’

How could Scribit.Pro leverage AI in the production of video accessibility?

‘The clone made from my voice contains emotion and intonation, making it a serious competitor for synthetic audio description voices. I expect that AI will partially take over the profession of video description for Scribit.Pro. It would be interesting for Scribit.Pro to develop its own AI model that can be trained to describe videos. The AI model must also be taught to consider the context of a video: including the subject, purpose, client, and target audience in the approach. Obviously, Scribit.Pro will always remain not only a teacher and final editor, but also the gatekeeper.’

What would be a desired development within AI for you, professionally or personally? Do you have a dream in the field of artificial intelligence?

‘Frankly, my dream has already become reality: an AI assistant with which I can walk down the street, assisting me with the information I need to get from A to B. When I’m looking for the entrance to a shop, I currently have to call a volunteer to see where the door is. I expect artificial intelligence to take over this task, with real-time descriptions of the environment, for instance. I hope that AI can replace assistive technology. Such AI exists, but still needs to be rolled out. So, my biggest wish is about to come true. But when it does, its stability – and thus reliability – is of utmost importance. Many AI models still crash too often at the moment. If I ask AI to notify me when the traffic light turns green, but there’s a server overload, I can’t rely on that system. Therefore, we must never become fully dependent on AI. In the event of a global computer crash, we’d be back to the time of hunters and gatherers. I want to be able to know myself, if necessary, when the traffic light turns green. Therefore, AI assisting blind and visually impaired pedestrians doesn’t mean the municipality can get rid of the tactile traffic signals.’


‘I think – and hope – that AI can replace assistive technology.’

Let’s hope policymakers continue to recognise the importance of this.

‘Shall we professionals promise each other, here and now, to always stand between AI and reality? We can be afraid of the consequences of a world taken over by artificial intelligence, but we can also simply say: we won’t let that happen. Great. Just like that, we’ve saved the world from destruction.’

Read more about our experiment with an AI-generated audio description voice.

Read our blog about AI in audio description.

Want to learn more?

Sign up to our newsletter