Skip to main content

Update on research into the (im)possibilities of AI

18 May 2026
A robot arm and a human arm reach out to each other. In the middle of the photo, the two index fingers touch.

Since 2024, Scribit.Pro has been exploring the possibilities of artificial intelligence (AI) within audio description (AD). Can AI support the description of visual content for people with a visual impairment? And how usable are these descriptions in practice?

In our earlier research, AI showed reasonable results for documentaries, while dramatic productions revealed significant limitations. Meanwhile, AI models have continued to evolve rapidly. We therefore repeated our study using a newer version of ChatGPT and the same productions as source material.

The latest results show that AI is making progress in certain areas. Language output has become richer and more natural, repetition occurs less frequently, and specific actions and details are recognized more accurately. At the same time, interpretative errors, hallucinations, and narrative inconsistencies remain a major challenge — especially in fiction and drama.

Where human audio describers consistently take context, atmosphere, emotion, and narrative structure into account, AI still tends to focus on isolated moments within a scene. As a result, it often misses the elements that are essential for a meaningful audiovisual experience.

At the same time, we also see new opportunities. Particularly in documentaries and informational content, AI could in the future potentially support the production process of audio description as a helpful tool.

In this update of our research, we compare the results of ChatGPT-4o with the latest generation of AI models and examine where real progress can be observed and where human audio description remains essential for now.

Read the research update

Sign up to our newsletter