Volume 8,Issue 8
This study explores the development of an automated audio description (AD) framework for local cultural promotional videos using a human-machine collaborative approach. The proposed framework integrates a multimodal large language model, Doubao, with human expertise to enhance AD production, particularly for videos featuring culturally rich content. By focusing on the example of the Fujian-based video “Where There Are Dreams, There Is Fu”, the study addresses two primary challenges in AD: cross-frame coherence and accurate cultural symbol interpretation. Through iterative human-machine collaboration, the model generates coherent, culturally grounded AD scripts that align with the cognitive patterns of visually impaired audiences. This research highlights the potential of GenAI-driven solutions in creating accessible content for public welfare organizations while maintaining cultural authenticity. The proposed framework offers a scalable, cost-effective approach to improving accessibility and promoting cultural heritage for visually impaired individuals.