FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator
This addresses the need for automated, domain-specific speech synthesis with acoustic environments, though it appears incremental as it builds on existing TTS and ASR methods.
The paper tackles generating Arabic speech with background crowd noise for football commentary, achieving a system that can produce such speech using only 15 minutes of commentator recording.
This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelization, followed by the commentator's speech synthesizer. Our pipeline included Arabic automatic speech recognition for data labeling, CTC segmentation, transcription vowelization to match speech, and fine-tuning the TTS. Our system is capable of generating speech with its acoustic environment within limited 15 minutes of football commentator recording. Our prototype is generalizable and can be easily applied to different domains and languages.