We present UrHiOdSynth, a three-language parallel S2ST dataset containing approximately 75 hours of speech across Urdu, Hindi, and Odia.
The corpus consists of 10,735 aligned sentence triplets, with an average utterance length of 8.45 seconds.
UrHiOdSynth represents the largest multi-domain resource offering aligned speech and text for S2ST in this language context.
Beyond speech-to-speech translation, the dataset supports tasks such as automatic speech recognition, speech-to-text translation, text-to-speech synthesis, and machine translation.