Skip to main content

AI Safety and Security

  • We present UrHiOdSynth, a three-language parallel S2ST dataset containing approximately 75 hours of speech across Urdu, Hindi, and Odia.
  • The corpus consists of 10,735 aligned sentence triplets, with an average utterance length of 8.45 seconds.
  • UrHiOdSynth represents the largest multi-domain resource offering aligned speech and text for S2ST in this language context.
  • Beyond speech-to-speech translation, the dataset supports tasks such as automatic speech recognition, speech-to-text translation, text-to-speech synthesis, and machine translation.

Leave a Reply

Your email address will not be published. Required fields are marked *