Bayesian Alignment Prior for Connectionist Temporal Classification Loss
Current methods to align speech audio with a transcription for both training speech recognition models and producing timestamps are quadratic in complexity. These methods work by considering all possible alignments between the audio and text, which scales poorly for long input sequences. This complexity becomes a limiting factor when processing downstream tasks, like forced alignment, which produces timestamps for a transcription in an input audio. We first propose a method for performing the currently quadratic alignment in linear memory, while only doubling the runtime complexity. We also note that in cases when silence is removed and speaking rate is consistent, the true alignment is a mostly straight line across time and transcription. We are experimenting with adding a Bayesian prior of phoneme length, scaled to speaking speed to the training loss to also constrict the alignment space that is used during training. We expect that doing so will improve the runtime efficiency of the algorithm and potentially stabilize training early on.
BYU Multilingual Corpus
We introduce the BYU Multilingual Corpus (BYU-MC), a parallel text dataset with English source segments translated into 98 languages, including many low-resource languages. Sourced from translation memories of The Church of Jesus Christ of Latter-day Saints, this corpus enables data creation for 9,702 translation directions.
When Scripts Diverge: Strengthening Low-Resource Neural Machine Translation Through Phonetic Cross-Lingual Transfer
We introduce a phonetic-based approach to enhance translation between closely related languages with different writing systems. By integrating IPA representations into Multilingual Neural Machine Translation models, we strengthen knowledge transfer from high-resource to low-resource languages without requiring a shared orthography.
Transfer Learning to Improve Translation of Low-resource Austronesian Languages
Translation of low-resource human languages, like Tok Pisin, is an open problem. The limited amount of training data limits the quality of translation models created with that data. However, translation performance for a low-resource language can be improved by also including data from higher-resource languages in a multilingual translation model. We explore this transfer learning effect by constructing multilingual models with data from higher-resource languages that have influenced or are similar to the low-resource language. We are currently exploring which combination of languages in a multilingual model serves to provide the best performance when translating from Tok Pisin to English.
MT Eval
Most automatic metrics for evaluating machine translation performance have been found to correlate little with human judgement of translation quality. While systems like COMET work to address the issue, we note the potential benefit of conducting manual evaluation of translation quality. We provide a simple tool, named MT Eval, for students and researchers to conduct ranked-choice evaluation to compare translation systems. We plan to continue expanding the tool's functionality and user experience.
Pathsay: Collecting Audio Data for African Languages from BYU Pathway Students
Aligned audio data is scarce for many languages, but especially for low-resource languages. In order to build models that perform translation, text-to-speech, or speech-to-text, significant quantities of aligned voice and text samples are required, but for many African languages this kind of data is not currently available. To help rectify this, we have developed a progressive web application that allows BYU Pathway students in Africa to record sentences in their native languages to create large audio dataset of African languages. This data will be used to train AI systems to break down language barriers. We intend to expand this project to other Pathway students globally in the future.
The Effects of Pretraining in Video-Guided Machine Translation
In this project, we explore how pretraining with a lexically rich dataset improves Video-Guided Machine Translation (VMT) models. By leveraging the newly introduced MAD dataset, we evaluate different video encoder architectures and assess the impact of transfer learning from video and text sources.
Semitic Root Encoding: Evaluating the Effect of Tokenizing Templatic Morphology on Machine Translation
In this study, we propose Semitic Root Encoding (SRE), a sub-word segmentation method designed to improve machine translation of Semitic languages by explicitly modeling their root-and-template morphology. We evaluate SRE’s impact on translation quality, generalization, and error generation, focusing on English-to-Arabic translation.
Preservation of Taiwanese Indigenous Languages through NMT
In this project, we work to preserve Taiwan’s endangered Indigenous languages by creating a parallel corpus for Amis, Atayal, Bunun, Paiwan, and Kanakanafu. We plan to fine-tune Meta’s NLLB-200 model on this dataset to demonstrate its effectiveness in training machine translation models for low-resource languages.
ToAll
Live interpretation of public events is a common occurance in our global community. As it is, the burden on interpretors can be significantly reduced by providing them with a pretranslated transcription. We have created an application, named ToAll, which provides automatic live interpretation by following such a transcription. The system detects the speaker's location in the transcription and uses speech-to-text to read out the prepared translation at that point. In cases where a translation has not been prepared in advance or when the speaker goes off script, machine translation is used to provide on-the-fly interpretation. This system provides greater control and reduced errors when compared to direct speech-to-speech translation, while still freeing interpretors to focus on other tasks. We are collaborating with BYU Speeches to provide access to this system at campus events to better support our international student body.