Fetching the paper…
Reading the bibliography…
What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200 languages, unified speech-to-speech translation models have yet to achieve similar strides.
Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu · 1951
Earlier work this paper cites.
The shifting relationships between speech and writing
Peter Elbow · 1985
Earlier work this paper cites.
Task requirements and media choice in collaborative writing
Robert Kraut, Jolene Galegher, Robert Fish, and Barbara Chalfonte · 1992
Earlier work this paper cites.
The relation of speech to reading and writing
Alvin M Liberman · 1992
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Janus-iii: speech-to-speech translation in multiple languages
Alon Lavie, Alexander H. Waibel, Lori S. Levin, Michael Finke, Donna Gates, Marsal Gavaldà, Torsten Zeppenfeld, and Puming Zhan · 1997
Earlier work this paper cites.
Language anxiety: Differentiating writing and speaking components
Yuh-show Cheng, Elaine K Horwitz, and Diane L Schallert · 1999
Earlier work this paper cites.
Morphological productivity across speech and writing
Ingo Plag, Christiane Dalton-Puffer, and Harald Baayen · 1999
Earlier work this paper cites.
Verbmobil: Foundations of speech-to-speech translation
Wolfgang Wahlster · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
The atr multilingual speech-to-speech translation system
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, Jin-Song Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto · 2005
Earlier work this paper cites.
The polyglot internet, October 2008
Ethan Zuckerman · 2008
Earlier work this paper cites.
Ethnologue: Languages of the World
M. Paul Lewis, editor · 2009
Earlier work this paper cites.
Language recognition via i-vectors and dimensionality reduction
Najim Dehak, Pedro A Torres-Carrasquillo, Douglas Reynolds, and Reda Dehak · 2011
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Sentence-based sentiment analysis for expressive text-to-speech
Alexandre Trilla and Francesc Alias · 2012
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel · 2013
Earlier work this paper cites.
Automatic language identification using deep neural networks
Ignacio Lopez-Moreno, Javier Gonzalez-Dominguez, Oldrich Plchot, David Martinez, Joaquin Gonzalez-Rodriguez, and Pedro Moreno · 2014
Earlier work this paper cites.
An end-to-end approach to language identification in short utterances using convolutional neural networks
Alicia Lozano-Diez, Ruben Zazo-Candil, Javier Gonzalez-Dominguez, Doroteo T Toledano, and Joaquin Gonzalez-Rodriguez · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
David Snyder, Guoguo Chen, and Daniel Povey · 2015
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Berard, Olivier Pietquin, Christophe Servan, and Laurent Besacier · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
The United Nations parallel corpus v1.0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen · 2016
Earlier work this paper cites.
Bidirectional modelling for short duration language identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah, and Julien Epps · 2017
Earlier work this paper cites.
The humanizing voice: Speech reveals, and text conceals, a more thoughtful mind in the midst of disagreement
Juliana Schroeder, Michael Kardas, and Nicholas Epley · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deliberation networks: Sequence generation beyond one-pass decoding
Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu · 2017
Earlier work this paper cites.
Tied multitask learning for neural speech translation
Antonios Anastasopoulos and David Chiang · 2018
Earlier work this paper cites.
End-to-end automatic speech translation of audiobooks
Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin · 2018
Earlier work this paper cites.
IIITH-ILSC Speech Database for Indain Language Identification
Ravi Kumar Vuddagiri, Krishna Gurugubelli, Priyam Jain, Hari Krishna Vydana, and Anil Kumar Vuppala · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Earlier work this paper cites.
Feature representation of short utterances based on knowledge distillation for spoken language identification
Peng Shen, Xugang Lu, Sheng Li, and Hisashi Kawai · 2018
Earlier work this paper cites.
Spoken language recognition using x-vectors
David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur · 2018
Earlier work this paper cites.
A comparative study on end-to-end speech to text translation
Parnia Bahar, Tobias Bieschke, and Hermann Ney · 2019
Earlier work this paper cites.
Utterance-level end-to-end language identification using attention-based cnn-blstm
Weicheng Cai, Danwei Cai, Shen Huang, and Ming Li · 2019
Earlier work this paper cites.
An analysis of gender bias studies in natural language processing
M.R. Costa-jussà · 2019
Earlier work this paper cites.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2019
Earlier work this paper cites.
One-to-many multilingual end-to-end speech translation
Mattia Antonino Di Gangi, Matteo Negri, and Marco Turchi · 2019
Earlier work this paper cites.
Multimodal language processing in human communication
Judith Holler and Stephen C Levinson · 2019
Earlier work this paper cites.
Multilingual end-to-end speech translation
Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe · 2019
Earlier work this paper cites.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Earlier work this paper cites.
End-to-end multi-task learning with attention
Shikun Liu, Edward Johns, and Andrew J Davison · 2019
Earlier work this paper cites.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham · 2019
Earlier work this paper cites.
Attentive single-tasking of multiple tasks
Kevis-Kokitsi Maninis, Ilija Radosavovic, and Iasonas Kokkinos · 2019
Earlier work this paper cites.
A new time-frequency attention mechanism for tdnn and cnn-lstm-tdnn, with application to language identification
Xiaoxiao Miao, Ian McLoughlin, and Yonghong Yan · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Two-Pass End-to-End Speech Recognition
Tara N. Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, Ian McGraw, and Chung-Cheng Chiu · 2019
Earlier work this paper cites.
Interactive learning of teacher-student model for short utterance spoken language identification
Peng Shen, Xugang Lu, Sheng Li, and Hisashi Kawai · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Towards end-to-end speech-to-text translation with two-pass decoding
Tzu-Wei Sung, Jun-You Liu, Hung-yi Lee, and Lin-shan Lee · 2019
Earlier work this paper cites.
Speech-to-speech translation between untranscribed unknown languages
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2019
Earlier work this paper cites.
Tuplemax loss for language identification
Li Wan, Prashant Sridhar, Yang Yu, Quan Wang, and Ignacio Lopez Moreno · 2019
Earlier work this paper cites.
Improving multilingual sentence embedding using bi-directional dual encoder with additive margin softmax
Yinfei Yang, Gustavo Hernandez Abrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel Cer, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil · 2019
Earlier work this paper cites.
Shallow-Fusion End-to-End Contextual Biasing
Ding Zhao, Tara N. Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang · 2019
Cited alongside, same era.
FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN
Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, Fei Huang, Kevin Knight, Xutai Ma, Ajay Nagesh, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Xing Shi, Sebastian Stüker, Marco Turchi, Alexander Waibel, and Changhan Wang · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Voice based e-mail for the visually impaired
Aishwarya Belekar, Shivani Sunka, Neha Bhawar, and Sudhir Bagade · 2020
Cited alongside, same era.
Gender in danger? evaluating speech translation technology on the MuST-SHE corpus
Luisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia A. Di Gangi, Roldano Cattoni, and Marco Turchi · 2020
Cited alongside, same era.
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna · 2022
Later among the works it cites.
Occgen: Selection of real-world multilingual parallel data balanced in gender within occupations
Marta Costa-jussà, Christine Basta, Oriol Domingo, and André Rubungo · 2022
Later among the works it cites.
Evaluating gender bias in speech translation
Marta R. Costa-jussà, Christine Basta, and Gerard I. Gállego · 2022
Later among the works it cites.
Interpreting gender bias in neural machine translation: Multilingual architecture matters
Marta R. Costa-jussà, Carlos Escolano, Christine Basta, Javier Ferrando, Roser Batlle, and Ksenia Kharitonova · 2022
Later among the works it cites.
High fidelity neural audio compression
Alexandre D’efossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · 2020
Cited alongside, same era.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Cited alongside, same era.
Deliberation model based two-pass end-to-end speech recognition
Ke Hu, Tara N. Sainath, Ruoming Pang, and Rohit Prabhavalkar · 2020
Cited alongside, same era.
Europarl-st: A multilingual corpus for speech translation of parliamentary debates
Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerda, Javier Jorge, Nahuel Roselló, Adria Giménez, Albert Sanchis, Jorge Civera, and Alfons Juan · 2020
Cited alongside, same era.
Later among the works it cites.
Toward fairness in speech recognition: Discovery and mitigation of performance disparities
Pranav Dheram, Murugesan Ramakrishnan, Anirudh Raju, I-Fan Chen, Brian King, Katherine Powell, Melissa Saboowala, Karan Shetty, and Andreas Stolcke · 2022
Later among the works it cites.
T-modules: Translation modules for zero-shot cross-modal machine translation
Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, and Holger Schwenk · 2022
Later among the works it cites.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang · 2022
Later among the works it cites.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan · 2022
Later among the works it cites.
Glottolog database 4.6, 2022
Harald Hammarström, Robert Forkel, Martin Haspelmath, and Sebastian Bank · 2022
Later among the works it cites.
Bitext mining using distilled sentence representations for low-resource languages
Kevin Heffernan, Onur Çelebi, and Holger Schwenk · 2022
Later among the works it cites.
From simultaneous to streaming machine translation by leveraging streaming history
Javier Iranzo-Sánchez, Jorge Civera, and Alfons Juan · 2022
Later among the works it cites.
CVSS corpus and massively multilingual speech-to-speech translation
Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, and Heiga Zen · 2022
Later among the works it cites.
Samu-xlsr: Semantically-aligned multimodal utterance-level cross-lingual speech representation
Sameer Khurana, Antoine Laurent, and James Glass · 2022
Later among the works it cites.
Findings of the 2022 conference on machine translation (wmt22)
Tom Kocmi, Rachel Bawden, OndÅ™ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, Maja Popović, and Mariya Shmatova · 2022
Later among the works it cites.
Direct speech-to-speech translation with discrete units
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, and Wei-Ning Hsu · 2022
Later among the works it cites.
Textless speech-to-speech translation on real data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, and Wei-Ning Hsu · 2022
Later among the works it cites.
Consistent human evaluation of machine translation across language pairs
Daniel Licht, Cynthia Gao, Janice Lam, Francisco Guzman, Mona Diab, and Philipp Koehn · 2022
Later among the works it cites.
Towards measuring fairness in speech recognition: Casual conversations dataset transcriptions
Chunxi Liu, Michael Picheny, Leda Sarı, Pooja Chitkara, Alex Xiao, Xiaohui Zhang, Mark Chou, Andres Alvarado, Caner Hazirbas, and Yatharth Saraf · 2022
Later among the works it cites.
Hey asr system! why aren’t you more inclusive? automatic speech recognition systems’ bias and proposed bias mitigation techniques. a literature review
Mikel K Ngueajio and Gloria Washington · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation, 2022
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia-Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang · 2022
Later among the works it cites.
Perturbation augmentation for fairer NLP
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapr · 2022
Later among the works it cites.
Investigating failures of automatic translationin the case of unambiguous gender
Adithya Renduchintala and Adina Williams · 2022
Later among the works it cites.
Streaming parrotron for on-device speech-to-speech conversion
Oleg Rybakov, Fadi Biadsy, Xia Zhang, Liyang Jiang, Phoenix Meadowlark, and Shivani Agrawal · 2022
Later among the works it cites.
A taxonomy and study of critical errors in machine translation
Khetam Al Sharou and Lucia Specia · 2022
Later among the works it cites.
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed · 2022
Later among the works it cites.
Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia · 2022
Later among the works it cites.
“I’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams · 2022
Later among the works it cites.
SHAS: Approaching optimal Segmentation for End-to-End Speech Translation
Ioannis Tsiamas, Gerard I. Gállego, José A. R. Fonollosa, and Marta R. Costa-jussà · 2022
Later among the works it cites.
Wav2vec-switch: Contrastive learning from original-noisy speech pairs for robust speech recognition
Yiming Wang, Jinyu Li, Heming Wang, Yao Qian, Chengyi Wang, and Yu Wu · 2022
Later among the works it cites.
Barack Wanjawa, Lilian Wanzare, Florence Indede, Owen McOnyango, Edward Ombui, and Lawrence Muchemi · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2022
Later among the works it cites.
SpeechUT: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training
Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, and Furu Wei · 2022
Later among the works it cites.
M-Adapter: Modality Adaptation for End-to-End Speech-to-Text Translation
Jinming Zhao, Hao Yang, Gholamreza Haffari, and Ehsan Shareghi · 2022
Later among the works it cites.
A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition
Qiu-Shi Zhu, Jie Zhang, Zi-Qiang Zhang, Ming-Hui Wu, Xin Fang, and Li-Rong Dai · 2022
Later among the works it cites.
World-readiness standards for learning languages, 2023
ACTFL · 2023
Closest in time.
FINDINGS OF THE IWSLT 2023 EVALUATION CAMPAIGN
Milind Agarwal, Sweta Agrawal, Antonios Anastasopoulos, Luisa Bentivogli, Ondřej Bojar, Claudia Borg, Marine Carpuat, Roldano Cattoni, Mauro Cettolo, Mingda Chen, William Chen, Khalid Choukri, Alexandra Chronopoulou, Anna Currey, Thierry Declerck, Qianqian Dong, Kevin Duh, Yannick Estève, Marcello Federico, Souhir Gahbiche, Barry Haddow, Benjamin Hsu, Phu Mon Htut, Hirofumi Inaguma, Dávid Javorský, John Judge, Yasumasa Kano, Tom Ko, Rishu Kumar, Pengwei Li, Xutai Ma, Prashant Mathur, Evgeny Matusov, Paul McNamee, John P. McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr. Ojha, John E. Ortega, Proyag Pal, Juan Pino, Lonneke van der Plas, Peter Polák, Elijah Rippeth, Elizabeth Salesky, Jiatong Shi, Matthias Sperber, Sebastian Stüker, Katsuhito Sudoh, Yun Tang, Brian Thompson, Kevin Tran, Marco Turchi, Alex Waibel, Mingxuan Wang, Shinji Watanabe, and Rodolfo Zevallos · 2023
Closest in time.
Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu, Juan Pino, and Changhan Wang · 2023
Closest in time.
Audiolm: A language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour · 2023
Closest in time.
BLASER: A text-free speech-to-speech translation evaluation metric
Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, and Marta R. Costa-jussà · 2023
Closest in time.
xSIM++: An improved proxy to bitext mining performance for low-resource languages
Mingda Chen, Kevin Heffernan, Onur Çelebi, Alexandre Mourachko, and Holger Schwenk · 2023
Closest in time.
Speech-to-speech translation for a real-world unwritten language
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, and Ann Lee · 2023
Closest in time.
MuSLAM: Multitask, multilingual speech and language models
Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey, and Ankur Bapna · 2023
Closest in time.
Marta R Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, and Carleigh Wood · 2023
Closest in time.
Toxicity in multilingual machine translation at scale, 2023
Marta R. Costa-jussà, Eric Smith, Christophe Ropers, Daniel Licht, Jean Maillard, Javier Ferrando, and Carlos Escolano · 2023
Closest in time.
SpeechMatrix: A large-scale mined corpus of multilingual speech-to-speech translations
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, and Holger Schwenk · 2023
Closest in time.
Resetox: Re-learning attention weights for toxicity mitigation in machine translation, 2023
Javier García Gilabert, Carlos Escolano, and Marta R. Costa-Jussà · 2023
Closest in time.
Multilingual speech-to-speech translation into multiple target languages
Hongyu Gong, Ning Dong, Sravya Popuri, Vedanuj Goswami, Ann Lee, and Juan Pino · 2023
Closest in time.
UnitY: Two-pass direct speech-to-speech translation with discrete units
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, and Juan Pino · 2023
Closest in time.
Coexistence of multiple writing systems: Classifying digraphia in post-socialist countries
Youngjoo Jung and Bora Kim · 2023
Closest in time.
The effectiveness of machine translation in foreign language education: a systematic review and meta-analysis
Sangmin-Michelle Lee · 2023
Closest in time.
Automatic pipeline for gender multilingual data characterisation at scale
Benjamin Muller, Belen Alastruey, Prangthip Hansanti, Elahe Kalbassi, Christophe Ropers, Eric Smith, Adina Williams, Luke Zettlemoyer, Pierre Andrews, and Marta R. Costa-jussà · 2023
Closest in time.
The casual conversations v2 dataset
Bilal Porgali, Vítor Albiero, Jordan Ryda, Cristian Canton Ferrer, and Caner Hazirbas · 2023
Closest in time.
Scaling speech technology to 1,000+ languages, 2023
Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang, Wei-Ning Hsu, Alexis Conneau, and Michael Auli · 2023
Closest in time.
Audiopalm: A large language model that can speak and listen
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Ryan Padfield, James Qin, Daniel Rozenberg, Tara N. Sainath, Johan Schalkwyk, Matthew Sharifi, Michelle D. Tadmor, Ramanovich, Marco Tagliasacchi, Alexandru Tudor, Mihajlo Velimirovi’c, Damien Vincent, Jiahui Yu, Yongqiang Wang, Victoria Zayats, Neil Zeghidour, Yu Zhang, Zhishuai Zhang, Lukás Zilka, and Christian Havnø Frank · 2023
Closest in time.
CoVoST 2 and Massively Multilingual Speech Translation
Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino · 2027
Closest in time.
Filtering and mining parallel data in a joint multilingual space
Holger Schwenk · 2037
Closest in time.