Fetching the paper…
Reading the bibliography…
Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue.
The effects of lag time on interpreter errors
D. Cokely · 1986
Earlier work this paper cites.
Task requirements and media choice in collaborative writing
R. Kraut, J. Galegher, R. Fish, and B. Chalfonte · 1992
Earlier work this paper cites.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Comparing prosody across many languages
F. Cummins, F. Gers, and J. Schmidhuber · 1999
Earlier work this paper cites.
Aac intervention in the intensive care unit: The children’s hospital boston model
J. Costello · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Accessing assets: Immigrant youth’s work as family translators or" para-phrasers"
M. F. Orellana, L. Dorner, and L. Pulido · 2003
Earlier work this paper cites.
Integration of immigrants: The role of language proficiency and experience
L. Delander, M. Hammarstedt, J. MÅnsson, and E. Nyberg · 2005
Earlier work this paper cites.
Prosody generation for speech-to-speech translation
P. Aguero, J. Adell, and A. Bonafonte · 2006
Earlier work this paper cites.
The construction of moral and social identity in immigrant children’s narratives-in-translation
I. G. Sánchez and M. F. Orellana · 2006
Earlier work this paper cites.
Stress-associated poor health among adult immigrants with a language barrier in the united states
H. Ding and L. Hargraves · 2009
Earlier work this paper cites.
Multiple uses of machine translation and computerised translation tools
J. Hutchins · 2009
Earlier work this paper cites.
Ethnologue: Languages of the World
M. P. Lewis, editor · 2009
Earlier work this paper cites.
Overcoming the language barrier with speech translation technology
S. Nakamura · 2009
Earlier work this paper cites.
"opensmile - the munich versatile and fast open-source audio feature extractor"
F. Eyben, M. Wöllmer, and S. Schuller, Björn · 2010
Earlier work this paper cites.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu · 2010
Earlier work this paper cites.
Acculturative stress in latino immigrants: The impact of social, socio-psychological and migration-related factors
K. Lueck and M. Wilson · 2011
Earlier work this paper cites.
Intent transfer in speech-to-speech machine translation
G. K. Anumanchipalli, L. C. Oliveira, and A. W. Black · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2012
Earlier work this paper cites.
Paralinguistics in speech and language—state-of-the-art and the challenge
B. Schuller, S. Steidl, A. Batliner, F. Burkhardt, L. Devillers, C. MüLler, and S. Narayanan · 2013
Earlier work this paper cites.
Collection and analysis of a japanese-english emphasized speech corpora
G. Neubig, S. Sakti, T. Toda, S. Nakamura, et al · 2014
Earlier work this paper cites.
The mind in the machine: Anthropomorphism increases trust in an autonomous vehicle
A. Waytz, J. Heafner, and N. Epley · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
M. Popović · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
D. Snyder, G. Chen, and D. Povey · 2015
Earlier work this paper cites.
L. J. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Can neural machine translation do simultaneous translation?
K. Cho and M. Esipova · 2016
Earlier work this paper cites.
Turn-taking in human communication–origins and implications for language processing
S. C. Levinson · 2016
Earlier work this paper cites.
World: A vocoder-based high-quality speech synthesis system for real-time applications
M. Morise, F. Yokomori, and K. Ozawa · 2016
Earlier work this paper cites.
Toxic comment classification challenge, 2017
cjadams, J. Sorensen, J. Elliott, L. Dixon, M. McDonald, nithum, and W. Cukierski · 2017
Earlier work this paper cites.
Toward expressive speech translation: A unified sequence-to-sequence lstms approach for translating words and emphasis
Q. T. Do, S. Sakti, and S. Nakamura · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger · 2017
Earlier work this paper cites.
Transfer learning across low-resource, related languages for neural machine translation
T. Q. Nguyen and D. Chiang · 2017
Earlier work this paper cites.
Online and linear-time attention by enforcing monotonic alignments
C. Raffel, M.-T. Luong, P. J. Liu, R. J. Weiss, and D. Eck · 2017
Earlier work this paper cites.
Monotonic chunkwise attention
C.-C. Chiu* and C. Raffel* · 2018
Earlier work this paper cites.
Subjective evaluation of speech quality with a crowdsourcing approach, 2018
ITU-T Recommendation P.808 · 2018
Earlier work this paper cites.
Introducing Parselmouth: A Python interface to Praat
Y. Jadoul, B. Thompson, and B. de Boer · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
J. Lee, E. Mansimov, and K. Cho · 2018
Earlier work this paper cites.
Tadam: Task dependent adaptive metric for improved few-shot learning
B. N. Oreshkin, P. Rodriguez, and A. Lacoste · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville · 2018
Earlier work this paper cites.
Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu · 2018
Earlier work this paper cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous · 2018
Earlier work this paper cites.
Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
N. Arivazhagan, C. Cherry, W. Macherey, C.-C. Chiu, S. Yavuz, R. Pang, W. Li, and C. Raffel · 2019
Earlier work this paper cites.
Margin-based parallel corpus mining with multilingual sentence embeddings
M. Artetxe and H. Schwenk · 2019
Earlier work this paper cites.
An analysis of gender bias studies in natural language processing
M. Costa-jussà · 2019
Earlier work this paper cites.
Equalizing gender bias in neural machine translation with word embeddings techniques
J. Escudé Font and M. R. Costa-jussà · 2019
Earlier work this paper cites.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
J. Johnson, M. Douze, and H. Jégou · 2019
Earlier work this paper cites.
Z. Liu and B. Mak · 2019
Cited alongside, same era.
FlowSeq: Non-autoregressive conditional sequence generation with generative flow
X. Ma, C. Zhou, X. Li, G. Neubig, and E. Hovy · 2019
Cited alongside, same era.
Model cards for model reporting
M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru · 2019
Cited alongside, same era.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Cited alongside, same era.
Analysis of deep learning architectures for cross-corpus speech emotion recognition
J. Parry, D. Palaz, G. Clarke, P. Lecomte, R. Mead, M. Berger, and G. Hofer · 2019
Cited alongside, same era.
Self-supervised learning with random-projection quantizer for speech recognition
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu · 2022
Later among the works it cites.
Fleurs: Few-shot learning evaluation of universal representations of speech
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna · 2022
Later among the works it cites.
Interpreting gender bias in neural machine translation: Multilingual architecture matters
M. R. Costa-jussà, C. Escolano, C. Basta, J. Ferrando, R. Batlle, and K. Kharitonova · 2022
Later among the works it cites.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Later among the works it cites.
Glottolog database 4.6, 2022
H. Hammarström, R. Forkel, M. Haspelmath, and S. Bank · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards empathetic open-domain conversation models: a new benchmark and dataset
H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
G. Stanovsky, N. A. Smith, and L. Zettlemoyer · 2019
Cited alongside, same era.
Speech-to-speech translation between untranscribed unknown languages
A. Tjandra, S. Sakti, and S. Nakamura · 2019
Cited alongside, same era.
FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN
E. Ansari, A. Axelrod, N. Bach, O. Bojar, R. Cattoni, F. Dalvi, N. Durrani, M. Federico, C. Federmann, J. Gu, F. Huang, K. Knight, X. Ma, A. Nagesh, M. Negri, J. Niehues, J. Pino, E. Salesky, X. Shi, S. Stüker, M. Turchi, A. Waibel, and C. Wang · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning for speech recognition
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli · 2020
Cited alongside, same era.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
B. Desplanques, J. Thienpondt, and K. Demuynck · 2020
Cited alongside, same era.
CVSS corpus and massively multilingual speech-to-speech translation
Y. Jia, M. Tadmor Ramanovich, Q. Wang, and H. Zen · 2022
Later among the works it cites.
Text-free prosody-aware generative spoken language modeling
E. Kharitonov, A. Lee, A. Polyak, Y. Adi, J. Copet, K. Lakhotia, T. A. Nguyen, M. Rivière, A. Mohamed, E. Dupoux, and W. Hsu · 2022
Later among the works it cites.
Direct speech-to-speech translation with discrete units
A. Lee, P.-J. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu · 2022
Later among the works it cites.
Textless speech-to-speech translation on real data
A. Lee, H. Gong, P.-A. Duquenne, H. Schwenk, P.-J. Chen, C. Wang, S. Popuri, Y. Adi, J. Pino, J. Gu, and W.-N. Hsu · 2022
Later among the works it cites.
Consistent human evaluation of machine translation across language pairs
D. Licht, C. Gao, J. Lam, F. Guzman, M. Diab, and P. Koehn · 2022
Later among the works it cites.
Hey asr system! why aren’t you more inclusive? automatic speech recognition systems’ bias and proposed bias mitigation techniques. a literature review
M. K. Ngueajio and G. Washington · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation, 2022
NLLB Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. Mejia-Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, S. Spruit, C. Tran, P. Andrews, N. F. Ayan, S. Bhosale, S. Edunov, A. Fan, C. Gao, V. Goswami, F. Guzmán, P. Koehn, A. Mourachko, C. Ropers, S. Saleem, H. Schwenk, and J. Wang · 2022
Later among the works it cites.
Over-generation cannot be rewarded: Length-adaptive average lagging for simultaneous speech translation
S. Papi, M. Gaido, M. Negri, and M. Turchi · 2022
Later among the works it cites.
Red teaming language models with language models
E. Perez, S. Huang, H. F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Later among the works it cites.
Disentangling prosody representations with unsupervised speech reconstruction
L. Qu, T. Li, C. Weber, T. Pekarek-Rosin, F. Ren, and S. Wermter · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
Diffuser: Discrete diffusion via edit-based reconstruction
M. Reid, V. J. Hellendoorn, and G. Neubig · 2022
Later among the works it cites.
Daft-exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis
J. Zaïdi, H. Seuté, B. van Niekerk, and M.-A. Carbonneau · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2022
Later among the works it cites.
SpeechUT: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training
Z. Zhang, L. Zhou, J. Ao, S. Liu, L. Dai, J. Li, and F. Wei · 2022
Later among the works it cites.
URL http://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm
Chinese ai governance rules, 2023 · 2023
Closest in time.
URL https://artificialintelligenceact.eu/
European ai act, 2023 · 2023
Closest in time.
URL https://www.whitehouse.gov/wp-content/uploads/2023/07/Ensuring-Safe-Secure-and-Trustworthy-AI.pdf
Ensuring safe, secure, and trustworthy ai, 2023 · 2023
Closest in time.
FINDINGS OF THE IWSLT 2023 EVALUATION CAMPAIGN
M. Agarwal, S. Agrawal, A. Anastasopoulos, L. Bentivogli, O. Bojar, C. Borg, M. Carpuat, R. Cattoni, M. Cettolo, M. Chen, W. Chen, K. Choukri, A. Chronopoulou, A. Currey, T. Declerck, Q. Dong, K. Duh, Y. Estève, M. Federico, S. Gahbiche, B. Haddow, B. Hsu, P. Mon Htut, H. Inaguma, D. Javorský, J. Judge, Y. Kano, T. Ko, R. Kumar, P. Li, X. Ma, P. Mathur, E. Matusov, P. McNamee, J. P. McCrae, K. Murray, M. Nadejde, S. Nakamura, M. Negri, H. Nguyen, J. Niehues, X. Niu, A. Kr. Ojha, J. E. Ortega, P. Pal, J. Pino, L. van der Plas, P. Polák, E. Rippeth, E. Salesky, J. Shi, M. Sperber, S. Stüker, K. Sudoh, Y. Tang, B. Thompson, K. Tran, M. Turchi, A. Waibel, M. Wang, S. Watanabe, and R. Zevallos · 2023
Closest in time.
Towards cross-language prosody transfer for dialog, 2023
J. E. Avila and N. G. Ward · 2023
Closest in time.
Audiolm: A language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
BLASER: A text-free speech-to-speech translation evaluation metric
M. Chen, P.-A. Duquenne, P. Andrews, J. Kao, A. Mourachko, H. Schwenk, and M. R. Costa-jussà · 2023
Closest in time.
Speech-to-speech translation for a real-world unwritten language
P.-J. Chen, K. Tran, Y. Yang, J. Du, J. Kao, Y.-A. Chung, P. Tomasello, P.-A. Duquenne, H. Schwenk, H. Gong, H. Inaguma, S. Popuri, C. Wang, J. Pino, W.-N. Hsu, and A. Lee · 2023
Closest in time.
Multilingual holistic bias: Extending descriptors and patterns to unveil demographic biases in languages at scale
M. R. Costa-jussà, P. Andrews, E. Smith, P. Hansanti, C. Ropers, E. Kalbassi, C. Gao, D. Licht, and C. Wood · 2023
Closest in time.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
D. Dale, E. Voita, L. Barrault, and M. R. Costa-jussà · 2023
Closest in time.
Polyvoice: Language models for speech to speech translation
Q. Dong, Z. Huang, Q. Tian, C. Xu, T. Ko, Y. Zhao, S. Feng, T. Li, K. Wang, X. Cheng, F. Yue, Y. Bai, X. Chen, L. Lu, Z. Ma, Y. Wang, M. Wang, and Y. Wang · 2023
Closest in time.
The stable signature: Rooting watermarks in latent diffusion models
P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon · 2023
Closest in time.
A holistic cascade system, benchmark, and human evaluation protocol for expressive speech-to-speech translation
W.-C. Huang, B. Peloquin, J. Kao, C. Wang, H. Gong, E. Salesky, Y. Adi, A. Lee, and P.-J. Chen · 2023
Closest in time.
Collaborative watermarking for adversarial speech synthesis
L. Juvela and X. Wang · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
A watermark for large language models
J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein · 2023
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, et al · 2023
Closest in time.
Dear: A deep-learning-based audio re-recording resilient watermarking
C. Liu, J. Zhang, H. Fang, Z. Ma, W. Zhang, and N. Yu · 2023
Closest in time.
Efficient monotonic multihead attention
X. Ma, A. Sun, S. Ouyang, H. Inaguma, and P. Tomasello · 2023
Closest in time.
Expresso: A benchmark and analysis of discrete expressive speech resynthesis, 2023
T. A. Nguyen, W.-N. Hsu, A. D’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, F. Kreuk, Y. Adi, and E. Dupoux · 2023
Closest in time.
Scaling speech technology to 1,000+ languages, 2023
V. Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y. Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli · 2023
Closest in time.
Hybrid transformers for music source separation
S. Rouard, F. Massa, and A. Défossez · 2023
Closest in time.
Seamlessm4t-massively multilingual & multimodal machine translation, 2023
Seamless Communication, L. Barrault, Y.-A. Chung, M. C. Meglioli, D. Dale, N. Dong, P.-A. Duquenne, H. Elsahar, H. Gong, K. Heffernan, J. Hoffman, C. Klaiber, P. Li, D. Licht, J. Maillard, A. Rakotoarison, K. R. Sadagopan, G. Wenzek, E. Ye, B. Akula, P.-J. Chen, N. E. Hachem, B. Ellis, G. M. Gonzalez, J. Haaheim, P. Hansanti, R. Howes, B. Huang, M.-J. Hwang, H. Inaguma, S. Jain, E. Kalbassi, A. Kallet, I. Kulikov, J. Lam, D. Li, X. Ma, R. Mavlyutov, B. Peloquin, M. Ramadan, A. Ramakrishnan, A. Sun, K. Tran, T. Tran, I. Tufanov, V. Vogeti, C. Wood, Y. Yang, B. Yu, P. Andrews, C. Balioglu, M. R. Costa-jussà, O. Celebi, M. Elbayad, C. Gao, F. Guzmán, J. Kao, A. Lee, A. Mourachko, J. Pino, S. Popuri, C. Ropers, S. Saleem, H. Schwenk, P. Tomasello, C. Wang, J. Wang, and S. Wang · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Closest in time.
Dialogs re-enacted across languages, 2023
N. G. Ward, J. E. Avila, E. Rivas, and D. Marco · 2023
Closest in time.
Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust
Y. Wen, J. Kirchenbauer, J. Geiping, and T. Goldstein · 2023
Closest in time.
CoVoST 2 and Massively Multilingual Speech Translation
C. Wang, A. Wu, J. Gu, and J. Pino · 2027
Closest in time.
Incremental Decoding and Training Methods for Simultaneous Translation in Neural Machine Translation
F. Dalvi, N. Durrani, H. Sajjad, and S. Vogel · 2079
Closest in time.