Fetching the paper…
Reading the bibliography…
The study of the attention mechanism has sparked interest in many fields, such as language modeling and machine translation.
Adding interpretable attention to neural translation models improves word alignment
Thomas Zenkel, Joern Wuebker, and John DeNero. 2019 · 1901
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philémon Brakel, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016 · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Sequence-Level Knowledge Distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
OPUS – parallel corpora for everyone
Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sequence-to-Sequence Models Can Directly Translate Foreign Speech
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Earlier work this paper cites.
A Call for Clarity in Reporting BLEU Scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Earlier work this paper cites.
An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018 · 2018
Earlier work this paper cites.
Thinking slow about latency evaluation for simultaneous machine translation
Colin Cherry and George Foster. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
Jointly learning to align and translate with transformer models
Sarthak Garg, Stephan Peitz, Udhyakumar Nallasamy, and Matthias Paulik. 2019 · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework
Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019 · 2019
Earlier work this paper cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Cited alongside, same era.
FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN
Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, Fei Huang, Kevin Knight, Xutai Ma, Ajay Nagesh, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Xing Shi, Sebastian Stüker, Marco Turchi, Alexander Waibel, and Changhan Wang. 2020 · 2020
Cited alongside, same era.
Losing heads in the lottery: Pruning transformer attention in neural machine translation
Maximiliana Behnke and Kenneth Heafield. 2020 · 2020
Cited alongside, same era.
End-to-end asr with adaptive span self-attention
Xuankai Chang, Aswin Shanmugam Subramanian, Pengcheng Guo, Shinji Watanabe, Yuya Fujita, and Motoi Omachi. 2020 · 2020
Cited alongside, same era.
Recent developments on espnet toolkit boosted by conformer
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang. 2021 · 2021
Later among the works it cites.
Simultaneous speech translation for live subtitling: from delay to display
Alina Karakanta, Sara Papi, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
The USTC-NELSLIP systems for simultaneous speech translation task at IWSLT 2021
Dan Liu, Mengge Du, Xiaoxi Li, Yuchen Hu, and Lirong Dai. 2021a · 2021
Later among the works it cites.
Cross attention augmented transducer networks for simultaneous translation
Dan Liu, Mengge Du, Xiaoxi Li, Ya Li, and Enhong Chen. 2021b · 2021
Later among the works it cites.
Streaming simultaneous speech translation with augmented memory transformer
Xutai Ma, Yongqiang Wang, Mohammad Javad Dousti, Philipp Koehn, and Juan Pino. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accurate word alignment induction from neural machine translation
Yun Chen, Yang Liu, Guanhua Chen, Xin Jiang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
On Target Segmentation for Direct Speech Translation
Mattia A. Di Gangi, Marco Gaido, Matteo Negri, and Marco Turchi. 2020 · 2020
Cited alongside, same era.
On Knowledge Distillation for Direct Speech Translation
Marco Gaido, Mattia A. Di Gangi, Matteo Negri, and Marco Turchi. 2021b · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020 · 2020
Cited alongside, same era.
End-to-end simultaneous translation system for IWSLT2020 using modality agnostic meta-learning
Hou Jeung Han, Mohd Abbas Zaidi, Sathish Reddy Indurthi, Nikhil Kumar Lakumarapu, Beomseok Lee, and Sangha Kim. 2020 · 2020
Cited alongside, same era.
Roles and utilization of attention heads in transformer-based neural language models
Jae-young Jo and Sung-Hyon Myaeng. 2020 · 2020
Cited alongside, same era.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Cited alongside, same era.
An empirical study of end-to-end simultaneous speech translation decoding strategies
Ha Nguyen, Yannick Estève, and Laurent Besacier. 2021 · 2021
Later among the works it cites.
Speechformer: Reducing information loss in direct speech translation
Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
Decision attentive regularization to improve simultaneous speech translation systems
Mohd Abbas Zaidi, Beomseok Lee, Nikhil Kumar Lakumarapu, Sangha Kim, and Chanwoo Kim. 2021 · 2021
Later among the works it cites.
RealTranS: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer
Xingshan Zeng, Liangyou Li, and Qun Liu. 2021 · 2021
Later among the works it cites.
Findings of the IWSLT 2022 evaluation campaign
Antonios Anastasopoulos, Loïc Barrault, Luisa Bentivogli, Marcely Zanon Boito, Ondřej Bojar, Roldano Cattoni, Anna Currey, Georgiana Dinu, Kevin Duh, Maha Elbayad, Clara Emmanuel, Yannick Estève, Marcello Federico, Christian Federmann, Souhir Gahbiche, Hongyu Gong, Roman Grundkiewicz, Barry Haddow, Benjamin Hsu, Dávid Javorský, Vĕra Kloudová, Surafel Lakew, Xutai Ma, Prashant Mathur, Paul McNamee, Kenton Murray, Maria Nǎdejde, Satoshi Nakamura, Matteo Negri, Jan Niehues, Xing Niu, John Ortega, Juan Pino, Elizabeth Salesky, Jiatong Shi, Matthias Sperber, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Yogesh Virkar, Alexander Waibel, Changhan Wang, and Shinji Watanabe. 2022 · 2022
Closest in time.
Uconv-conformer: High reduction of input sequence length for end-to-end speech recognition
Andrei Andrusenko, Rauf Nasretdinov, and Aleksei Romanenko. 2022 · 2022
Closest in time.
Exploring Continuous Integrate-and-Fire for Adaptive Simultaneous Speech Translation
Chih-Chiang Chang and Hung-Yi Lee. 2022 · 2022
Closest in time.
Towards opening the black box of neural machine translation: Source and target interpretations of the transformer
Javier Ferrando, Gerard I Gállego, Belen Alastruey, Carlos Escolano, and Marta R Costa-jussà. 2022 · 2022
Closest in time.
Efficient yet competitive speech translation: FBK@IWSLT2022
Marco Gaido, Sara Papi, Dennis Fucci, Giuseppe Fiameni, Matteo Negri, and Marco Turchi. 2022b · 2022
Closest in time.
Squeezeformer: An efficient transformer for automatic speech recognition
Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer. 2022 · 2022
Closest in time.
Attention weights accurately predict language representations in the brain
Mathis Lamarre, Catherine Chen, and Fatma Deniz. 2022 · 2022
Closest in time.
Align, write, re-order: Explainable end-to-end speech translation via operation sequence generation
Motoi Omachi, Brian Yan, Siddharth Dalmia, Yuya Fujita, and Shinji Watanabe. 2022 · 2022
Closest in time.
Does simultaneous speech translation need simultaneous models?
Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2022a · 2022
Closest in time.
CUNI-KIT system for simultaneous speech translation task at IWSLT 2022
Peter Polák, Ngoc-Quan Pham, Tuan Nam Nguyen, Danni Liu, Carlos Mullov, Jan Niehues, Ondřej Bojar, and Alexander Waibel. 2022 · 2022
Closest in time.
Cross-Modal Decision Regularization for Simultaneous Speech Translation
Mohd Abbas Zaidi, Beomseok Lee, Sangha Kim, and Chanwoo Kim. 2022 · 2022
Closest in time.
Information-transport-based policy for simultaneous translation
Shaolei Zhang and Yang Feng. 2022 · 2022
Closest in time.
Reproducibility is nothing without correctness: The importance of testing code in nlp
Sara Papi, Marco Gaido, Matteo Negri, and Andrea Pilzer. 2023 · 2023
Closest in time.