Fetching the paper…
Reading the bibliography…
We propose Speculative Decoding (SpecDec), for the first time ever, to formally study exploiting the idea of speculative execution to accelerate autoregressive (AR) decoding.
Encode, tag, realize: High-precision text editing
Eric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, and Aliaksei Severyn. 2019 · 1909
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Glancing transformer for non-autoregressive neural machine translation
Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, and Lei Li. 2021 · 2003
Earlier work this paper cites.
Gector–grammatical error correction: tag, not rewrite
Kostiantyn Omelianchuk, Vitaliy Atrasevych, Artem Chernodub, and Oleksandr Skurzhanskyi. 2020 · 2005
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindrich Libovický and Jindrich Helcl. 2018 · 2018
Earlier work this paper cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. 2018 · 2018
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Improving the efficiency of grammatical error correction with erroneous span detection and correction
Mengyun Chen, Tao Ge, Xingxing Zhang, Furu Wei, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Aligned cross entropy for non-autoregressive machine translation
Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Non-autoregressive machine translation with disentangled context transformer
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu. 2020 · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Edgeformer: A parameter-efficient transformer for on-device seq2seq generation
Tao Ge, Si-Qing Chen, and Furu Wei. 2022a · 2022
Closest in time.
Directed acyclic transformer for non-autoregressive machine translation
Fei Huang, Hao Zhou, Yang Liu, Hang Li, and Minlie Huang. 2022 · 2022
Closest in time.
Non-autoregressive neural machine translation: A call for clarity
Robin M. Schmidt, Telmo Pires, Stephan Peitz, and Jonas Lööf. 2022 · 2022
Closest in time.
Non-monotonic latent alignments for ctc-based non-autoregressive machine translation
Chenze Shao and Yang Feng. 2022 · 2022
Closest in time.
Sustainable AI: environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga Behram, Jinshi Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin S. Lee, Bugra Akyildiz, Maximilian Balandat, Joe Spisak, Ravi Jain, Mike Rabbat, and Kim M. Hazelwood. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho. 2020 · 2020
Cited alongside, same era.
Non-autoregressive translation by learning target categorical codes
Yu Bao, Shujian Huang, Tong Xiao, Dongqi Wang, Xinyu Dai, and Jiajun Chen. 2021 · 2021
Cited alongside, same era.
Learning to rewrite for non-autoregressive neural machine translation
Xinwei Geng, Xiaocheng Feng, and Bing Qin. 2021 · 2021
Cited alongside, same era.
Fully non-autoregressive neural machine translation: Tricks of the trade
Jiatao Gu and Xiang Kong. 2021 · 2021
Cited alongside, same era.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
Closest in time.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023 · 2023
Closest in time.
Speculative decoding with big little decoder
Sehoon Kim, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, Amir Gholami, and Kurt Keutzer. 2023 · 2023
Closest in time.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Closest in time.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2023 · 2023
Closest in time.
Accelerating transformer inference for translation via parallel decoding
Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, and Emanuele Rodolà. 2023 · 2023
Closest in time.
Accelerating LLM inference with staged speculative decoding
Benjamin Spector and Chris Re. 2023 · 2023
Closest in time.
Inference with reference: Lossless acceleration of large language models
Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023 · 2023
Closest in time.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. 2023 · 2023
Closest in time.
Distillspec: Improving speculative decoding via knowledge distillation
Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. 2023 · 2023
Closest in time.