Fetching the paper…
Reading the bibliography…
In this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer. 2019 · 1911
Earlier work this paper cites.
Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters
John Bridle. 1989 · 1989
Earlier work this paper cites.
The cache performance and optimizations of blocked algorithms
Monica D. Lam, Edward E. Rothberg, and Michael E. Wolf. 1991 · 1991
Earlier work this paper cites.
GLU variants improve transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
Slice sampling
Radford M. Neal. 2003 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Program optimization space pruning for a multithreaded gpu
Shane Ryoo, Christopher I Rodrigues, Sam S Stone, Sara S Baghsorkhi, Sain-Zee Ueng, John A Stratton, and Wen-mei W Hwu. 2008 · 2008
Earlier work this paper cites.
TED-LIUM: an automatic speech recognition dedicated corpus
Anthony Rousseau, Paul Deléglise, and Yannick Estève. 2012 · 2012
Earlier work this paper cites.
Professional CUDA C Programming
J. Cheng, M. Grossman, and T. McKercher. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Efficient softmax approximation for GPUs
Édouard Grave, Armand Joulin, Moustapha Cissé, David Grangier Facebook AI Research, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
Svd-softmax: Fast softmax approximation on large vocabulary neural networks
Kyuhong Shim, Minjae Lee, Iksoo Choi, Yoonho Boo, and Wonyong Sung. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Hardware-aware softmax approximation for deep neural networks
Xue Geng, Jie Lin, Bin Zhao, Anmin Kong, Mohamed M. Sabry Aly, and Vijay Ramaseshan Chandrasekhar. 2018 · 2018
Earlier work this paper cites.
Dissecting the NVIDIA volta GPU architecture via microbenchmarking
Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P. Scarpazza. 2018 · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Maxim Milakov and Natalia Gimelshein. 2018 · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. 2018 · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Root mean square layer normalization
Biao Zhang and Rico Sennrich. 2019 · 2019
Cited alongside, same era.
Common Voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020 · 2020
Cited alongside, same era.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Narain Sohl-Dickstein, and Roman Novak. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling
Sanchit Gandhi, Patrick von Platen, and Alexander M Rush. 2023 · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Later among the works it cites.
Softmax output approximation for activation memory-efficient training of attention-based networks
Changhyeon Lee and Seulki Lee. 2023 · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Later among the works it cites.
PaSS: Parallel speculative sampling
Giovanni Monea, Armand Joulin, and Edouard Grave. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
NVIDIA A100 tensor core GPU architecture
NVIDIA Corporation. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush. 2021 · 2021
Cited alongside, same era.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart Van Baalen, and Tijmen Blankevoort. 2021 · 2021
Cited alongside, same era.
Self-attention does not need o ( n 2 ) o(n^{2}) memory
Markus N Rabe and Charles Staats. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Later among the works it cites.
A study on relu and softmax in transformer
Kai Shen, Junliang Guo, Xu Tan, Siliang Tang, Rui Wang, and Jiang Bian. 2023 · 2023
Later among the works it cites.
Accelerating LLM inference with staged speculative decoding
Benjamin Spector and Chris Re. 2023 · 2023
Later among the works it cites.
Replacing softmax with relu in vision transformers
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. 2023 · 2023
Later among the works it cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
Josh Achiam et al. 2024 · 2024
Closest in time.
Medusa: Simple LLM inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. 2024 · 2024
Closest in time.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2024 · 2024
Closest in time.
Sophie Fischer, Carlos Gemmell, Niklas Tecklenburg, Iain Mackie, Federico Rossetto, and Jeffrey Dalton. 2024 · 2024
Closest in time.
The unreasonable ineffectiveness of the deeper layers
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A Roberts. 2024 · 2024
Closest in time.
Optimizing parallel reduction in cuda
Mark Harris. 2007 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024 · 2024
Closest in time.
Theory, analysis, and best practices for sigmoid self-attention
Jason Ramapuram, Federico Danieli, Eeshan Dhekane, Floris Weers, Dan Busbridge, Pierre Ablin, Tatiana Likhomanenko, Jagrit Digani, Zijin Gu, Amitis Shidani, and Russ Webb. 2024 · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024 · 2024
Closest in time.
Multi-candidate speculative decoding
Sen Yang, Shujian Huang, Xinyu Dai, and Jiajun Chen. 2024 · 2024
Closest in time.
Distillspec: Improving speculative decoding via knowledge distillation
Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. 2024 · 2024
Closest in time.