Fetching the paper…
Reading the bibliography…
Recent advances in state-of-the-art DNN architecture design have been moving toward Transformer models.
Llvm: A compilation framework for lifelong program analysis & transformation
Chris Lattner and Vikram Adve · 2004
Earlier work this paper cites.
The libm library and floatingpoint arithmetic in hp-ux for itanium-based systems
James W Thomas, John P Okada, Peter Markstein, and Ren-Chang Li · 2004
Earlier work this paper cites.
A parameterized floating-point exponential function for fpgas
Jérémie Detrey and Florent de Dinechin · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Ansor: Generating High-Performance Tensor Programs for Deep Learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica · 2006
Earlier work this paper cites.
A practical automatic polyhedral parallelizer and locality optimizer
Uday Bondhugula, Albert Hartono, Jagannathan Ramanujam, and Ponnuswamy Sadayappan · 2008
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Samuel Williams, Andrew Waterman, and David Patterson · 2009
Earlier work this paper cites.
Polly-polyhedral optimization in llvm
Tobias Grosser, Hongbin Zheng, Raghesh Aloor, Andreas Simbürger, Armin Größlinger, and Louis-Noël Pouchet · 2011
Earlier work this paper cites.
Neural Acceleration for General-Purpose Approximate Programs
Hadi Esmaeilzadeh, Adrian Sampson, Luis Ceze, and Doug Burger · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
When polyhedral transformations meet simd code generation
Martin Kong, Richard Veras, Kevin Stock, Franz Franchetti, Louis-Noël Pouchet, and Ponnuswamy Sadayappan · 2013
Earlier work this paper cites.
Predictive modeling in a polyhedral optimization space
Eunjung Park, John Cavazos, Louis-Noël Pouchet, Cédric Bastoul, Albert Cohen, and P Sadayappan · 2013
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Fredo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Opentuner: An extensible framework for program autotuning
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe · 2014
Earlier work this paper cites.
Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning
Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam · 2014
Earlier work this paper cites.
DaDianNao: A Machine-learning Supercomputer
Yunji Chen, Tao Luo, Shaoli Liu, Shijin Zhang, Liqiang He, Jia Wang, Ling Li, Tianshi Chen, Zhiwei Xu, Ninghui Sun, and Olivier Temam · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
A 240 g-ops/s mobile coprocessor for deep neural networks
Vinayak Gokhale, Jonghoon Jin, Aysegul Dundar, Berin Martini, and Eugenio Culurciello · 2014
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Caffe: Convolutional Architecture for Fast Feature Embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross B. Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
Hardware implementation of the exponential function using taylor series
Peter Nilsson, Ateeq Ur Rahman Shaik, Rakesh Gangarajaiah, and Erik Hertz · 2014
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Pencil: A platform-neutral compute intermediate language for accelerator programming
Riyadh Baghdadi, Ulysse Beaugnon, Albert Cohen, Tobias Grosser, Michael Kruse, Chandan Reddy, Sven Verdoolaege, Adam Betts, Alastair F. Donaldson, Jeroen Ketema, Javed Absar, Sven Van Haastregt, Alexey Kravets, Anton Lokhmotov, Robert David, and Elnar Hajiyev · 2015
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Shidiannao: Shifting vision processing closer to the sensor
Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Cnvlutin: Ineffectual-neuron-free deep neural network computing
Jorge Albericio, Patrick Judd, Tayler Hetherington, Tor Aamodt, Natalie Enright Jerger, and Andreas Moshovos · 2016
Earlier work this paper cites.
Fused-layer cnn accelerators
Manoj Alwani, Han Chen, Michael Ferdman, and Peter Milder · 2016
Earlier work this paper cites.
Designing neural network architectures using reinforcement learning
Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar · 2016
Earlier work this paper cites.
The pluto+ algorithm: A practical approach for parallelization and locality optimization of affine loop nests
Uday Bondhugula, Aravind Acharya, and Albert Cohen · 2016
Earlier work this paper cites.
Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks
Y. Chen, J. Emer, and V. Sze · 2016
Earlier work this paper cites.
Eyeriss: A Spatial Architecture for Energy-efficient Dataflow for Convolutional Neural Networks
Yu-Hsin Chen, Joel Emer, and Vivienne Sze · 2016
Earlier work this paper cites.
ASAP7: A 7-nm FinFET Predictive Process Design Kit
L.T. Clark, V. Vashishtha, L. Shifren, A. Gujia, S. Sinha, B. Cline, C. Ramamurthya, and G. Yeric · 2016
Earlier work this paper cites.
The log-sum-exp trick in machine learning, 2016
Robert Eisele · 2016
Earlier work this paper cites.
Eie: Efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
SqueezeNet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
The max trick when computing softmax, 2016
James McCaffrey · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Automatically scheduling halide image processing pipelines
Ravi Teja Mullapudi, Andrew Adams, Dillon Sharlet, Jonathan Ragan-Kelley, and Kayvon Fatahalian · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Cambricon-x: An accelerator for sparse neural networks
Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Using dataflow to optimize energy efficiency of deep neural network accelerators
Yu-Hsin Chen, Joel Emer, and Vivienne Sze · 2017
Earlier work this paper cites.
Tetris: Scalable and Efficient Neural Network Acceleration with 3D Memory
Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
New movidius myriad x vpu packs a custom neural compute engine, 2017
J Hruska · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs.(2017)
Shankar Iyer, Nikhil Dandekar, and Kornl Csernai · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon · 2017
Earlier work this paper cites.
The tensor algebra compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2017
Earlier work this paper cites.
Hierarchical representations for efficient architecture search
Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Scnn: An accelerator for compressed-sparse convolutional neural networks
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally · 2017
Earlier work this paper cites.
Efficient processing of deep neural networks: A tutorial and survey, 2017
Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel Emer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
https://cloud.google.com/edge-tpu/
Edge TPU · 2018
Earlier work this paper cites.
An approach for finding permutations quickly: Fusion and dimension matching
Aravind Acharya, Uday Bondhugula, and Albert Cohen · 2018
Earlier work this paper cites.
Polyhedral auto-transformation with no integer linear programming
Aravind Acharya, Uday Bondhugula, and Albert Cohen · 2018
Earlier work this paper cites.
Snapea: Predictive early activation for reducing computation in deep convolutional neural networks
Vahideh Akhlaghi, Amir Yazdanbakhsh, Kambiz Samadi, Rajesh K. Gupta, and Hadi Esmaeilzadeh · 2018
Earlier work this paper cites.
Understanding and simplifying one-shot architecture search
Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le · 2018
Earlier work this paper cites.
Proxylessnas: Direct neural architecture search on target task and hardware
Han Cai, Ligeng Zhu, and Song Han · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Dpp-net: Device-aware progressive search for pareto-optimal neural architectures
Jin-Dong Dong, An-Chieh Cheng, Da-Cheng Juan, Wei Wei, and Min Sun · 2018
Earlier work this paper cites.
Lectures on randomized numerical linear algebra
Petros Drineas and Michael W Mahoney · 2018
Earlier work this paper cites.
Gap-8: A risc-v soc for ai at the edge of the iot
Eric Flamand, Davide Rossi, Francesco Conti, Igor Loi, Antonio Pullini, Florent Rotenberg, and Luca Benini · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Firesim: Fpga-accelerated cycle-exact scale-out system simulation in the public cloud
S. Karandikar, H. Mao, D. Kim, D. Biancolin, A. Amid, D. Lee, N. Pemberton, E. Amaro, C. Schmidt, A. Chopra, Q. Huang, K. Kovacs, B. Nikolic, R. Katz, J. Bachrach, and K. Asanovic · 2018
Earlier work this paper cites.
Adaptive tiling: Applying fixed-size systolic arrays to sparse convolutional neural networks
H. T. Kung, Bradley McDanel, and Sai Qian Zhang · 2018
Earlier work this paper cites.
Progressive neural architecture search
Chenxi Liu, Barret Zoph, Maxim Neumann, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, and Kevin Murphy · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Earlier work this paper cites.
Neural architecture optimization
Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, and Tie-Yan Liu · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax, 2018
Maxim Milakov and Natalia Gimelshein · 2018
Earlier work this paper cites.
TensorRT: https://developer.nvidia.com/tensorrt, 2018
NVIDIA · 2018
Earlier work this paper cites.
Efficient neural architecture search via parameters sharing
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Glow: Graph lowering compiler techniques for neural networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Garret Catron, Summer Deng, Roman Dzhabarov, Nick Gibson, James Hegeman, Meghan Lele, Roman Levenstein, et al · 2018
Earlier work this paper cites.
MobilenetV2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
The NVIDIA Deep Learning Accelerator
Frans Sijstermans · 2018
Cited alongside, same era.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions, 2018
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
A high-speed and low-complexity architecture for softmax function in deep learning
Meiqi Wang, Siyuan Lu, Danyang Zhu, Jun Lin, and Zhongfeng Wang · 2018
Cited alongside, same era.
Efficient conformer: Progressive downsampling and grouped attention for automatic speech recognition
Maxime Burchi and Valentin Vielzeuf · 2021
Later among the works it cites.
Glit: Neural architecture search for global and local image transformer
Boyu Chen, Peixia Li, Chuming Li, Baopu Li, Lei Bai, Chen Lin, Ming Sun, Junjie Yan, and Wanli Ouyang · 2021
Later among the works it cites.
Accelerating transformer networks through recomposing softmax layers
Jaewan Choi, Hailong Li, Byeongho Kim, Seunghwan Hwang, and Jung Ho Ahn · 2021
Later among the works it cites.
Nvidia a100 tensor core gpu: Performance and innovation
Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky · 2021
Later among the works it cites.
Hardware acceleration of sparse and irregular tensor computations of ml models: A survey and insights
Shail Dave, Riyadh Baghdadi, Tony Nowatzki, Sasikanth Avancha, Aviral Shrivastava, and Baoxin Li · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer · 2018
Cited alongside, same era.
Netadapt: Platform-aware neural network adaptation for mobile applications
Tien-Ju Yang, Andrew Howard, Bo Chen, Xiao Zhang, Alec Go, Mark Sandler, Vivienne Sze, and Hartwig Adam · 2018
Cited alongside, same era.
Practical block-wise neural network architecture generation
Zhao Zhong, Junjie Yan, Wei Wu, Jing Shao, and Cheng-Lin Liu · 2018
Cited alongside, same era.
Cambricon-s: Addressing irregularity in sparse neural networks through a cooperative software/hardware approach
Xuda Zhou, Zidong Du, Qi Guo, Shaoli Liu, Chengsi Liu, Chao Wang, Xuehai Zhou, Ling Li, Tianshi Chen, and Yunji Chen · 2018
Cited alongside, same era.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Cited alongside, same era.
Learning to optimize halide with tree search and random programs
Andrew Adams, Karima Ma, Luke Anderson, Riyadh Baghdadi, Tzu-Mao Li, Michael Gharbi, Benoit Steiner, Steven Johnson, Kayvon Fatahalian, Frédo Durand, and Jonathan Ragan-Kelley · 2019
Cited alongside, same era.
Tiramisu: A polyhedral compiler for expressing fast and portable code
R. Baghdadi, J. Ray, M. B. Romdhane, E. D. Sozzo, A. Akkas, Y. Zhang, P. Suriana, S. Kamil, and S. Amarasinghe · 2019
Cited alongside, same era.
Gemmini: Enabling systematic deep-learning architecture evaluation via full-stack integration
Hasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali, Vighnesh Iyer, Pranav Prakash, Jerry Zhao, Daniel Grubb, Harrison Liew, Howard Mao, Albert Ou, Colin Schmidt, Samuel Steffl, John Wright, Ion Stoica, Jonathan Ragan-Kelley, Krste Asanovic, Borivoje Nikolic, and Yakun Sophia Shao · 2021
Later among the works it cites.
A survey of quantization methods for efficient neural network inference, 2021
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
Nasvit: Neural architecture search for efficient vision transformers with gradient conflict aware supernet training
Chengyue Gong, Dilin Wang, Meng Li, Xinlei Chen, Zhicheng Yan, Yuandong Tian, Vikas Chandra, et al · 2021
Later among the works it cites.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W. Lee · 2021
Later among the works it cites.
Mind mappings: enabling efficient algorithm-accelerator mapping space search
Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W Fletcher · 2021
Later among the works it cites.
Cosa: Scheduling by constrained optimization for spatial accelerators
Qijing Huang, Minwoo Kang, Grace Dinh, Thomas Norell, Aravind Kalaiah, James Demmel, John Wawrzynek, and Yakun Sophia Shao · 2021
Later among the works it cites.
A learned performance model for tensor processing units
Sam Kaufman, Phitchaya Phothilimthana, Yanqi Zhou, Charith Mendis, Sudip Roy, Amit Sabne, and Mike Burrows · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
Learned token pruning for transformers
Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer · 2021
Later among the works it cites.
Graphcore
Simon Knowles · 2021
Later among the works it cites.
Bert busters: Outlier dimensions that disrupt transformers
Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky · 2021
Later among the works it cites.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush · 2021
Later among the works it cites.
The deep learning compiler: A comprehensive survey
Mingzhen Li, Yi Liu, Xiaoyan Liu, Qingxiao Sun, Xin You, Hailong Yang, Zhongzhi Luan, Lin Gan, Guangwen Yang, and Depei Qian · 2021
Later among the works it cites.
Analytical characterization and design space exploration for optimization of cnns
Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, and P Sadayappan · 2021
Later among the works it cites.
Searching for efficient multi-stage vision transformers
Yi-Lun Liao, Sertac Karaman, and Vivienne Sze · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Ebert: Efficient bert inference with dynamic structured pruning
Zejian Liu, Fanrong Li, Gang Li, and Jian Cheng · 2021
Later among the works it cites.
Tenet: A framework for modeling tensor dataflow based on relation-centric notation
Liqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia, Zizhang Luo, Jieming Yin, Jason Cong, and Yun Liang · 2021
Later among the works it cites.
Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang · 2021
Later among the works it cites.
Token pooling in vision transformers
Dmitrii Marin, Jen-Hao Rick Chang, Anurag Ranjan, Anish Prabhu, Mohammad Rastegari, and Oncel Tuzel · 2021
Later among the works it cites.
Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer
Sachin Mehta and Mohammad Rastegari · 2021
Later among the works it cites.
Zigzag: Enlarging joint architecture-mapping design space exploration for dnn accelerators
Linyan Mei, Pouya Houshmand, Vikram Jain, Sebastian Giraldo, and Marian Verhelst · 2021
Later among the works it cites.
Ioopt: automatic derivation of i/o complexity bounds for affine programs
Auguste Olivry, Guillaume Iooss, Nicolas Tollenaere, Atanas Rountev, P Sadayappan, and Fabrice Rastello · 2021
Later among the works it cites.
Demystifying bert: Implications for accelerator design
Suchita Pati, Shaizeen Aga, Nuwan Jayasena, and Matthew D Sinclair · 2021
Later among the works it cites.
Sambanova sn10 rdu: Accelerating software 2.0 with dataflow
Raghu Prabhakar and Sumti Jairath · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Later among the works it cites.
A comprehensive survey of neural architecture search: Challenges and solutions
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Xiaojiang Chen, and Xin Wang · 2021
Later among the works it cites.
Consistent accelerated inference via confident adaptive transformers
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay · 2021
Later among the works it cites.
Neural architecture search and hardware accelerator co-search: A survey
Lukas Sekanina · 2021
Later among the works it cites.
Searching for efficient transformers for language modeling
David So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le · 2021
Later among the works it cites.
Softermax: Hardware/software co-design of an efficient softmax for transformers
Jacob R. Stevens, Rangharajan Venkatesan, Steve Dai, Brucek Khailany, and Anand Raghunathan · 2021
Later among the works it cites.
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M. Rush, David Brooks, and Gu-Yeon Wei · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Alphanet: improved training of supernets with alpha-divergence
Dilin Wang, Chengyue Gong, Meng Li, Qiang Liu, and Vikas Chandra · 2021
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han · 2021
Later among the works it cites.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2021
Later among the works it cites.
Nas-bert: task-agnostic and adaptive-size bert compression with neural architecture search
Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu · 2021
Later among the works it cites.
https://openai.com/blog/chatgpt/ , 2022
Chatgpt: Optimizing language models for dialogue · 2022
Later among the works it cites.
The groq software-defined scale-out tensor streaming multiprocessor: From chips-to-systems architectural overview
Dennis Abts, John Kim, Garrin Kimmell, Matthew Boyd, Kris Kang, Sahil Parmar, Andrew Ling, Andrew Bitar, Ibrahim Ahmed, and Jonathan Ross · 2022
Later among the works it cites.
Hardware approximate techniques for deep neural network accelerators: A survey
Giorgos Armeniakos, Georgios Zervakis, Dimitrios Soudris, and Jörg Henkel · 2022
Later among the works it cites.
Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Han Cai, Chuang Gan, and Song Han · 2022
Later among the works it cites.
Mobile-former: Bridging mobilenet and transformer
Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Later among the works it cites.
Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design, 2022
Hongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D. Lane, and Mohamed S. Abdelfattah · 2022
Later among the works it cites.
An algorithm-hardware co-optimized framework for accelerating n:m sparse transformers
Chao Fang, Aojun Zhou, and Zhongfeng Wang · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Exocompilation for productive programming of hardware accelerators
Yuka Ikarashi, Gilbert Louis Bernstein, Alex Reinking, Hasan Genc, and Jonathan Ragan-Kelley · 2022
Later among the works it cites.
Demystifying map space exploration for npus, 2022
Sheng-Chun Kao, Angshuman Parashar, Po-An Tsai, and Tushar Krishna · 2022
Later among the works it cites.
An optimized dataflow for mitigating attention performance bottlenecks
Sheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, and Tushar Krishna · 2022
Later among the works it cites.
A 17–95.6 tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization for transformers in 5nm
Ben Keller, Rangharajan Venkatesan, Steve Dai, Stephen G. Tell, Brian Zimmer, William J. Dally, C. Thomas Gray, and Brucek Khailany · 2022
Later among the works it cites.
Squeezeformer: An efficient transformer for automatic speech recognition
Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer · 2022
Later among the works it cites.
Integer-only zero-shot quantization for efficient speech recognition
Sehoon Kim, Amir Gholami, Zhewei Yao, Nicholas Lee, Patrick Wang, Aniruddha Nrusimha, Bohan Zhai, Tianren Gao, Michael W Mahoney, and Kurt Keutzer · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh · 2022
Later among the works it cites.
Fp8 quantization: The power of the exponent
Andrey Kuzmin, Mart Van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, and Tijmen Blankevoort · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers
Woosuk Kwon, Sehoon Kim, Michael W Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami · 2022
Later among the works it cites.
Efficientformer: Vision transformers at mobilenet speed
Yanyu Li, Geng Yuan, Yang Wen, Eric Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren · 2022
Later among the works it cites.
Accelerating attention through gradient-based learned runtime pruning
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang · 2022
Later among the works it cites.
I-vit: Integer-only quantization for efficient vision transformer inference, 2022
Zhikai Li and Qingyi Gu · 2022
Later among the works it cites.
Cerebras architecture deep dive: First look inside the hw/sw co-design for deep learning: Cerebras systems
Sean Lie · 2022
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2022
Later among the works it cites.
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, et al · 2022
Later among the works it cites.
Dota: Detect and omit weak attentions for scalable transformer acceleration
Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2022
Later among the works it cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q Tran, Yi Tay, and Donald Metzler · 2022
Later among the works it cites.
Salo: An efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences
Guan Shen, Jieru Zhao, Quan Chen, Jingwen Leng, Chao Li, and Minyi Guo · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Later among the works it cites.
Speformer: An efficient hardware-software cooperative solution for sparse spectral transformer
Yang Sun, Wei Hu, Fang Liu, Min Jiang, FeiHu Huang, and Dian Xu · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen · 2022
Later among the works it cites.
Searching for burgerformer with micro-meso-macro space design
Longxing Yang, Yu Hu, Shun Lu, Zihao Sun, Jilin Mei, Yinhe Han, and Xiaowei Li · 2022
Later among the works it cites.
Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer architecture
Tao Yang, Hui Ma, Xiaoling Li, Fangxin Liu, Yilong Zhao, Zhezhi He, and Li Jiang · 2022
Later among the works it cites.
Nn-lut: Neural approximation of non-linear operations for efficient transformer inference
Joonsang Yu, Junki Park, Seongmin Park, Minsoo Kim, Sihwa Lee, Dong Hyun Lee, and Jungwook Choi · 2022
Later among the works it cites.
Hessian-aware pruning and optimal neural implant
Shixing Yu, Zhewei Yao, Amir Gholami, Zhen Dong, Sehoon Kim, Michael W Mahoney, and Kurt Keutzer · 2022
Later among the works it cites.
A full-stack search technique for domain optimized deep learning accelerators
Dan Zhang, Safeen Huda, Ebrahim Songhori, Kartik Prabhu, Quoc Le, Anna Goldie, and Azalia Mirhoseini · 2022
Later among the works it cites.
Energon: Towards efficient acceleration of transformers using dynamic sparse attention
Zhe Zhou, Junlin Liu, Zhenyu Gu, and Guangyu Sun · 2022
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper · 2023
Closest in time.
Big little transformer decoder
Sehoon Kim, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, Amir Gholami, and Kurt Keutzer · 2023
Closest in time.