Fetching the paper…
Reading the bibliography…
In recent years, Dynamic Sparse Training (DST) has emerged as an alternative to post-training pruning for generating efficient models.
Solving elliptic problems using ELLPACK
John R Rice and Ronald F Boisvert · 1985
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork · 1992
Earlier work this paper cites.
Enhancing navigation on wikipedia with social tags
Arkaitz Zubiaga · 2012
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Inferring networks of substitutable and complementary products
Julian McAuley, Rahul Pandey, and Jure Leskovec · 2015
Earlier work this paper cites.
The extreme classification repository: Multi-label datasets and code, 2016
K. Bhatia, K. Dahiya, H. Jain, P. Kar, A. Mittal, Y. Prabhu, and M. Varma · 2016
Earlier work this paper cites.
Pd-sparse: A primal and dual sparse approach to extreme multiclass and multilabel classification
Ian En-Hsu Yen, Xiangru Huang, Pradeep Ravikumar, Kai Zhong, and Inderjit Dhillon · 2016
Earlier work this paper cites.
Extreme multi-label loss functions for recommendation, tagging, ranking & other missing label applications
Himanshu Jain, Yashoteja Prabhu, and Manik Varma · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Dismec: Distributed sparse machines for extreme multi-label classification
Rohit Babbar and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Earlier work this paper cites.
Sparse dnns with improved adversarial robustness
Yiwen Guo, Chao Zhang, Changshui Zhang, and Yurong Chen · 2018
Earlier work this paper cites.
Parabel: Partitioned label trees for extreme classification with application to dynamic search advertising
Yashoteja Prabhu, Anil Kag, Shrutendra Harsola, Rahul Agrawal, and Manik Varma · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2018
Earlier work this paper cites.
Attentionxml: Label tree-based attention-aware deep model for high-performance extreme multi-label text classification
Ronghui You, Zihan Zhang, Ziye Wang, Suyang Dai, Hiroshi Mamitsuka, and Shanfeng Zhu · 2019
Earlier work this paper cites.
Data scarcity, robustness and extreme multi-label classification
Rohit Babbar and Bernhard Schölkopf · 2019
Earlier work this paper cites.
Extreme classification in log memory using count-min sketch: A case study of amazon search with 50m products
Tharun Kumar Reddy Medini, Qixuan Huang, Yiqiu Wang, Vijai Mohan, and Anshumali Shrivastava · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V Sanh · 2019
Earlier work this paper cites.
Nest: A neural network synthesis tool based on a grow-and-prune paradigm
Xiaoliang Dai, Hongxu Yin, and Niraj K Jha · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu · 2019
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
Proving the lottery ticket hypothesis: Pruning is all you need
Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Cited alongside, same era.
Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity
Shiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen, Ghada Sokar, Elena Mocanu, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu · 2022
Later among the works it cites.
Gradient flow in sparse neural networks and how lottery tickets win
Utku Evci, Yani Ioannou, Cem Keskin, and Yann Dauphin · 2022
Later among the works it cites.
Cascadexml: Rethinking transformers for end-to-end multi-resolution training in extreme multi-label classification
Siddhant Kharbanda, Atmadeep Banerjee, Erik Schultheis, and Rohit Babbar · 2022
Later among the works it cites.
Speeding-up one-versus-all training for extreme classification via mean-separating initialization
Erik Schultheis and Rohit Babbar · 2022
Later among the works it cites.
On missing labels, long-tails and propensities in extreme multi-label classification
Erik Schultheis, Marek Wydmuch, Rohit Babbar, and Krzysztof Dembczynski · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev, and Martin Jaggi · 2020
Cited alongside, same era.
Sparse gpu kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Cited alongside, same era.
Towards structured dynamic sparse pre-training of bert
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr, Ivan Chelombiev, Daniel Justus, and Carlo Luschi · 2021
Cited alongside, same era.
Training adversarially robust sparse networks via bayesian connectivity sampling
Ozan Özdenizci and Robert Legenstein · 2021
Cited alongside, same era.
A winning hand: Compressing deep networks can improve out-of-distribution robustness
James Diffenderfer, Brian Bartoldson, Shreya Chaganti, Jize Zhang, and Bhavya Kailkhura · 2021
Cited alongside, same era.
Dense for the price of sparse: Improved performance of sparsely initialized networks via a subspace offset
Ilan Price and Jared Tanner · 2021
Cited alongside, same era.
Truly sparse neural networks at scale
Selima Curci, Decebal Constantin Mocanu, and Mykola Pechenizkiyi · 2021
Cited alongside, same era.
Ngame: Negative mining-aware mini-batching for extreme classification
Kunal Dahiya, Nilesh Gupta, Deepak Saini, Akshay Soni, Yajun Wang, Kushal Dave, Jian Jiao, Gururaj K, Prasenjit Dey, Amit Singh, et al · 2023
Later among the works it cites.
Towards memory-efficient training for extremely large output spaces–learning with 670k labels on a single commodity gpu
Erik Schultheis and Rohit Babbar · 2023
Later among the works it cites.
Renee: End-to-end training of extreme classification models
Vidit Jain, Jatin Prakash, Deepak Saini, Jian Jiao, Ramachandran Ramjee, and Manik Varma · 2023
Later among the works it cites.
Sparsity may cry: Let us fail (current) sparse neural networks together!
Shiwei Liu, Tianlong Chen, Zhenyu Zhang, Xuxi Chen, Tianjin Huang, Ajay Jaiswal, and Zhangyang Wang · 2023
Later among the works it cites.
Enhancing tail performance in extreme classifiers by label variance reduction
Anirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal, Mudit Dhawan, Neelabh Madan, Deepesh Hada, Vidit Jain, SONU MEHTA, Yashoteja Prabhu, et al · 2023
Later among the works it cites.
Jaxpruner: A concise library for sparsity research
Joo Hyung Lee, Wonpyo Park, Nicole Elyse Mitchell, Jonathan Pilault, Johan Samir Obando Ceron, Han-Byul Kim, Namhoon Lee, Elias Frantar, Yun Long, Amir Yazdanbakhsh, Woohyun Han, Shivani Agrawal, Suvinay Subramanian, Xin Wang, Sheng-Chun Kao, Xingyao Zhang, Trevor Gale, Aart J.C. Bik, Milen Ferev, Zhonglin Han, Hong-Seok Kim, Yann Dauphin, Gintare Karolina Dziugaite, Pablo Samuel Castro, and Utku Evci · 2023
Later among the works it cites.
Venom: A vectorized n: M format for unleashing the power of sparse tensor cores
Roberto L Castro, Andrei Ivanov, Diego Andrade, Tal Ben-Nun, Basilio B Fraguela, and Torsten Hoefler · 2023
Later among the works it cites.
Shiwei Liu and Zhangyang Wang · 2023
Later among the works it cites.
Uniform sparsity in deep neural networks
Saurav Muralidharan · 2023
Later among the works it cites.
Inceptionxml: A lightweight framework with synchronized negative sampling for short text extreme classification
Siddhant Kharbanda, Atmadeep Banerjee, Devaansh Gupta, Akash Palrecha, and Rohit Babbar · 2023
Later among the works it cites.
Progressive gradient flow for robust n: M sparsity training in transformers
Abhimanyu Rajeshkumar Bambhaniya, Amir Yazdanbakhsh, Suvinay Subramanian, Sheng-Chun Kao, Shivani Agrawal, Utku Evci, and Tushar Krishna · 2024
Closest in time.
Meta-classifier free negative sampling for extreme multilabel classification
Mohammadreza Qaraei and Rohit Babbar · 2024
Closest in time.
Compressing LLMs: The truth is rarely pure and never simple
Ajay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang, Zhangyang Wang, and Yinfei Yang · 2024
Closest in time.
Generalized test utilities for long-tail performance in extreme multi-label classification
Erik Schultheis, Marek Wydmuch, Wojciech Kotlowski, Rohit Babbar, and Krzysztof Dembczynski · 2024
Closest in time.
Dynamic sparse training with structured sparsity
Mike Lasby, Anna Golubeva, Utku Evci, Mihai Nica, and Yani Ioannou · 2024
Closest in time.
How to prune your language model: Recovering accuracy on the “sparsity may cry” benchmark
Eldar Kurtic, Torsten Hoefler, and Dan Alistarh · 2024
Closest in time.
Fantastic weights and how to find them: Where to prune in dynamic sparse training
Aleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, and Jacek Tabor · 2024
Closest in time.