Fetching the paper…
Reading the bibliography…
The popularity of deep learning has led to the curation of a vast number of massive and multifarious datasets.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Discriminability-based transfer between neural networks
Lorien Y Pratt · 1992
Earlier work this paper cites.
Descent approaches for quadratic bilevel programming
Luis Vicente, Gilles Savard, and Joaquim Júdice · 1994
Earlier work this paper cites.
No free lunch theorems for optimization
D.H. Wolpert and W.G. Macready · 1997
Earlier work this paper cites.
An optimal algorithm for approximate nearest neighbor searching fixed dimensions
Sunil Arya, David M Mount, Nathan S Netanyahu, Ruth Silverman, and Angela Y Wu · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
The EM algorithm and extensions
Geoffrey J McLachlan and Thriyambakam Krishnan · 2007
Earlier work this paper cites.
Differential privacy: A survey of results
Cynthia Dwork · 2008
Earlier work this paper cites.
Herding dynamical weights to learn
Max Welling · 2009
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Accelerating the super-resolution convolutional neural network
Chao Dong, Chen Change Loy, and Xiaoou Tang · 2016
Earlier work this paper cites.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konečnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik · 2016
Earlier work this paper cites.
Recommendations as treatments: Debiasing learning and evaluation
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims · 2016
Earlier work this paper cites.
Practical coreset constructions for machine learning
Olivier Bachem, Mario Lucic, and Andreas Krause · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
Meta-sgd: Learning to learn quickly for few-shot learning
Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Earlier work this paper cites.
Fair and diverse dpp-based data summarization
Elisa Celis, Vijay Keswani, Damian Straszak, Amit Deshpande, Tarun Kathuria, and Nisheeth Vishnoi · 2018
Earlier work this paper cites.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Earlier work this paper cites.
Dynamic few-shot visual learning without forgetting
Spyros Gidaris and Nikos Komodakis · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Self-attentive sequential recommendation
Wang-Cheng Kang and Julian McAuley · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros · 2018
Earlier work this paper cites.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Earlier work this paper cites.
Stock movement prediction from tweets and historical prices
Yumo Xu and Shay B Cohen · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Earlier work this paper cites.
Graph neural networks for social recommendation
Wenqi Fan, Yao Ma, Qing Li, Yuan He, Eric Zhao, Jiliang Tang, and Dawei Yin · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
DARTS: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Earlier work this paper cites.
Understanding and correcting pathologies in the training of learned optimizers
Luke Metz, Niru Maheswaranathan, Jeremy Nixon, Daniel Freeman, and Jascha Sohl-Dickstein · 2019
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy · 2019
Cited alongside, same era.
Continual lifelong learning with neural networks: A review
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter · 2019
Cited alongside, same era.
Sequential variational autoencoders for collaborative filtering
Noveen Sachdeva, Giuseppe Manco, Ettore Ritacco, and Vikram Pudi · 2019
Cited alongside, same era.
Embarrassingly shallow autoencoders for sparse data
Harald Steck · 2019
Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies
Paul Vicol, Luke Metz, and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi · 2021
Later among the works it cites.
Condensed composite memory continual learning
Felix Wiewel and Bin Yang · 2021
Later among the works it cites.
Dataset condensation with differentiable siamese augmentation
Bo Zhao and Hakan Bilen · 2021
Later among the works it cites.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2021
Later among the works it cites.
If influence functions are the answer, then what is the question?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon · 2019
Cited alongside, same era.
Session-based recommendation with graph neural networks
Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan · 2019
Cited alongside, same era.
The adversarial robustness of sampling
Omri Ben-Eliezer and Eylon Yogev · 2020
Cited alongside, same era.
Flexible dataset distillation: Learn labels instead of images
Ondrej Bohdal, Yongxin Yang, and Timothy Hospedales · 2020
Cited alongside, same era.
Coresets via bilevel optimization for continual learning and streaming
Zalán Borsos, Mojmir Mutny, and Andreas Krause · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Spagnn: Spatially-aware graph neural networks for relational behavior forecasting from sensor data
Sergio Casas, Cole Gulino, Renjie Liao, and Raquel Urtasun · 2020
Cited alongside, same era.
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger Grosse · 2022
Later among the works it cites.
Submodularity in machine learning and artificial intelligence
Jeff Bilmes · 2022
Later among the works it cites.
Dataset distillation by matching training trajectories
George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu · 2022
Later among the works it cites.
Private set generation with discriminative information
Dingfan Chen, Raouf Kerkouche, and Mario Fritz · 2022
Later among the works it cites.
Remember the past: Distilling datasets into addressable memories for neural networks
Zhiwei Deng and Olga Russakovsky · 2022
Later among the works it cites.
Privacy for free: How does dataset condensation help privacy?
Tian Dong, Bo Zhao, and Lingjuan Lyu · 2022
Later among the works it cites.
Cramming: Training a language model on a single gpu in one day, 2022
Jonas Geiping and Tom Goldstein · 2022
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning
Chengcheng Guo, Bo Zhao, and Yanbing Bai · 2022
Later among the works it cites.
Chasing carbon: The elusive environmental footprint of computing
Udit Gupta, Young Geun Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin S Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Fedsynth: Gradient compression via synthetic data in federated learning
Shengyuan Hu, Jack Goetz, Kshitiz Malik, Hongyuan Zhan, Zhe Liu, and Yue Liu · 2022
Later among the works it cites.
Dataset condensation via efficient synthetic-data parameterization
Jang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun, Hwanjun Song, Joonhyun Jeong, Jung-Woo Ha, and Hyun Oh Song · 2022
Later among the works it cites.
Compressed gastric image generation based on soft-label dataset distillation for medical data sharing
Guang Li, Ren Togo, Takahiro Ogawa, and Miki Haseyama · 2022
Later among the works it cites.
Efficient dataset distillation using random feature approximation
Noel Loo, Ramin Hasani, Alexander Amini, and Daniela Rus · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Gapformer: Fast autoregressive transformers meet rnns for personalized adaptive cruise control
Noveen Sachdeva, Ziran Wang, Kyungtae Han, Rohit Gupta, and Julian McAuley · 2022
Later among the works it cites.
Sample condensation in online continual learning
Mattia Sangermano, Antonio Carta, Andrea Cossu, and Davide Bacciu · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2022
Later among the works it cites.
When less is more: Simplifying inputs aids neural network understanding
Robin Tibor Schirrmeister, Rosanne Liu, Sara Hooker, and Tonio Ball · 2022
Later among the works it cites.
Federated learning via decentralized dataset distillation in resource-constrained edge environments
Rui Song, Dai Liu, Dave Zhenyu Chen, Andreas Festag, Carsten Trinitis, Martin Schulz, and Alois Knoll · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
On implicit bias in overparameterized bilevel optimization
Paul Vicol, Jonathan P Lorraine, Fabian Pedregosa, David Duvenaud, and Roger B Grosse · 2022
Later among the works it cites.
Cafe: Learning to condense dataset by aligning features
Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You · 2022
Later among the works it cites.
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Later among the works it cites.
Feddm: Iterative distribution matching for communication-efficient federated learning
Yuanhao Xiong, Ruochen Wang, Minhao Cheng, Felix Yu, and Cho-Jui Hsieh · 2022
Later among the works it cites.
Synthesizing informative training samples with gan
Bo Zhao and Hakan Bilen · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos · 2023
Closest in time.
Data pruning and neural scaling laws: fundamental limitations of score-based algorithms
Fadhel Ayed and Soufiane Hayou · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2023
Closest in time.