Fetching the paper…
Reading the bibliography…
Transformers achieve state-of-the-art performance for natural language processing tasks by pre-training on large-scale text corpora.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q. V.; and Salakhutdinov, R. 2019 · 1901
Earlier work this paper cites.
Liu, H.; Trott, A.; Socher, R.; and Xiong, C. 2019a · 1902
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Dong, L.; Yang, N.; Wang, W.; Wei, F.; Liu, X.; Wang, Y.; Gao, J.; Zhou, M.; and Hon, H.-W. 2019 · 1905
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
Song, K.; Tan, X.; Qin, T.; Lu, J.; and Liu, T.-Y. 2019 · 1905
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Strubell, E.; Ganesh, A.; and McCallum, A. 2019 · 1906
Earlier work this paper cites.
When to use parametric models in reinforcement learning?
van Hasselt, H.; Hessel, M.; and Aslanides, J. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2019 · 1907
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W.; Potapenko, A.; Jayakumar, S. M.; and Lillicrap, T. P. 2019 · 1911
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M.; and Cohen, N. J. 1989 · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. 1992 · 1992
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N.; Kaiser, Ł.; and Levskaya, A. 2020 · 2001
Earlier work this paper cites.
PAC bounds for multi-armed bandit and Markov decision processes
Even-Dar, E.; Mannor, S.; and Mansour, Y. 2002 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K.; Luong, M.-T.; Le, Q. V.; and Manning, C. D. 2020 · 2003
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2005
Earlier work this paper cites.
Pure exploration in multi-armed bandits problems
Bubeck, S.; Munos, R.; and Stoltz, G. 2009 · 2009
Cited alongside, same era.
Efficient transformers: A survey
Tay, Y.; Dehghani, M.; Bahri, D.; and Metzler, D. 2020b · 2009
Cited alongside, same era.
Fastformers: Highly efficient transformer models for natural language understanding
Kim, Y. J.; and Awadalla, H. H. 2020 · 2010
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W.; Li, L.; Reyzin, L.; and Schapire, R. 2011 · 2011
Cited alongside, same era.
Efficient per-example gradient computations
Goodfellow, I. 2015 · 2015
Cited alongside, same era.
Huang, C.-Z. A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A. M.; Hoffman, M. D.; Dinculescu, M.; and Eck, D. 2018 · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A.; and Fleuret, F. 2018 · 2018
Later among the works it cites.
The effects of memory replay in reinforcement learning
Liu, R.; and Zou, J. 2018 · 2018
Later among the works it cites.
Image transformer
Parmar, N.; Vaswani, A.; Uszkoreit, J.; Kaiser, L.; Shazeer, N.; Ku, A.; and Tran, D. 2018 · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
Schaul, T.; Quan, J.; Antonoglou, I.; and Silver, D. 2015 · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, P.; and Zhang, T. 2015 · 2015
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Wang, Z.; Bapst, V.; Heess, N.; Mnih, V.; Munos, R.; Kavukcuoglu, K.; and de Freitas, N. 2016 · 2016
Cited alongside, same era.
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, P.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L.; Jamieson, K.; DeSalvo, G.; Rostamizadeh, A.; and Talwalkar, A. 2017 · 2017
Cited alongside, same era.
Memory replay gans: Learning to generate new categories without forgetting
Wu, C.; Herranz, L.; Liu, X.; van de Weijer, J.; Raducanu, B.; et al. 2018 · 2018
Later among the works it cites.
Video action transformer network
Girdhar, R.; Carreira, J.; Doersch, C.; and Zisserman, A. 2019 · 2019
Later among the works it cites.
A transformer model for retrosynthesis
Karpov, P.; Godin, G.; and Tetko, I. V. 2019 · 2019
Later among the works it cites.
A bandit approach to maximum inner product search
Liu, R.; Wu, T.; and Mozafari, B. 2019 · 2019
Later among the works it cites.
Revisiting fundamentals of experience replay
Fedus, W.; Ramachandran, P.; Agarwal, R.; Bengio, Y.; Larochelle, H.; Rowland, M.; and Dabney, W. 2020 · 2020
Later among the works it cites.
A Method for Optimizing Opaque Filter Queries
He, W.; Anderson, M. R.; Strome, M.; and Cafarella, M. 2020 · 2020
Later among the works it cites.
Pop Music Transformer: Beat-based modeling and generation of expressive Pop piano compositions
Huang, Y.-S.; and Yang, Y.-H. 2020 · 2020
Later among the works it cites.
Adam with bandit sampling for deep learning
Liu, R.; Wu, T.; and Mozafari, B. 2020 · 2020
Later among the works it cites.
Dynamic experience replay
Luo, J.; and Li, H. 2020 · 2020
Later among the works it cites.
Attentive experience replay
Sun, P.; Zhou, W.; and Li, H. 2020 · 2020
Later among the works it cites.
Diagnosing bottlenecks in deep q-learning algorithms
Fu, J.; Kumar, A.; Soh, M.; and Levine, S. 2019 · 2030
Closest in time.