Meta-dataset: A dataset of datasets for learning to learn from few examples
Original
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, et al · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan C. Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Xiaohua Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, André Susano Pinto, Maxim Neumann, A. Dosovitskiy, Lucas Beyer, Olivier Bachem, M. Tschannen, Marcin Michalski, O. Bousquet, S. Gelly, and N. Houlsby · 2019
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Original
Alexei Baevski, H. Zhou, Abdel rahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben-Zaken, Shauli Ravfogel, and Yoav Goldberg · 2020
Later among the works it cites.
Language models are few-shot learners
Original
T. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, J. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. Henighan, R. Child, A. Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, J. Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Original
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, J. Davis, Tamás Sarlós, David Belanger, Lucy J. Colwell, and Adrian Weller · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Original
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
A. Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, M. Dehghani, Matthias Minderer, G. Heigold, S. Gelly, Jakob Uszkoreit, and N. Houlsby · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of nlp leaderboard design
Original
Kawin Ethayarajh and Dan Jurafsky · 2020
Later among the works it cites.
Training batchnorm and only batchnorm: On the expressive power of random features in cnns
Original
Jonathan Frankle, D. Schwab, and Ari S. Morcos · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Original
Jean-Bastien Grill, Florian Strub, Florent Altch’e, C. Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, B. A. Pires, Zhaohan Daniel Guo, M. G. Azar, Bilal Piot, K. Kavukcuoglu, R. Munos, and Michal Valko · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick · 2020
Later among the works it cites.
Scaling laws for autoregressive generative modeling
Original
T. Henighan, J. Kaplan, Mor Katz, Mark Chen, Christopher Hesse, J. Jackson, Heewoo Jun, T. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, A. Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish · 2020
Later among the works it cites.
Scaling laws for neural language models
Original
J. Kaplan, Sam McCandlish, T. Henighan, T. Brown, Benjamin Chess, R. Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Andy T. Liu, Shuwen Yang, Po-Han Chi, Po-Chun Hsu, and Hung yi Lee · 2020
Later among the works it cites.
Why You Should Do NLP Beyond English
Sebastian Ruder · 2020
Later among the works it cites.
Towards learning a universal non-semantic representation of speech
Original
Joel Shor, A. Jansen, R. Maor, Oran Lang, Omry Tuval, F. D. C. Quitry, M. Tagliasacchi, Ira Shavitt, D. Emanuel, and Yinnon A. Haviv · 2020
Later among the works it cites.
Towards domain-agnostic contrastive learning
Original
Vikas Verma, Minh-Thang Luong, Kenji Kawaguchi, Hieu Pham, and Quoc V. Le · 2020
Later among the works it cites.
Are all languages created equal in multilingual bert?
Shijie Wu and Mark Dredze · 2020
Later among the works it cites.
Graph contrastive learning with augmentations
Original
Yuning You, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen · 2020
Later among the works it cites.
Documenting the english colossal clean crawled corpus
Original
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner · 2021
Closest in time.
Structured dataset documentation: a datasheet for chexpert
Original
Christian Garbin, P. Rajpurkar, Jeremy A. Irvin, M. Lungren, and Oge Marques · 2021
Closest in time.
Ast: Audio spectrogram transformer
Original
Yuan Gong, Yu-An Chung, and J. Glass · 2021
Closest in time.
Tabbie: Pretrained representations of tabular data
Original
H. Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer · 2021
Closest in time.
Perceiver: General perception with iterative attention
Original
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2021
Closest in time.
i-mix: A domain-agnostic strategy for contrastive representation learning
Kibok Lee, Yian Zhu, Kihyuk Sohn, Chun-Liang Li, Jinwoo Shin, and Honglak Lee · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Original
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Original
Xiang Lisa Li and Percy Liang · 2021
Closest in time.
Gpt understands, too
Original
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2021
Closest in time.
C5t5: Controllable generation of organic molecules with transformers
Original
Daniel Rothchild, Alex Tamkin, Julie Yu, Ujval Misra, and Joseph Gonzalez · 2021
Closest in time.
Superb: Speech processing universal performance benchmark
Original
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al · 2021
Closest in time.
Oolong: Investigating what makes crosslingual transfer hard with controlled studies
Original
Zhengxuan Wu, Isabel Papadimitriou, and Alex Tamkin · 2022
Closest in time.
Benchmd: A benchmark for modality-agnostic learning on medical images and sensors
Kathryn Wantlin, Chenwei Wu, Shih-Cheng Huang, Oishi Banerjee, Farah Dadabhoy, Veeral Vipin Mehta, Ryan Wonhee Han, Fang Cao, Raja R. Narayan, Errol Colak, Adewole S. Adamson, Laura Heacock, Geoffrey H. Tison, Alex Tamkin, and Pranav Rajpurkar · 2023
Closest in time.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A. Efros · 2023
Closest in time.