Fetching the paper…
Reading the bibliography…
Fine-tuning large-scale pretrained models has led to tremendous progress in well-studied modalities such as vision and NLP.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Dolan, E. D. and Moré, J. J · 2002
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q · 2009
Earlier work this paper cites.
Fast and robust earth mover’s distances
Pele, O. and Werman, M · 2009
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M · 2013
Earlier work this paper cites.
Openml: networked science in machine learning
Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
Zhou, J. and Troyanskaya, O. G · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Shi, W., Caballero, J., Huszár, F., Totz, J., Aitken, A. P., Bishop, R., Rueckert, D., and Wang, Z · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y · 2017
Earlier work this paper cites.
Catboost: unbiased boosting with categorical features
Ostroumova, L., Gusev, G., Vorobev, A., Dorogush, A. V., and Gulin, A · 2017
Earlier work this paper cites.
Spherical cnns
Cohen, T., Geiger, M., Köhler, J., and Welling, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A. and Narasimhan, K · 2018
Earlier work this paper cites.
Deep visual domain adaptation: A survey
Wang, M. and Deng, W · 2018
Earlier work this paper cites.
DEEPCON: protein contact prediction using dilated convolutional neural networks with dropout
Adhikari, B · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Supervised multimodal bitransformers for classifying images and text
Kiela, D., Bhooshan, S., Firooz, H., and Testuggine, D · 2019
Earlier work this paper cites.
Auto-deeplab: Hierarchical neural architecture search for semantic image segmentation
Liu, C., Chen, L.-C., Schroff, F., Adam, H., Hua, W., Yuille, A. L., and Fei-Fei, L · 2019
Earlier work this paper cites.
Moment matching for multi-source domain adaptation
Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B · 2019
Earlier work this paper cites.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Raissi, M., Perdikaris, P., and Karniadakis, G. E · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Earlier work this paper cites.
Heterogeneous domain adaptation via soft transfer network
Yao, Y., Zhang, Y., Li, X., and Ye, Y · 2019
Earlier work this paper cites.
Geometric dataset distances via optimal transport
Alvarez-Melis, D. and Fusi, N · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., rahman Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Rocket: exceptionally fast and accurate time series classification using random convolutional kernels
Dempster, A., Petitjean, F., and Webb, G. I · 2020
Cited alongside, same era.
Autogluon-tabular: Robust and accurate automl for structured data
Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., and Smola, A · 2020
Cited alongside, same era.
Rethinking neural operations for diverse tasks
Roberts, N. C., Khodak, M., Dao, T., Li, L., Re, C., and Talwalkar, A · 2021
Later among the works it cites.
Don’t sweep your learning rate under the rug: A closer look at cross-modal transfer of pretrained transformers
Rothermel, D., Li, M., Rocktaschel, T., and Foerster, J. N · 2021
Later among the works it cites.
Converting tabular data into images for deep learning with convolutional neural networks
Zhu, Y., Brettin, T. S., Xia, F., Partin, A., Shukla, M., Yoo, H. S., Evrard, Y. A., Doroshow, J. H., and Stevens, R. L · 2021
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R. C., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D. M., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fang, J., Sun, Y., Zhang, Q., Li, Y., Liu, W., and Wang, X · 2020
Cited alongside, same era.
Holmes: Health online model ensemble serving for deep learning models in intensive care units
Hong, S., Xu, Y., Khare, A., Priambada, S., Maher, K. O., Aljiffry, A., Sun, J., and Tumanov, A · 2020
Cited alongside, same era.
Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization
Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Zhao, T · 2020
Cited alongside, same era.
semg gesture recognition with a simple model of attention
Josephs, D., Drake, C., Heroy, A. M., and Santerre, J · 2020
Cited alongside, same era.
Automl-zero: Evolving machine learning algorithms from scratch
Real, E., Liang, C., So, D. R., and Le, Q. V · 2020
Cited alongside, same era.
Class-imbalanced domain adaptation: An empirical odyssey
Tan, S., Peng, X., and Saenko, K · 2020
Cited alongside, same era.
deepcr: Cosmic ray rejection with deep learning
Zhang, K. and Bloom, J. S · 2020
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Later among the works it cites.
Exploring visual prompts for adapting large-scale models
Bahng, H., Jahanian, A., Sankaranarayanan, S., and Isola, P · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., et al · 2022
Later among the works it cites.
Lift: Language-interfaced fine-tuning for non-language machine learning tasks
Dinh, T., Zeng, Y., Zhang, R., Lin, Z., Rajput, S., Gira, M., yong Sohn, J., Papailiopoulos, D., and Lee, K · 2022
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2022
Later among the works it cites.
Tabpfn: A transformer that solves small tabular classification problems in a second
Hollmann, N., Muller, S., Eggensperger, K., and Hutter, F · 2022
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O. J., Botvinick, M., Zisserman, A., Vinyals, O., and Carreira, J · 2022
Later among the works it cites.
Visual prompt tuning
Jia, M., Tang, L., Chen, B.-C., Cardie, C., Belongie, S. J., Hariharan, B., and Lim, S. N · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Later among the works it cites.
Surgical fine-tuning improves adaptation to distribution shifts
Lee, Y., Chen, A. S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C · 2022
Later among the works it cites.
Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Li, F., Zhang, H., Xu, H.-S., Liu, S., Zhang, L., Ni, L. M., and yeung Shum, H · 2022
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2022
Later among the works it cites.
Frozen pretrained transformers as universal computation engines
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2022
Later among the works it cites.
Can wikipedia help offline reinforcement learning?
Reid, M., Yamada, Y., and Gu, S. S · 2022
Later among the works it cites.
Efficient architecture search for diverse tasks
Shen, J., Khodak, M., and Talwalkar, A · 2022
Later among the works it cites.
Pdebench: An extensive benchmark for scientific machine learning
Takamoto, M., Praditia, T., Leiteritz, R., MacKinlay, D., Alesiani, F., Pflüger, D., and Niepert, M · 2022
Later among the works it cites.
NAS-bench-360: Benchmarking neural architecture search on diverse tasks
Tu, R., Roberts, N., Khodak, M., Shen, J., Sala, F., and Talwalkar, A · 2022
Later among the works it cites.
Contrastive learning rivals masked image modeling in fine-tuning via feature distillation
Wei, Y., Hu, H., Xie, Z., Zhang, Z., Cao, Y., Bao, J., Chen, D., and Guo, B · 2022
Later among the works it cites.
A generalist agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A. D., Heess, N. M. O., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2023
Closest in time.
Reprogramming pretrained language models for protein sequence representation learning
Vinod, R., Chen, P.-Y., and Das, P · 2023
Closest in time.