Fetching the paper…
Reading the bibliography…
Averaging the parameters of models that have the same architecture and initialization can provide a means of combining their respective capabilities.
Where is the information in a deep neural network?
Achille, A., Paolini, G., and Soatto, S · 1905
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
Fisher, R. A · 1922
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Actively searching for an effective neural network ensemble
Opitz, D. W. and Shavlik, J. W · 1996
Earlier work this paper cites.
Neural learning in structured parameter spaces-natural riemannian gradient
Amari, S · 1997
Earlier work this paper cites.
Specification uncertainty and model averaging
Bartels, L. M · 1997
Earlier work this paper cites.
Model selection and model averaging in phylogenetics: advantages of akaike information criterion and bayesian approaches over likelihood ratio tests
Posada, D. and Buckley, T. R · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Sided and symmetrized bregman centroids
Nielsen, F. and Nock, R · 2009
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q · 2009
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Cited alongside, same era.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Later among the works it cites.
A surprisingly robust trick for winograd schema challenge
Kocijan, V., Cretu, A.-M., Camburu, O.-M., Yordanov, Y., and Lukasiewicz, T · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
S2orc: The semantic scholar open research corpus
Lo, K., Wang, L. L., Neumann, M., Kinney, R., and Weld, D. S · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grosse, R. and Martens, J · 2016
Cited alongside, same era.
Chemprot-3.0: a global chemical biology diseases mapping
Kringelum, J., Kjaerulff, S. K., Brunak, S., Lund, O., Oprea, T. I., and Taboureau, O · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Cited alongside, same era.
First quora dataset release: Question pairs, 2017
Iyer, S., Dandekar, N., and Csernai, K · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Cited alongside, same era.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Later among the works it cites.
Transfer learning in natural language processing
Ruder, S., Peters, M. E., Swayamdipta, S., and Wolf, T · 2019
Later among the works it cites.
Learning robust global representations by penalizing local predictive power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
An elementary introduction to information geometry
Nielsen, F · 2020
Later among the works it cites.
English intermediate-task training improves zero-shot cross-lingual transfer too
Phang, J., Htut, P. M., Pruksachatkun, Y., Liu, H., Vania, C., Kann, K., Calixto, I., and Bowman, S. R · 2020
Later among the works it cites.
Pruksachatkun, Y., Phang, J., Liu, H., Htut, P. M., Zhang, X., Pang, R. Y., Vania, C., Kann, K., and Bowman, S. R · 2020
Later among the works it cites.
Exploring and predicting transferability across nlp tasks
Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M · 2020
Later among the works it cites.
Federated learning with matched averaging
Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., and Khazaeni, Y · 2020
Later among the works it cites.
Laplace redux-effortless bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P · 2021
Closest in time.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Li, M., Kim, J. W., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L · 2021
Closest in time.
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L · 2022
Closest in time.