Fetching the paper…
Reading the bibliography…
Research on neural networks has focused on understanding a single model trained on a single dataset.
On the algebraic structure of feedforward network weight spaces
Hecht-Nielsen, R · 1990
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
Chen, A. M., Lu, H.-m., and Hecht-Nielsen, R · 1993
Earlier work this paper cites.
Evidence for universality and cultural variation of differential emotion response patterning
Scherer, K. R. and Wallbott, H. G · 1994
Earlier work this paper cites.
Learning question classifiers
Li, X. and Roth, D · 2002
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Pang, B. and Lee, L · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., and Magnini, B · 2006
Earlier work this paper cites.
SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter
Basile, V., Bosco, C., Fersini, E., Nozza, D., Patti, V., Rangel Pardo, F. M., Rosso, P., and Sanguinetti, M · 2007
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, W. B · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
The sixth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
The winograd schema challenge
Levesque, H. J., Davis, E., and Morgenstern, L · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
The Winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Malo, P., Sinha, A., Korhonen, P., Wallenius, J., and Takala, P · 2014
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
He, R. and McAuley, J · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D. M., Diab, M. T., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Cited alongside, same era.
Emotion intensities in tweets
Mohammad, S. M. and Bravo-Marquez, F · 2017
Cited alongside, same era.
SemEval-2018 Task 2: Multilingual Emoji Prediction
Barbieri, F., Camacho-Collados, J., Ronzano, F., Espinosa-Anke, L., Ballesteros, M., Basile, V., Patti, V., and Saggion, H · 2018
Cited alongside, same era.
e-snli: Natural language inference with natural language explanations
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Parameter norm growth during training of transformers
Merrill, W., Ramanujan, V., Goldberg, Y., Schwartz, R., and Smith, N. A · 2020
Later among the works it cites.
Linear mode connectivity in multitask and continual learning
Mirzadeh, S. I., Farajtabar, M., Gorur, D., Pascanu, R., and Ghasemzadeh, H · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2020
Later among the works it cites.
Investigating societal biases in a poetry composition system
Sheng, E. and Uthus, D · 2020
Later among the works it cites.
Loss surface simplexes for mode connecting volumes and fast ensembling
Benton, G., Maddox, W., Lotfi, S., and Wilson, A. G. G · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Cited alongside, same era.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Phang, J., Févry, T., and Bowman, S. R · 2018
Cited alongside, same era.
Collecting diverse natural language inference problems for sentence representation evaluation
Poliak, A., Haldar, A., Rudinger, R., Hu, J. E., Pavlick, E., White, A. S., and Durme, B. V · 2018
Cited alongside, same era.
SemEval-2018 task 3: Irony detection in English tweets
Van Hee, C., Lefever, E., and Hoste, V · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Cited alongside, same era.
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2021
Later among the works it cites.
Merging models with fisher-weighted averaging
Matena, M. and Raffel, C · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Ben Zaken, E., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning
Don-Yehiya, S., Venezian, E., Raffel, C., Slonim, N., Katz, Y., and Choshen, L · 2022
Later among the works it cites.
Measuring causal effects of data statistics on language model’s ‘factual’ predictions
Elazar, Y., Kassner, N., Ravfogel, S., Feder, A., Ravichander, A., Mosbach, M., Belinkov, Y., Schütze, H., and Goldberg, Y · 2022
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Later among the works it cites.
Repair: Renormalizing permuted activations for interpolation repair
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., and Neyshabur, B · 2022
Later among the works it cites.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Later among the works it cites.
Improving generalization of pre-trained language models via stochastic weight averaging
Lu, P., Kobyzev, I., Rezagholizadeh, M., Rashid, A., Ghodsi, A., and Langlais, P · 2022
Later among the works it cites.
Exploring mode connectivity for pre-trained language models
Qin, Y., Qian, C., Yi, J., Chen, W., Lin, Y., Han, X., Liu, Z., Sun, M., and Zhou, J · 2022
Later among the works it cites.
Recycling diverse models for out-of-distribution generalization
Ram’e, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2022
Later among the works it cites.
Revisiting sequential information bottleneck: New implementation and evaluation
Toledo, A., Venezian, E., and Slonim, N · 2022
Later among the works it cites.
Uncertainty-aware natural language inference with stochastic weight averaging
Talman, A., Celikkanat, H., Virpioja, S., Heinonen, M., and Tiedemann, J · 2023
Closest in time.
Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Closest in time.
SemEval-2017 task 4: Sentiment analysis in Twitter
Rosenthal, S., Farra, N., and Nakov, P · 2088
Closest in time.