Fetching the paper…
Reading the bibliography…
The advent of large-scale pre-trained language models has contributed greatly to the recent progress in natural language processing.
1909
Earlier work this paper cites.
B. W. Matthews, “Comparison of the predicted and observed secondary structure of t4 phage lysozyme,” Biochimica et Biophysica Acta (BBA)-Protein Structure , vol. 405, no. 2, pp. 442–451, 1975
1975
Earlier work this paper cites.
A. N. Tikhonov and V. Y. Arsenin, “Solutions of ill-posed problems,” New York , vol. 1, no. 30, p. 487, 1977
1977
Earlier work this paper cites.
J. Sietsma and R. J. Dow, “Creating artificial neural networks that generalize,” Neural networks , vol. 4, no. 1, pp. 67–79, 1991
1991
Earlier work this paper cites.
C. M. Bishop, “Training with noise is equivalent to tikhonov regularization,” Neural computation , vol. 7, no. 1, pp. 108–116, 1995
1995
Earlier work this paper cites.
H. Federer et al. , “Geometric measure theory,” 1996
1996
Earlier work this paper cites.
S. T. Roweis and L. K. Saul, “Nonlinear dimensionality reduction by locally linear embedding,” science , vol. 290, no. 5500, pp. 2323–2326, 2000
2000
Earlier work this paper cites.
L. Cayton, “Algorithms for manifold learning,” Univ. of California at San Diego Tech. Rep , vol. 12, no. 1-17, p. 1, 2005
2005
Earlier work this paper cites.
W. Dolan and C. Brockett, “Automatically constructing a corpus of sentential paraphrases,” in IWP@IJCNLP , 2005
2005
Earlier work this paper cites.
I. Dagan, O. Glickman, and B. Magnini, “The pascal recognising textual entailment challenge,” in MLCW , 2005
2005
Earlier work this paper cites.
X. Huo, X. S. Ni, and A. K. Smith, “A survey of manifold-based learning methods,” Recent advances in data mining of enterprise data , pp. 691–745, 2007
2007
Earlier work this paper cites.
D. Giampiccolo, B. Magnini, I. Dagan, and W. Dolan, “The third pascal recognizing textual entailment challenge,” in ACL-PASCAL@ACL , 2007
2007
Earlier work this paper cites.
T. Lin and H. Zha, “Riemannian manifold learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 30, no. 5, pp. 796–809, 2008
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
D. Erhan, P.-A. Manzagol, Y. Bengio, S. Bengio, and P. Vincent, “The difficulty of training deep architectures and the effect of unsupervised pre-training,” in AISTATS , 2009
2009
Earlier work this paper cites.
D. Erhan, A. C. Courville, Y. Bengio, and P. Vincent, “Why does unsupervised pre-training help deep learning?” J. Mach. Learn. Res. , vol. 11, pp. 625–660, 2010
2010
Earlier work this paper cites.
2011
Earlier work this paper cites.
L. Wan, M. D. Zeiler, S. Zhang, Y. LeCun, and R. Fergus, “Regularization of neural networks using dropconnect,” in ICML , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” in 2nd International Conference on Learning Representations, ICLR 2014 , 2014
2014
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP , 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. M. Dai and Q. V. Le, “Semi-supervised sequence learning,” in NIPS , 2015
2015
Earlier work this paper cites.
G. Tsatsaronis, G. Balikas, P. Malakasiotis, I. Partalas, M. Zschunke, M. R. Alvers, D. Weissenborn, A. Krithara, S. Petridis, D. Polychronopoulos et al. , “An overview of the bioasq large-scale biomedical semantic indexing and question answering competition,” BMC bioinformatics , vol. 16, no. 1, pp. 1–28, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Z. Wen, S.-C. Fuh, and A. Mircea, “Neurips 2019 reproducibility challenge: Controllable unsupervised text attribute transfer via editing entangled latent representation,” 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:211094926
2019
Later among the works it cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi, “Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern recognition , 2017
2017
Cited alongside, same era.
X. Li, Y. Grandvalet, and F. Davoine, “Explicit inductive bias for transfer learning with convolutional networks,” in ICML , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang, “Stronger generalization bounds for deep nets via a compression approach,” in International Conference on Machine Learning . PMLR, 2018, pp. 254–263
2018
Cited alongside, same era.
G. Zhang, C. Wang, B. Xu, and R. Grosse, “Three mechanisms of weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Cited alongside, same era.
C. Zhu, Y. Cheng, Z. Gan, S. Sun, T. Goldstein, and J. jing Liu, “Freelb: Enhanced adversarial training for natural language understanding,” arXiv: Computation and Language , 2020
2020
Later among the works it cites.
H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao, “Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization,” in ACL , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “Spanbert: Improving pre-training by representing and predicting spans,” Transactions of the Association for Computational Linguistics , 2020
2020
Later among the works it cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020. [Online]. Available: http://jmlr.org/papers/v21/20-074.html
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Cai, J. Huang, Y. Bian, and K. Church, “Isotropy in the contextual embedding space: Clusters and manifolds,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Dong, A. T. Luu, M. Lin, S. Yan, and H. Zhang, “How should pre-trained language models be fine-tuned towards adversarial robustness?” Advances in Neural Information Processing Systems , vol. 34, pp. 4356–4369, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Zhou, L. Liao, Y. Gao, R. Wang, and H. Huang, “Topicbert: A topic-enhanced neural language model fine-tuned for sentiment classification,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Later among the works it cites.
Y. Ro, J. Choi, B. Heo, and J. Y. Choi, “Rollback ensemble with multiple local minima in fine-tuning deep learning networks,” IEEE transactions on neural networks and learning systems , vol. 33, no. 9, pp. 4648–4660, 2021
2021
Later among the works it cites.
2022
Closest in time.
X. Li, H. Hang, C. Xu, and D. Dou, “Method and apparatus for transfer learning,” Dec. 15 2022, uS Patent App. 17/820,321
2022
Closest in time.
2022
Closest in time.