Fetching the paper…
Reading the bibliography…
Adversarial training (AT) is one of the most reliable methods for defending against adversarial attacks in machine learning.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019b · 1901
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 1902
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith. 2019a · 1903
Earlier work this paper cites.
Correlating neural and symbolic representations of language
Grzegorz Chrupała and Afra Alishahi. 2019 · 1905
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. 2019 · 1905
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019a · 1905
Earlier work this paper cites.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019b · 1905
Earlier work this paper cites.
Visualizing and understanding the effectiveness of bert
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 1908
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 1909
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2019 · 1911
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
On the shortest arborescence of a directed graph
Yoeng-Jin Chu. 1965 · 1965
Earlier work this paper cites.
Optimum branchings
Jack Edmonds. 1968 · 1968
Earlier work this paper cites.
Algebraic connectivity of graphs
Miroslav Fiedler. 1973 · 1973
Earlier work this paper cites.
Computers and intractability , volume 174
Michael R Garey and David S Johnson. 1979 · 1979
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Eigenvalues and the max-cut problem
Bojan Mohar and Svatopluk Poljak. 1990 · 1990
Earlier work this paper cites.
The laplacian spectrum of graphs
Ortrud R Oellermann and Allen J Schwenk. 1991 · 1991
Earlier work this paper cites.
Fast is better than free: Revisiting adversarial training
Eric Wong, Leslie Rice, and J Zico Kolter. 2020 · 2001
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2003
Earlier work this paper cites.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Schlichtkrull, Wilker Aziz, and Ivan Titov. 2020 · 2004
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. 2020 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting bert
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2004
Cited alongside, same era.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2020 · 2005
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Cited alongside, same era.
Non-projective dependency parsing using spanning tree algorithms
Ryan McDonald, Fernando Pereira, Kiril Ribarov, and Jan Hajic. 2005 · 2005
Cited alongside, same era.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Adversarial training for relation extraction
Yi Wu, David Bamman, and Stuart Russell. 2017 · 2017
Later among the works it cites.
Robust multilingual part-of-speech tagging via adversarial training
Michihiro Yasunaga, Jungo Kasai, and Dragomir Radev. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bo Pang and Lillian Lee. 2005 · 2005
Cited alongside, same era.
Adversarial training for commonsense inference
Lis Pereira, Xiaodong Liu, Fei Cheng, Masayuki Asahara, and Ichiro Kobayashi. 2020 · 2005
Cited alongside, same era.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020 · 2006
Cited alongside, same era.
Revisiting few-sample bert fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2006
Cited alongside, same era.
Do adversarially robust imagenet models transfer better?
Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor, and Aleksander Madry. 2020 · 2007
Cited alongside, same era.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta. 2020a · 2008
Cited alongside, same era.
Posterior differential regularization with f-divergence for improving model robustness
Hao Cheng, Xiaodong Liu, Lis Pereira, Yaoliang Yu, and Jianfeng Gao. 2020 · 2010
Cited alongside, same era.
Pareto probing: Trading off accuracy for complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020 · 2010
Cited alongside, same era.
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis. 2018 · 2018
Later among the works it cites.
Do latent tree learning models identify meaningful structure in sentences?
Adina Williams, Samuel R Bowman, et al. 2018 · 2018
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Later among the works it cites.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Adversarial learning guarantees for linear hypotheses and neural networks
Pranjal Awasthi, Natalie Frank, and Mehryar Mohri. 2020 · 2020
Later among the works it cites.
Seqvat: Virtual adversarial training for semi-supervised sequence labeling
Luoxin Chen, Weitong Ruan, Xinyue Liu, and Jianhua Lu. 2020 · 2020
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2020 · 2020
Later among the works it cites.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020 · 2020
Later among the works it cites.
On the proper role of linguistically-oriented deep net analysis in linguistic theorizing
Marco Baroni. 2021 · 2021
Closest in time.
Low-complexity probing via finding subnetworks
Steven Cao, Victor Sanh, and Alexander Rush. 2021 · 2021
Closest in time.
Bert & family eat word salad: Experiments with text understanding
Ashim Gupta, Giorgi Kvernadze, and Vivek Srikumar. 2021 · 2021
Closest in time.
The low-rank simplicity bias in deep networks
Minyoung Huh, Hossein Mobahi, Richard Zhang, Brian Cheung, Pulkit Agrawal, and Phillip Isola. 2021 · 2021
Closest in time.
Is sparse attention more interpretable?
Clara Meister, Stefan Lazov, Isabelle Augenstein, and Ryan Cotterell. 2021 · 2021
Closest in time.
The rediscovery hypothesis: Language models need to meet linguistics
Vassilina Nikoulina, Maxat Tezekbayev, Nuradil Kozhakhmet, Madina Babazhanova, Matthias Gallé, and Zhenisbek Assylbekov. 2021 · 2021
Closest in time.
Targeted adversarial training for natural language understanding
Lis Pereira, Xiaodong Liu, Hao Cheng, Hoifung Poon, Jianfeng Gao, and Ichiro Kobayashi. 2021 · 2021
Closest in time.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Closest in time.
On the inductive bias of masked language modeling: From statistical to syntactic dependencies
Tianyi Zhang and Tatsunori Hashimoto. 2021 · 2021
Closest in time.