Fetching the paper…
Reading the bibliography…
Statistical natural language inference (NLI) models are susceptible to learning dataset bias: superficial cues that happen to associate with the label on a particular dataset, but are not useful in general, e.g., negation words indicate contradiction.
Theoretically principled trade-off between robustness and accuracy
H. Zhang, Y. Yu, J. Jiao, E. P. Xing, L. E. Ghaoui, and M. I. Jordan. 2019b · 1901
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R. T. McCoy, E. Pavlick, and T. Linzen. 2019 · 1902
Earlier work this paper cites.
H. Gonen and Y. Goldberg. 2019 · 1903
Earlier work this paper cites.
WINOGRANDE: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi. 2019 · 1907
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
H. Shimodaira. 2000 · 2000
Earlier work this paper cites.
Analysis of representations for domain adaptation
S. Ben-David, J. Blitzer, K. Crammer, and F. Pereira. 2006 · 2006
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. Efros. 2011 · 2011
Earlier work this paper cites.
On causal and anticausal learning
B. Scholkopf, D. Janzing, J. Peters, E. Sgouritsa, K. Zhang, and J. Mooij. 2012 · 2012
Earlier work this paper cites.
Domain adaptation under target and conditional shift
K. Zhang, B. Schölkopf, K. Muandet, and Z. Wang. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning. 2015 · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
T. Bolukbasi, K. Chang, J. Y. Zou, V. Saligrama, and A. T. Kalai. 2016 · 2016
Cited alongside, same era.
Natural language inference by tree-based convolution and heuristic matching
L. Mou, R. Men, G. Li, Y. Xu, L. Zhang, R. Yan, and Z. Jin. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit. 2016 · 2016
Cited alongside, same era.
Enhanced LSTM for natural language inference
Q. Chen, X. Zhu, Z. Ling, S. Wei, H. Jiang, and D. Inkpen. 2017 · 2017
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
R. Jia and P. Liang. 2017 · 2017
Cited alongside, same era.
The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task
Stress test evaluation for natural language inference
A. Naik, A. Ravichander, N. Sadeh, C. Rose, and G. Neubig. 2018 · 2018
Later among the works it cites.
Hypothesis only baselines in natural language inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. V. Durme. 2018 · 2018
Later among the works it cites.
Semantically equivalent adversarial rules for debugging NLP models
M. T. Ribeiro, S. Singh, and C. Guestrin. 2018 · 2018
Later among the works it cites.
Good-enough compositional data augmentation
J. Andreas. 2019 · 2019
Closest in time.
Don’t take the premise for granted: Mitigating artifacts in natural language inference
Y. Belinkov, A. Poliak, S. M. Shieber, B. V. Durme, and A. M. Rush. 2019 · 2019
Closest in time.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Schwartz, M. Sap, Y. Konstas, L. Zilles, Y. Choi, and N. A. Smith. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. R. Bowman. 2017 · 2017
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
J. Duchi and H. Namkoong. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
M. Glockner, V. Shwartz, and Y. Goldberg. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith. 2018 · 2018
Cited alongside, same era.
Does distributionally robust supervised learning give robust classifiers?
W. Hu, G. Niu, I. Sato, and M. Sugiyama. 2018 · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
D. Kaushik and Z. C. Lipton. 2018 · 2018
Cited alongside, same era.
C. Clark, M. Yatskar, and L. Zettlemoyer. 2019 · 2019
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2019 · 2019
Closest in time.
Training on synthetic noise improves robustness to natural noise in machine translation
V. Karpukhin, O. Levy, J. Eisenstein, and M. Ghazvininejad. 2019 · 2019
Closest in time.
Inoculation by fine-tuning: A method for analyzing challenge datasets
N. F. Liu, R. Schwartz, and N. A. Smith. 2019 · 2019
Closest in time.
Robustness may be at odds with accuracy
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. 2019 · 2019
Closest in time.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. 2019 · 2019
Closest in time.
Gender bias in contextualized word embeddings
J. Zhao, T. Wang, M. Yatskar, R. Cotterell, V. Ordonez, and K. Chang. 2019 · 2019
Closest in time.