Fetching the paper…
Reading the bibliography…
Adversarial evaluation stress tests a model's understanding of natural language.
The Little Foxes revived
Elizabeth Hardwick. 1967 · 1967
Earlier work this paper cites.
Lillian Hellman’s "The Little Foxes" and the new south creed: An ironic view of southern history
Ritchie D. Watson. 1996 · 1996
Earlier work this paper cites.
Natural language question answering: The view from here
Lynette Hirschman and Rob Gaizauskas. 2001 · 2001
Earlier work this paper cites.
Writing good quizbowl questions: A quick primer
Paul Lujan and Seth Teitler. 2003 · 2003
Earlier work this paper cites.
Prisoner of Trebekistan: A Decade in Jeopardy!
Bob Harris. 2006 · 2006
Earlier work this paper cites.
Brainiac: adventures in the curious, competitive, compulsive world of trivia buffs
Ken Jennings. 2006 · 2006
Earlier work this paper cites.
Building Watson: An Overview of the DeepQA Project
David Ferrucci, Eric Brown, Jennifer Chu-Carroll, James Fan, David Gondek, Aditya A. Kalyanpur, Adam Lally, J. William Murdock, Eric Nyberg, John Prager, Nico Schlaefer, and Chris Welty. 2010 · 2010
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
To search or to ask: the routing of information needs between traditional search engines and social networks
Anne Oeldorf-Hirsch, Brent Hecht, Meredith Ringel Morris, Jaime Teevan, and Darren Gergle. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Elasticsearch: The Definitive Guide
Clinton Gormley and Zachary Tong. 2015 · 2015
Earlier work this paper cites.
Removing the training wheels: A coreference dataset that entertains humans and challenges computers
Anupam Guha, Mohit Iyyer, Danny Bouman, and Jordan Boyd-Graber. 2015 · 2015
Earlier work this paper cites.
Deep unordered composition rivals syntactic methods for text classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III. 2015 · 2015
Earlier work this paper cites.
A thorough examination of the CNN/Daily Mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
Interpretese vs. translationese: The uniqueness of human strategies in simultaneous interpretation
He He, Jordan Boyd-Graber, and Hal Daumé III. 2016 · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016 · 2016
Cited alongside, same era.
Why should I trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Build it, break it, fix it: Contesting secure development
Andrew Ruef, Michael Hicks, James Parker, Dave Levin, Michelle L. Mazurek, and Piotr Mardziel. 2016 · 2016
Cited alongside, same era.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
How much reading does reading comprehension require? A critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Closest in time.
Methods for interpreting and understanding deep neural networks
Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. 2018 · 2018
Closest in time.
Did the model understand the question?
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018 · 2018
Closest in time.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Closest in time.
Semantically equivalent adversarial rules for debugging nlp models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards linguistically generalizable NLP systems: A workshop and shared task
Allyson Ettinger, Sudha Rao, Hal Daumé III, and Emily M. Bender. 2017 · 2017
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Cited alongside, same era.
The craft of writing pyramidal quiz questions: Why writing quiz bowl questions is an intellectual task
Ike Jose. 2017 · 2017
Cited alongside, same era.
Starcraft II: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, John Quan, Stephen Gaffney, Stig Petersen, Karen Simonyan, Tom Schaul, Hado van Hasselt, David Silver, Timothy P. Lillicrap, Kevin Calderone, Paul Keet, Anthony Brunasso, David Lawrence, Anders Ekermo, Jacob Repp, and Rodney Tsing. 2017 · 2017
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Cited alongside, same era.
Human-Computer Question Answering: The Case for Quizbowl . Springer
Jordan Boyd-Graber, Shi Feng, and Pedro Rodriguez. 2018 · 2018
Cited alongside, same era.
How To Become A Centaur
Nicky Case. 2018 · 2018
Cited alongside, same era.
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Closest in time.
Interpreting neural networks with nearest neighbors
Eric Wallace, Shi Feng, and Jordan Boyd-Graber. 2018 · 2018
Closest in time.
Studio ousia’s quiz bowl question answering system
Ikuya Yamada, Ryuji Tamaki, Hiroyuki Shindo, and Yoshiyasu Takefuji. 2018 · 2018
Closest in time.
QANet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le. 2018 · 2018
Closest in time.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Closest in time.
Record: Bridging the gap between human and machine commonsense reading comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018 · 2018
Closest in time.
Generating natural adversarial examples
Zhengli Zhao, Dheeru Dua, and Sameer Singh. 2018 · 2018
Closest in time.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Closest in time.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Closest in time.
Natural Questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Rhinehart, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, et al. 2019 · 2019
Closest in time.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Closest in time.