Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al (2020) Language models are few-shot learners. In: Neuromuscular junction. Handbook of experimental pharmacology. Curran Associates Inc., Red Hook, NY, USA, pp 1877–1901
1901
Earlier work this paper cites.
Ziegler DM, Stiennon N, Wu J, Brown TB, Radford A, Amodei D, et al (2019) Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593
Original
1909
Earlier work this paper cites.
Liu X, Chen Q, Deng C, Zeng H, Chen J, Li D, et al (2018) LCQMC: A large-scale Chinese question matching corpus. In: Bender EM, Derczynski L, Isabelle P (eds) Proceedings of the 27th International Conference on Computational Linguistics. ACL, Santa Fe, New Mexico, USA, pp 1952–1962
1962
Earlier work this paper cites.
Hu B, Chen Q, Zhu F (2015) LCSTS: A large scale Chinese short text summarization dataset. In: Màrquez L, Callison-Burch C, Su J (eds) Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. ACL, Lisbon, Portugal, pp 1967–1972, 10.18653/v1/D15-1229
1972
Earlier work this paper cites.
Grishman R, Sundheim B (1996) Message Understanding Conference-6: A brief history. In: Proceedings of the 16th Conference on Computational Linguistics, vol 1. ACL, USA, pp 466–471
1996
Earlier work this paper cites.
Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Computation 9(8):1735–1780
1997
Earlier work this paper cites.
Cer D, Diab M, Agirre E, Lopez-Gazpio I, Specia L (2017) SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In: Bethard S, Carpuat M, Apidianaki M, Mohammad SM, Cer D, Jurgens D (eds) Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). ACL, Vancouver, Canada, pp 1–14, 10.18653/v1/S17-2001
2001
Earlier work this paper cites.
Xu L, Dong Q, Liao Y, Yu C, Tian Y, Liu W, et al (2020a) CLUENER2020: Fine-grained named entity recognition dataset and benchmark for Chinese. arXiv preprint arXiv:2001.04351
Original
2001
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, Zhu WJ (2002) BLEU: A method for automatic evaluation of machine translation. In: Isabelle P, Charniak E, Lin D (eds) Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. ACL, Philadelphia, Pennsylvania, USA, pp 311–318, 10.3115/1073083.1073135
2002
Earlier work this paper cites.
Siddhant A, Hu J, Johnson M, Firat O, Ruder S (2020) XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. arXiv preprint arXiv:2003.11080
Original
2003
Earlier work this paper cites.
Tjong Kim Sang EF, De Meulder F (2003) Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In: Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pp 142–147
2003
Earlier work this paper cites.
Xu L, Zhang X, Dong Q (2020c) CLUECorpus2020: A large-scale Chinese corpus for pre-training language model. arXiv preprint arXiv:2003.01355
Original
2003
Earlier work this paper cites.
Lin CY (2004) ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out. ACL, Barcelona, Spain, pp 74–81
2004
Earlier work this paper cites.
Dolan WB, Brockett C (2005) Automatically constructing a corpus of sentential paraphrases. In: Proceedings of the Third International Workshop on Paraphrasing (IWP2005), pp 9–16
2005
Earlier work this paper cites.
Bar-Haim R, Dagan I, Dolan B, Ferro L, Giampiccolo D, Magnini B, et al (2006) The second PASCAL recognising textual entailment challenge. URL https://api.semanticscholar.org/CorpusID:13385138
2006
Earlier work this paper cites.
Dagan I, Glickman O, Magnini B (2006) The PASCAL recognising textual entailment challenge. In: Quiñonero-Candela J, Dagan I, Magnini B, d’Alché Buc F (eds) Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 177–190
2006
Earlier work this paper cites.
Levow GA (2006) The third international Chinese language processing bakeoff: Word segmentation and named entity recognition. In: Ng HT, Kwong OO (eds) Proceedings of the Fifth SIGHAN Workshop on Chinese Language Processing. ACL, Sydney, Australia, pp 108–117
2006
Earlier work this paper cites.
Paek T (2006) Reinforcement learning for spoken dialogue systems: Comparing strengths and weaknesses for practical deployment. In: Proc. Dialog-on-Dialog Workshop, Interspeech
2006
Earlier work this paper cites.
Giampiccolo D, Magnini B, Dagan I, Dolan B (2007) The third PASCAL recognizing textual entailment challenge. In: Sekine S, Inui K, Dagan I, Dolan B, Giampiccolo D, Magnini B (eds) Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing. ACL, Prague, pp 1–9
2007
Earlier work this paper cites.
Bentivogli L, Clark P, Dagan I, Giampiccolo D (2009) The fifth PASCAL recognizing textual entailment challenge. TAC 7:8
2009
Earlier work this paper cites.
Go A, Bhayani R, Huang L (2009) Twitter sentiment classification using distant supervision. CS224N project report, Stanford 1(12)
2009
Earlier work this paper cites.
Zhao W, Shang M, Liu Y, Wang L, Liu J (2020) Ape210K: A large-scale and template-rich dataset of math word problems. arXiv preprint arXiv:2009.11506
Original
2009
Earlier work this paper cites.
Eisele A, Chen Y (2010) MultiUN: A multilingual corpus from United Nation documents. In: Calzolari N, Choukri K, Maegaard B, Mariani J, Odijk J, Piperidis S, et al (eds) Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10). ELRA, Valletta, Malta, pp 2868–2872
2010
Earlier work this paper cites.
Maas AL, Daly RE, Pham PT, Huang D, Ng AY, Potts C (2011) Learning word vectors for sentiment analysis. In: Lin D, Matsumoto Y, Mihalcea R (eds) Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies. ACL, Portland, Oregon, USA, pp 142–150
2011
Earlier work this paper cites.
Roemmele M, Bejan CA, Gordon AS (2011) Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In: 2011 AAAI Spring Symposium Series
2011
Earlier work this paper cites.
Levesque H, Davis E, Morgenstern L (2012) The winograd schema challenge. In: Thirteenth international conference on the principles of knowledge representation and reasoning, pp 552–561
2012
Earlier work this paper cites.
Rahman A, Ng V (2012) Resolving complex cases of definite pronouns: The winograd schema challenge. In: Tsujii J, Henderson J, Paşca M (eds) Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. ACL, Jeju Island, Korea, pp 777–789
2012
Earlier work this paper cites.
Weischedel R, Palmer M, Marcus M, Hovy E, Pradhan S, Ramshaw L, et al (2012) OntoNotes release 5.0 with OntoNotes DB tool v0.999 beta. Linguistic Data Consortium pp 1–53
2012
Earlier work this paper cites.
Richardson M, Burges CJ, Renshaw E (2013) MCTest: A challenge dataset for the open-domain machine comprehension of text. In: Yarowsky D, Baldwin T, Korhonen A, Livescu K, Bethard S (eds) Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. ACL, Seattle, Washington, USA, pp 193–203
2013
Earlier work this paper cites.
Socher R, Perelygin A, Wu J, Chuang J, Manning CD, Ng A, et al (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Yarowsky D, Baldwin T, Korhonen A, Livescu K, Bethard S (eds) Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. ACL, Seattle, Washington, USA, pp 1631–1642
2013
Earlier work this paper cites.
Wu SH, Liu CL, Lee LH (2013) Chinese spelling check evaluation at SIGHAN bake-off 2013. In: Yu LC, Tseng YH, Zhu J, Ren F (eds) Proceedings of the Seventh SIGHAN Workshop on Chinese Language Processing. Asian Federation of Natural Language Processing, Nagoya, Japan, pp 35–42
2013
Earlier work this paper cites.
Yu LC, Lee LH, Tseng YH, Chen HH (2014) Overview of SIGHAN 2014 bake-off for Chinese spelling check. In: Sun L, Zong C, Zhang M, Levow GA (eds) Proceedings of the Third CIPS-SIGHAN Joint Conference on Chinese Language Processing. Association for Computational Linguistics, Wuhan, China, pp 126–132, 10.3115/v1/W14-6820
2014
Earlier work this paper cites.
Bowman SR, Angeli G, Potts C, Manning CD (2015) A large annotated corpus for learning natural language inference. In: Màrquez L, Callison-Burch C, Su J (eds) Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. ACL, Lisbon, Portugal, pp 632–642, 10.18653/v1/D15-1075
2015
Earlier work this paper cites.
Peng N, Dredze M (2015) Named entity recognition for Chinese social media with jointly trained embeddings. In: Màrquez L, Callison-Burch C, Su J (eds) Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. ACL, Lisbon, Portugal, pp 548–554, 10.18653/v1/D15-1064
2015
Earlier work this paper cites.
Rush AM, Chopra S, Weston J (2015) A neural attention model for abstractive sentence summarization. In: Màrquez L, Callison-Burch C, Su J (eds) Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. ACL, Lisbon, Portugal, pp 379–389, 10.18653/v1/D15-1044
2015
Earlier work this paper cites.
Tseng YH, Lee LH, Chang LP, Chen HH (2015) Introduction to SIGHAN 2015 bake-off for Chinese spelling check. In: Yu LC, Sui Z, Zhang Y, Ng V (eds) Proceedings of the Eighth SIGHAN Workshop on Chinese Language Processing. ACL, Beijing, China, pp 32–37, 10.18653/v1/W15-3106
2015
Earlier work this paper cites.
Zhang X, Zhao J, LeCun Y (2015) Character-level convolutional networks for text classification. In: Cortes C, Lawrence N, Lee D, Sugiyama M, Garnett R (eds) Advances in Neural Information Processing Systems, vol 28. Curran Associates, Inc., pp 1–9
2015
Earlier work this paper cites.
Zhu Y, Kiros R, Zemel R, Salakhutdinov R, Urtasun R, Torralba A, et al (2015) Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp 19–27
2015
Earlier work this paper cites.
Caballero E, OpenAI , Sutskever I (2016) Description2Code dataset. https://github.com/ethancaballero/description2code
2016
Earlier work this paper cites.
Mostafazadeh N, Chambers N, He X, Parikh D, Batra D, Vanderwende L, et al (2016) A corpus and cloze evaluation for deeper understanding of commonsense stories. In: Knight K, Nenkova A, Rambow O (eds) Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL, San Diego, California, pp 839–849, 10.18653/v1/N16-1098
2016
Earlier work this paper cites.
Nguyen T, Rosenberg M, Song X, Gao J, Tiwary S, Majumder R, et al (2016) MS MARCO: A human generated machine reading comprehension dataset. choice 2640:660
2016
Earlier work this paper cites.
Paperno D, Kruszewski G, Lazaridou A, Pham NQ, Bernardi R, Pezzelle S, et al (2016) The LAMBADA dataset: Word prediction requiring a broad discourse context. In: Erk K, Smith NA (eds) Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Berlin, Germany, pp 1525–1534, 10.18653/v1/P16-1144
2016
Earlier work this paper cites.
Rajpurkar P, Zhang J, Lopyrev K, Liang P (2016) SQuAD: 100,000+ questions for machine comprehension of text. In: Su J, Duh K, Carreras X (eds) Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. ACL, Austin, Texas, pp 2383–2392, 10.18653/v1/D16-1264
2016
Earlier work this paper cites.
Wang L, Ling W (2016) Neural network-based abstract generation for opinions and arguments. In: Knight K, Nenkova A, Rambow O (eds) Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL, San Diego, California, pp 47–57, 10.18653/v1/N16-1007
2016
Earlier work this paper cites.
Ziemski M, Junczys-Dowmunt M, Pouliquen B (2016) The United Nations parallel corpus v1.0. In: Calzolari N, Choukri K, Declerck T, Goggi S, Grobelnik M, Maegaard B, et al (eds) Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16). ELRA, Portorož, Slovenia, pp 3530–3534
2016
Earlier work this paper cites.
Cettolo M, Federico M, Bentivogli L, Niehues J, Stüker S, Sudoh K, et al (2017) Overview of the IWSLT 2017 evaluation campaign. In: Proceedings of the 14th International Workshop on Spoken Language Translation, pp 2–14
2017
Earlier work this paper cites.
Christiano PF, Leike J, Brown T, Martic M, Legg S, Amodei D (2017) Deep reinforcement learning from human preferences. In: Guyon I, Luxburg UV, Bengio S, Wallach H, Fergus R, Vishwanathan S, et al (eds) Advances in Neural Information Processing Systems, vol 30. Curran Associates, Inc., pp 1–9
2017
Earlier work this paper cites.
Derczynski L, Nichols E, van Erp M, Limsopatham N (2017) Results of the WNUT2017 shared task on novel and emerging entity recognition. In: Derczynski L, Xu W, Ritter A, Baldwin T (eds) Proceedings of the 3rd Workshop on Noisy User-generated Text. ACL, Copenhagen, Denmark, pp 140–147, 10.18653/v1/W17-4418
2017
Earlier work this paper cites.
Gardent C, Shimorina A, Narayan S, Perez-Beltrachini L (2017) Creating training corpora for NLG micro-planners. In: Barzilay R, Kan MY (eds) Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Vancouver, Canada, pp 179–188, 10.18653/v1/P17-1017
2017
Earlier work this paper cites.
Joshi M, Choi E, Weld D, Zettlemoyer L (2017) TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In: Barzilay R, Kan MY (eds) Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Vancouver, Canada, pp 1601–1611, 10.18653/v1/P17-1147
2017
Earlier work this paper cites.
Lai G, Xie Q, Liu H, Yang Y, Hovy E (2017) RACE: Large-scale reading comprehension dataset from examinations. In: Palmer M, Hwa R, Riedel S (eds) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. ACL, Copenhagen, Denmark, pp 785–794, 10.18653/v1/D17-1082
2017
Earlier work this paper cites.
Ling W, Yogatama D, Dyer C, Blunsom P (2017) Program induction by rationale generation: Learning to solve and explain algebraic word problems. In: Barzilay R, Kan MY (eds) Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Vancouver, Canada, pp 158–167, 10.18653/v1/P17-1015
2017
Earlier work this paper cites.
Novikova J, Dušek O, Rieser V (2017) The E2E dataset: New challenges for end-to-end generation. In: Jokinen K, Stede M, DeVault D, Louis A (eds) Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. ACL, Saarbrücken, Germany, pp 201–206, 10.18653/v1/W17-5525
2017
Earlier work this paper cites.
Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347
Original
2017
Earlier work this paper cites.
See A, Liu PJ, Manning CD (2017) Get to the point: Summarization with pointer-generator networks. In: Barzilay R, Kan MY (eds) Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Vancouver, Canada, pp 1073–1083, 10.18653/v1/P17-1099
2017
Earlier work this paper cites.
Wang Y, Liu X, Shi S (2017) Deep neural solver for math word problems. In: Palmer M, Hwa R, Riedel S (eds) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. ACL, Copenhagen, Denmark, pp 845–854, 10.18653/v1/D17-1088
2017
Earlier work this paper cites.
Welbl J, Liu NF, Gardner M (2017) Crowdsourcing multiple choice science questions. In: Derczynski L, Xu W, Ritter A, Baldwin T (eds) Proceedings of the 3rd Workshop on Noisy User-generated Text. ACL, Copenhagen, Denmark, pp 94–106, 10.18653/v1/W17-4413
2017
Earlier work this paper cites.
Xu B, Xu Y, Liang J, Xie C, Liang B, Cui W, et al (2017) CN-DBpedia: A never-ending Chinese knowledge extraction system. In: Benferhat S, Tabia K, Ali M (eds) Advances in Artificial Intelligence: From Theory to Practice. Springer International Publishing, Cham, pp 428–438
2017
Earlier work this paper cites.
Yan Z, Duan N, Chen P, Zhou M, Zhou J, Li Z (2017) Building task-oriented dialogue systems for online shopping. Proceedings of the AAAI Conference on Artificial Intelligence 31(1). 10.1609/aaai.v31i1.11182
2017
Earlier work this paper cites.
Zhang Y, Zhong V, Chen D, Angeli G, Manning CD (2017) Position-aware attention and supervised data improve slot filling. In: Palmer M, Hwa R, Riedel S (eds) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. ACL, Copenhagen, Denmark, pp 35–45, 10.18653/v1/D17-1004
2017
Earlier work this paper cites.
Chen J, Chen Q, Liu X, Yang H, Lu D, Tang B (2018) The BQ corpus: A large-scale domain-specific Chinese corpus for sentence semantic equivalence identification. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 4946–4951, 10.18653/v1/D18-1536
2018
Earlier work this paper cites.
Choi E, He H, Iyyer M, Yatskar M, Yih Wt, Choi Y, et al (2018) QuAC: Question answering in context. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 2174–2184, 10.18653/v1/D18-1241
2018
Earlier work this paper cites.
Clark P, Cowhey I, Etzioni O, Khot T, Sabharwal A, Schoenick C, et al (2018) Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457
Original
2018
Earlier work this paper cites.
Conneau A, Kiela D (2018) SentEval: An evaluation toolkit for universal sentence representations. In: Calzolari N, Choukri K, Cieri C, Declerck T, Goggi S, Hasida K, et al (eds) Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). European Language Resources Association (ELRA), Miyazaki, Japan, pp 1699–1704
2018
Earlier work this paper cites.
Conneau A, Rinott R, Lample G, Williams A, Bowman S, Schwenk H, et al (2018) XNLI: Evaluating cross-lingual sentence representations. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 2475–2485, 10.18653/v1/D18-1269
2018
Earlier work this paper cites.
Grusky M, Naaman M, Artzi Y (2018) Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In: Walker M, Ji H, Stent A (eds) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). ACL, New Orleans, Louisiana, pp 708–719, 10.18653/v1/N18-1065
2018
Earlier work this paper cites.
Han X, Zhu H, Yu P, Wang Z, Yao Y, Liu Z, et al (2018) FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 4803–4809, 10.18653/v1/D18-1514
2018
Earlier work this paper cites.
He W, Liu K, Liu J, Lyu Y, Zhao S, Xiao X, et al (2018) DuReader: A Chinese machine reading comprehension dataset from real-world applications. In: Choi E, Seo M, Chen D, Jia R, Berant J (eds) Proceedings of the Workshop on Machine Reading for Question Answering. ACL, Melbourne, Australia, pp 37–46, 10.18653/v1/W18-2605
2018
Earlier work this paper cites.
Khashabi D, Chaturvedi S, Roth M, Upadhyay S, Roth D (2018) Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In: Walker M, Ji H, Stent A (eds) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). ACL, New Orleans, Louisiana, pp 252–262, 10.18653/v1/N18-1023
2018
Earlier work this paper cites.
Koupaee M, Wang WY (2018) Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305
Original
2018
Earlier work this paper cites.
McCann B, Keskar NS, Xiong C, Socher R (2018) The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730
Original
2018
Earlier work this paper cites.
Mihaylov T, Clark P, Khot T, Sabharwal A (2018) Can a suit of armor conduct electricity? A new dataset for open book question answering. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 2381–2391, 10.18653/v1/D18-1260
2018
Earlier work this paper cites.
Narayan S, Cohen SB, Lapata M (2018) Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 1797–1807, 10.18653/v1/D18-1206
2018
Earlier work this paper cites.
Rajpurkar P, Jia R, Liang P (2018) Know what you don’t know: Unanswerable questions for SQuAD. In: Gurevych I, Miyao Y (eds) Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). ACL, Melbourne, Australia, pp 784–789, 10.18653/v1/P18-2124
2018
Earlier work this paper cites.
Romanov A, Shivade C (2018) Lessons from natural language inference in the clinical domain. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 1586–1596, 10.18653/v1/D18-1187
2018
Earlier work this paper cites.
Saha A, Aralikatte R, Khapra MM, Sankaranarayanan K (2018) DuoRC: Towards complex language understanding with paraphrased reading comprehension. In: Gurevych I, Miyao Y (eds) Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Melbourne, Australia, pp 1683–1693, 10.18653/v1/P18-1156
2018
Earlier work this paper cites.
Trinh TH, Le QV (2018) A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847
Original
2018
Earlier work this paper cites.
Wang A, Singh A, Michael J, Hill F, Levy O, Bowman S (2018) GLUE: A multi-task benchmark and analysis platform for natural language understanding. In: Linzen T, Chrupała G, Alishahi A (eds) Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. ACL, Brussels, Belgium, pp 353–355, 10.18653/v1/W18-5446
2018
Earlier work this paper cites.
Williams A, Nangia N, Bowman S (2018) A broad-coverage challenge corpus for sentence understanding through inference. In: Walker M, Ji H, Stent A (eds) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). ACL, New Orleans, Louisiana, pp 1112–1122, 10.18653/v1/N18-1101
2018
Earlier work this paper cites.
Xie Q, Lai G, Dai Z, Hovy E (2018) Large-scale cloze test dataset created by teachers. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 2344–2356, 10.18653/v1/D18-1257
2018
Earlier work this paper cites.
Yang Y, Yih Wt, Meek C (2015) WikiQA: A challenge dataset for open-domain question answering. In: Màrquez L, Callison-Burch C, Su J (eds) Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. ACL, Lisbon, Portugal, pp 2013–2018, 10.18653/v1/D15-1237
2018
Earlier work this paper cites.
Yang Z, Qi P, Zhang S, Bengio Y, Cohen W, Salakhutdinov R, et al (2018) HotpotQA: A dataset for diverse, explainable multi-hop question answering. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 2369–2380, 10.18653/v1/D18-1259
2018
Earlier work this paper cites.
Yu T, Zhang R, Yang K, Yasunaga M, Wang D, Li Z, et al (2018) Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J (eds) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, Brussels, Belgium, pp 3911–3921, 10.18653/v1/D18-1425
2018
Earlier work this paper cites.
Zhang S, Zhang X, Wang H, Guo L, Liu S (2018b) Multi-scale attentive interaction networks for Chinese medical question answer selection. IEEE Access 6:74061–74071. 10.1109/ACCESS.2018.2883637
2018
Earlier work this paper cites.
Zhang Y, Yang J (2018) Chinese NER using lattice LSTM. In: Gurevych I, Miyao Y (eds) Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Melbourne, Australia, pp 1554–1564, 10.18653/v1/P18-1144
2018
Earlier work this paper cites.
Amini A, Gabriel S, Lin S, Koncel-Kedziorski R, Choi Y, Hajishirzi H (2019) MathQA: Towards interpretable math word problem solving with operation-based formalisms. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). ACL, Minneapolis, Minnesota, pp 2357–2367, 10.18653/v1/N19-1245
2019
Earlier work this paper cites.
Clark C, Lee K, Chang MW, Kwiatkowski T, Collins M, Toutanova K (2019) BoolQ: Exploring the surprising difficulty of natural yes/no questions. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). ACL, Minneapolis, Minnesota, pp 2924–2936, 10.18653/v1/N19-1300
2019
Earlier work this paper cites.
Cui Y, Liu T, Che W, Xiao L, Chen Z, Ma W, et al (2019) A span-extraction dataset for Chinese machine reading comprehension. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, Hong Kong, China, pp 5883–5889, 10.18653/v1/D19-1600
2019
Earlier work this paper cites.
Dasigi P, Liu NF, Marasović A, Smith NA, Gardner M (2019) Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, Hong Kong, China, pp 5925–5932, 10.18653/v1/D19-1606
2019
Earlier work this paper cites.
De Marneffe MC, Simons M, Tonhauser J (2019) The CommitmentBank: Investigating projection in naturally occurring discourse. In: proceedings of Sinn und Bedeutung, pp 107–124
2019
Earlier work this paper cites.
Devlin J, Chang MW, Lee K, Toutanova K (2019) BERT: Pre-training of deep bidirectional Transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, pp 4171–4186
2019
Earlier work this paper cites.
Dua D, Wang Y, Dasigi P, Stanovsky G, Singh S, Gardner M (2019) DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). ACL, Minneapolis, Minnesota, pp 2368–2378, 10.18653/v1/N19-1246
2019
Earlier work this paper cites.
Fabbri A, Li I, She T, Li S, Radev D (2019) Multi-News: A large-scale multi-document summarization dataset and abstractive hierarchical model. In: Korhonen A, Traum D, Màrquez L (eds) Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. ACL, Florence, Italy, pp 1074–1084, 10.18653/v1/P19-1102
2019
Earlier work this paper cites.
Gliwa B, Mochol I, Biesek M, Wawer A (2019) SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. In: Wang L, Cheung JCK, Carenini G, Liu F (eds) Proceedings of the 2nd Workshop on New Frontiers in Summarization. ACL, Hong Kong, China, pp 70–79, 10.18653/v1/D19-5409
2019
Earlier work this paper cites.
Gokaslan A, Cohen V (2019) OpenWebText corpus. http://Skylion007.github.io/OpenWebTextCorpus
2019
Earlier work this paper cites.
He J, Fu M, Tu M (2019) Applying deep matching networks to Chinese medical question answering: A study and a dataset. BMC medical informatics and decision making 19(2):91–100
2019
Earlier work this paper cites.
Huang L, Le Bras R, Bhagavatula C, Choi Y (2019) Cosmos QA: Machine reading comprehension with contextual commonsense reasoning. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, Hong Kong, China, pp 2391–2401, 10.18653/v1/D19-1243
2019
Earlier work this paper cites.
Jie Z, Xie P, Lu W, Ding R, Li L (2019) Better modeling of incomplete annotations for named entity recognition. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). ACL, Minneapolis, Minnesota, pp 729–734, 10.18653/v1/N19-1079
2019
Earlier work this paper cites.
Jin Q, Dhingra B, Liu Z, Cohen W, Lu X (2019) PubMedQA: A dataset for biomedical research question answering. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, Hong Kong, China, pp 2567–2577, 10.18653/v1/D19-1259
2019
Earlier work this paper cites.
Kwiatkowski T, Palomaki J, Redfield O, Collins M, Parikh A, Alberti C, et al (2019) Natural Questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics 7:452–466. 10.1162/tacl_a_00276
2019
Earlier work this paper cites.
Lin K, Tafjord O, Clark P, Gardner M (2019) Reasoning over paragraph effects in situations. In: Fisch A, Talmor A, Jia R, Seo M, Choi E, Chen D (eds) Proceedings of the 2nd Workshop on Machine Reading for Question Answering. ACL, Hong Kong, China, pp 58–62, 10.18653/v1/D19-5808
2019
Earlier work this paper cites.
Min Q, Shi Y, Zhang Y (2019) A pilot study for Chinese SQL semantic parsing. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, Hong Kong, China, pp 3652–3658, 10.18653/v1/D19-1377
2019
Earlier work this paper cites.
Pilehvar MT, Camacho-Collados J (2019) WiC: The word-in-context dataset for evaluating context-sensitive meaning representations. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). ACL, Minneapolis, Minnesota, pp 1267–1273, 10.18653/v1/N19-1128
2019
Earlier work this paper cites.
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I, et al (2019) Language models are unsupervised multitask learners. OpenAI blog 1(8):1–24
2019
Earlier work this paper cites.