Fetching the paper…
Reading the bibliography…
We introduce Korean Language Understanding Evaluation (KLUE) benchmark.
A survey on dialogue systems: Recent advances and new frontiers
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang · 1931
Earlier work this paper cites.
An iterative design methodology for user-friendly natural language office information applications
J. F. Kelley · 1984
Earlier work this paper cites.
KAIST tree bank project for Korean: Present and future development
Key sun Choi, Young S. Han, Young G. Han, and Oh W. Kwon · 1994
Earlier work this paper cites.
Okapi at trec-3
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al · 1995
Earlier work this paper cites.
Message Understanding Conference- 6: A brief history
Ralph Grishman and Beth Sundheim · 1996
Earlier work this paper cites.
Overview of MUC-7
Nancy A. Chinchor · 1998
Earlier work this paper cites.
An activity based approach to pragmatics
Jens Allwood · 2000
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2001
Earlier work this paper cites.
Accurate unlexicalized parsing
Dan Klein and Christopher D. Manning · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
The automatic content extraction (ACE) program – tasks, data, and evaluation
George Doddington, Alexis Mitchell, Mark Przybocki, Lance Ramshaw, Stephanie Strassel, and Ralph Weischedel · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Incorporating non-local information into information extraction systems by Gibbs sampling
Jenny Rose Finkel, Trond Grenager, and Christopher Manning · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
Marie-Catherine de Marneffe, Bill MacCartney, and Christopher D. Manning · 2006
Earlier work this paper cites.
Korean treebank annotations version 2.0
Na-Rae Han, Shijong Ryu, Sook-Hee Chae, Seung-yun Yang, Seunghun Lee, and Martha Palmer · 2006
Earlier work this paper cites.
MeCab: Yet another part-of-speech and morphological analyzer, 2006
Taku Kudo · 2006
Earlier work this paper cites.
The Stanford typed dependencies representation
Marie-Catherine de Marneffe and Christopher D. Manning · 2008
Earlier work this paper cites.
Newspapers and copyright, 2009
Korea Copyright Commission · 2009
Earlier work this paper cites.
Overview of the TAC 2009 knowledge base population track
Paul McNamee and Hoa Trang Dang · 2009
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
Mike Mintz, Steven Bills, Rion Snow, and Daniel Jurafsky · 2009
Earlier work this paper cites.
SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz · 2010
Earlier work this paper cites.
Modeling relations and their mentions without labeled text
Sebastian Riedel, Limin Yao, and Andrew McCallum · 2010
Earlier work this paper cites.
Computing Krippendorff’s alpha-reliability
K. Krippendorff · 2011
Earlier work this paper cites.
Named entity recognition in tweets: An experimental study
Alan Ritter, Sam Clark, Mausam, and Oren Etzioni · 2011
Earlier work this paper cites.
A universal part-of-speech tagset
Slav Petrov, Dipanjan Das, and Ryan McDonald · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
The third dialog state tracking challenge
Matthew Henderson, Blaise Thomson, and Jason D Williams · 2014
Earlier work this paper cites.
KoNLPy: Korean natural language processing in Python
Eunjeong L. Park and Sungzoon Cho · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
SemEval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe · 2016
Earlier work this paper cites.
Simple and accurate dependency parsing using bidirectional LSTM feature representations
Eliyahu Kiperwasser and Yoav Goldberg · 2016
Earlier work this paper cites.
MS MARCO: A human generated MAchine reading COmprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Earlier work this paper cites.
Universal Dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Results of the WNUT16 named entity recognition shared task
Benjamin Strauss, Bethany Toma, Alan Ritter, Marie-Catherine de Marneffe, and Wei Xu · 2016
Earlier work this paper cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes · 2017
Earlier work this paper cites.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D. Manning · 2017
Earlier work this paper cites.
Frames: a corpus for adding memory to goal-oriented dialogue systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman · 2017
Earlier work this paper cites.
Key-value retrieval networks for task-oriented dialogue
Mihail Eric, Lakshmi Krishnan, Francois Charette, and Christopher D. Manning · 2017
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Universal Dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, Claudia Bedini, Núria Bertomeu Castelló, and Jungmee Lee · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
NewsQA: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2017
Cited alongside, same era.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young · 2017
Cited alongside, same era.
Jingjing Xu, Ji Wen, Xu Sun, and Qi Su · 2017
Cited alongside, same era.
Position-aware attention and supervised data improve slot filling
Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning · 2017
Cited alongside, same era.
MRC AI Dataset
National Information Society Agency · 2018
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Later among the works it cites.
MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines
Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur · 2020
Later among the works it cites.
KorNLI and KorSTS: New benchmark datasets for Korean natural language understanding
Jiyeon Ham, Yo Joong Choe, Kyubyong Park, Ilji Choi, and Hyungjoon Soh · 2020
Later among the works it cites.
Annotation issues in Universal Dependencies for Korean and Japanese
Ji Yoon Han, Tae Hwan Oh, Lee Jin, and Hansaem Kim · 2020
Later among the works it cites.
OCNLI: Original Chinese Natural Language Inference
Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra Kübler, and Lawrence Moss · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emily M. Bender and Batya Friedman · 2018
Cited alongside, same era.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić · 2018
Cited alongside, same era.
Building Universal Dependency treebanks in Korean
Jayeol Chun, Na-Rae Han, Jena D. Hwang, and Jinho D. Choi · 2018
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg · 2018
Cited alongside, same era.
FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation
Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2018
Cited alongside, same era.
A French corpus and annotation schema for named entity recognition and relation extraction of financial news
Ali Jabbari, Olivier Sauvage, Hamada Zeine, and Hamza Chergui · 2020
Later among the works it cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
ParsiNLU: a suite of language understanding challenges for persian
Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, et al · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Later among the works it cites.
KorQuAD 2.0: Korean QA dataset for web document machine comprehension
Youngmin Kim, Seungyoung Lim, Hyunjeong Lee, Soyoon Park, and Myungji Kim · 2020
Later among the works it cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang · 2020
Later among the works it cites.
Dynamic data selection for curriculum learning via ability estimation
John P. Lalor and Hong Yu · 2020
Later among the works it cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
FlauBERT: Unsupervised language model pre-training for French
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, and Didier Schwab · 2020
Later among the works it cites.
KcBERT: Korean comments BERT
Junbum Lee · 2020
Later among the works it cites.
KR-BERT: A small-scale Korean-specific language model
Sangah Lee, Hansol Jang, Yunmee Baik, Suzi Park, and Hyopil Shin · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
CoCo: Controllable counterfactuals for evaluating dialogue state trackers
Shiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li, Tong Niu, Nazneen Rajani, Xifeng Yan, Yingbo Zhou, and Caiming Xiong · 2020
Later among the works it cites.
XGLUE: A new benchmark datasetfor cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou · 2020
Later among the works it cites.
DialoGLUE: A natural language understanding benchmark for task-oriented dialogue
Shikib Mehri, Mihail Eric, and Dilek Hakkani-Tur · 2020
Later among the works it cites.
Syntactic data augmentation increases robustness to inference heuristics
Junghyun Min, R. Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen · 2020
Later among the works it cites.
BEEP! Korean corpus of online news comments for toxic speech detection
Jihyung Moon, Won Ik Cho, and Junbum Lee · 2020
Later among the works it cites.
Effective crowdsourcing of multiple tasks for comprehensive knowledge extraction
Sangha Nam, Minho Lee, Donghwan Kim, Kijong Han, Kuntae Kim, Sooji Yoon, Eun-kyung Kim, and Key-Sun Choi · 2020
Later among the works it cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Later among the works it cites.
NIKL CORPORA 2020 (v.1.0), 2020
National Institute of Korean Languages · 2020
Later among the works it cites.
Analysis of the Penn Korean Universal Dependency treebank (PKT-UD): Manual revision to build robust parsing model in Korean
Tae Hwan Oh, Ji Yoon Han, Hyonsu Choe, Seokwon Park, Han He, Jinho D. Choi, Na-Rae Han, Jena D. Hwang, and Hansaem Kim · 2020
Later among the works it cites.
KoELECTRA: Pretrained ELECTRA model for Korean
Jangwon Park · 2020
Later among the works it cites.
An empirical study of tokenization strategies for various Korean NLP tasks
Kyubyong Park, Joohong Lee, Seongbo Jang, and Dawoon Jung · 2020
Later among the works it cites.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al · 2020
Later among the works it cites.
RiSAWOZ: A large-scale multi-domain Wizard-of-Oz dataset with rich semantic annotations for task-oriented dialogue modeling
Jun Quan, Shian Zhang, Qian Cao, Zizhong Li, and Deyi Xiong · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan · 2020
Later among the works it cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych · 2020
Later among the works it cites.
What do models learn from question answering datasets?
Priyanka Sen and Amir Saffari · 2020
Later among the works it cites.
RussianSuperGLUE: A Russian language understanding evaluation benchmark
Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, and Andrey Evlampiev · 2020
Later among the works it cites.
Asking Crowdworkers to Write Entailment Examples: The Best of Bad options
Clara Vania, Ruijie Chen, and Samuel R. Bowman · 2020
Later among the works it cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2020
Later among the works it cites.
IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Later among the works it cites.
CLUE: A Chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan · 2020
Later among the works it cites.
Dialogue-based relation extraction
Dian Yu, Kai Sun, Claire Cardie, and Dong Yu · 2020
Later among the works it cites.
CrossWOZ: A large-scale Chinese cross-domain task-oriented dialogue dataset
Qi Zhu, Kaili Huang, Zheng Zhang, Xiaoyan Zhu, and Minlie Huang · 2020
Later among the works it cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George E Dahl · 2021
Closest in time.
FairFil: Contrastive neural debiasing method for pretrained text encoders
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, and Lawrence Carin · 2021
Closest in time.
Ethnologue: Languages of the World
David M. Eberhard and Charles D. Simons, Gary F. Fenning · 2021
Closest in time.
DeBERTa: Decoding-enhanced BERT with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
On-the-fly controlled text generation with experts and anti-experts, 2021
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi · 2021
Closest in time.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2023
Closest in time.
SemEval-2015 task 2: Semantic textual similarity, English, Spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Iñigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe · 2045
Closest in time.