Fetching the paper…
Reading the bibliography…
Typically, tokenization is the very first step in most text processing works.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
KorQuAD 1.0: Korean qa dataset for machine reading comprehension
Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
21st century sejong project-compiling korean corpora
Beom-mo Kang and Hung-gyu Kim. 2001 · 2001
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2004
Earlier work this paper cites.
KorNLI and KorSTS: New benchmark datasets for korean natural language understanding
Jiyeon Ham, Yo Joong Choe, Kyubyong Park, Ilji Choi, and Hyungjoon Soh. 2020 · 2004
Earlier work this paper cites.
Mecab: Yet another part-of-speech and morphological analyzer
Taku Kudo. 2006 · 2006
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Overview of the 2nd workshop on Asian translation
Toshiaki Nakazawa, Hideya Mino, Isao Goto, Graham Neubig, Sadao Kurohashi, and Eiichiro Sumita. 2015 · 2015
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Overview of the 3rd workshop on Asian translation
Toshiaki Nakazawa, Chenchen Ding, Hideya Mino, Isao Goto, Graham Neubig, and Sadao Kurohashi. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Overview of the 4th workshop on Asian translation
Toshiaki Nakazawa, Shohei Higashiyama, Chenchen Ding, Hideya Mino, Isao Goto, Hideto Kazawa, Yusuke Oda, Graham Neubig, and Sadao Kurohashi. 2017 · 2017
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Later among the works it cites.
Overview of the 6th workshop on Asian translation
Toshiaki Nakazawa, Nobushige Doi, Shohei Higashiyama, Chenchen Ding, Raj Dabre, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, and Sadao Kurohashi. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural machine translation for morphologically rich languages with improved sub-word units and synthetic data
Mārcis Pinnis, Rihards Krišlauks, Daiga Deksne, and Toms Miks. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Meaningless yet meaningful: Morphology grounded subword-level NMT
Tamali Banerjee and Pushpak Bhattacharyya. 2018 · 2018
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Overview of the 5th workshop on Asian translation
Toshiaki Nakazawa, Katsuhito Sudoh, Shohei Higashiyama, Chenchen Ding, Raj Dabre, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, and Sadao Kurohashi. 2018 · 2018
Cited alongside, same era.
Parallel corpus filtering and korean-optimized subword tokenization for machine translation
Chanjun Park, Gyeongmin kim, and Heuiseok Lim. 2019a
Cited in the paper.
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
KNU-HYUNDAI’s NMT system for scientific paper and patent tasks on WAT 2019
Cheoneum Park, Young-Jun Jung, Kihoon Kim, Geonyeong Kim, Jae-Won Jeon, Seongmin Lee, Junseok Kim, and Changki Lee. 2019b · 2019
Later among the works it cites.
Morphology-aware word-segmentation in dialectal Arabic adaptation of neural machine translation
Ahmed Tawfik, Mahitab Emam, Khaled Essam, Robert Nabil, and Hany Hassan. 2019 · 2019
Later among the works it cites.
PAWS-x: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019 · 2019
Later among the works it cites.
Jamo pair encoding: Subcharacter representation-based extreme Korean vocabulary compression for efficient subword tokenization
Sangwhan Moon and Naoaki Okazaki. 2020 · 2020
Closest in time.