Fetching the paper…
Reading the bibliography…
Word or word-fragment based Language Models (LM) are typically preferred over character-based ones in many downstream applications.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019 · 1911
Earlier work this paper cites.
Blimp: A benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2019 · 1912
Earlier work this paper cites.
Contextual correlates of synonymy
Herbert Rubenstein and John B Goodenough. 1965 · 1965
Earlier work this paper cites.
Contextual correlates of semantic similarity
George A Miller and Walter G Charles. 1991 · 1991
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
The emergence of grammaticality in connectionist networks
Joseph Allen and Mark S Seidenberg. 1999 · 1999
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2006
Earlier work this paper cites.
Verb similarity on the taxonomy of WordNet
Dongqiang Yang and David Martin Powers. 2006 · 2006
Earlier work this paper cites.
A study on similarity and relatedness using distributional and wordnet-based approaches
Eneko Agirre, Enrique Alfonseca, Keith Hall, Jana Kravalova, Marius Pasca, and Aitor Soroa. 2009 · 2009
Earlier work this paper cites.
A bayesian framework for word segmentation: Exploring the effects of context
S. Goldwater, Tom Griffiths, and M. Johnson. 2009 · 2009
Cited alongside, same era.
Wuggy: A multilingual pseudoword generator
Emmanuel Keuleers and Marc Brysbaert. 2010 · 2010
Cited alongside, same era.
A word at a time: computing word relatedness using temporal semantic analysis
Kira Radinsky, Eugene Agichtein, Evgeniy Gabrilovich, and Shaul Markovitch. 2011 · 2011
Cited alongside, same era.
Distributional semantics in technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, and Nam-Khanh Tran. 2012 · 2012
Cited alongside, same era.
Large-scale learning of word relatedness with constraints
Guy Halawi, Gideon Dror, Evgeniy Gabrilovich, and Yehuda Koren. 2012 · 2012
Cited alongside, same era.
Better word representations with recursive neural networks for morphology
Minh-Thang Luong, Richard Socher, and Christopher D Manning. 2013 · 2013
Comparing character-level neural language models using a lexical decision task
Gaël Godais, Tal Linzen, and Emmanuel Dupoux. 2017 · 2017
Later among the works it cites.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Yu-An Chung and James Glass. 2018 · 2018
Later among the works it cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Can LSTM learn to capture agreement? the case of basque
Shauli Ravfogel, Francis M Tyers, and Yoav Goldberg. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An unsupervised model for instance level subcategorization acquisition
Simon Baker, Roi Reichart, and Anna Korhonen. 2014 · 2014
Cited alongside, same era.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Cited alongside, same era.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur. 2015 · 2015
Cited alongside, same era.
Simverb-3500: A large-scale evaluation set of verb similarity
Daniela Gerz, Ivan Vulić, Felix Hill, Roi Reichart, and Anna Korhonen. 2016 · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Michael Hahn and Marco Baroni. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Later among the works it cites.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.