Fetching the paper…
Reading the bibliography…
Contextualized embeddings such as BERT can serve as strong input representations to NLP tasks, outperforming their static embeddings counterparts such as skip-gram, CBOW and GloVe.
Visualizing and measuring the geometry of bert
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Todd Pearce, Fernanda Vi’egas, and Martin Wattenberg. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
Placing search in context: The concept revisited
Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2001 · 2001
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Distributional semantics in technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, and Nam-Khanh Tran. 2012 · 2012
Earlier work this paper cites.
SemEval-2012 task 2: Measuring degrees of relational similarity
David Jurgens, Saif Mohammad, Peter Turney, and Keith Holyoak. 2012 · 2012
Earlier work this paper cites.
Better word representations with recursive neural networks for morphology
Thang Luong, Richard Socher, and Christopher Manning. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
A unified model for word sense representation and disambiguation
Xinxiong Chen, Zhiyuan Liu, and Maosong Sun. 2014 · 2014
Earlier work this paper cites.
Less grammar, more features
David Hall, Greg Durrett, and Dan Klein. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Retrofitting word vectors to semantic lexicons
Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, and Noah A. Smith. 2015 · 2015
Cited alongside, same era.
Ontologically grounded multi-sense representation learning for semantic vector space models
Sujay Kumar Jauhar, Chris Dyer, and Eduard Hovy. 2015 · 2015
Cited alongside, same era.
Do multi-sense embeddings improve natural language understanding?
Jiwei Li and Dan Jurafsky. 2015 · 2015
Cited alongside, same era.
Two/too simple adaptations of Word2Vec for syntax problems
Wang Ling, Chris Dyer, Alan W. Black, and Isabel Trancoso. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Dependency based embeddings for sentence classification tasks
Alexandros Komninos and Suresh Manandhar. 2016 · 2016
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Closest in time.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Closest in time.
KagNet: Knowledge-aware graph networks for commonsense reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation
Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, and Eneko Agirre. 2018 · 2018
Cited alongside, same era.
QuAC: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Bill Yuchen Lin, Xinyue Chen, Jamin Chen, and Xiang Ren. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Incorporating syntactic and semantic information in word embeddings using graph convolutional networks
Shikhar Vashishth, Manik Bhandari, Prateek Yadav, Piyush Rai, Chiranjib Bhattacharyya, and Partha Talukdar. 2019 · 2019
Closest in time.
BERT post-training for review reading comprehension and aspect-based sentiment analysis
Hu Xu, Bing Liu, Lei Shu, and Philip Yu. 2019 · 2019
Closest in time.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Closest in time.
Specializing word embeddings for similarity or relatedness
Douwe Kiela, Felix Hill, and Stephen Clark. 2015a · 2048
Closest in time.
Specializing word embeddings for similarity or relatedness
Douwe Kiela, Felix Hill, and Stephen Clark. 2015b · 2048
Closest in time.