Fetching the paper…
Reading the bibliography…
Contextualized word embeddings, i.e.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Distributional semantics and linguistic theory
Gemma Boleda. 2019 · 1905
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019a · 1905
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
On the Validity of Self-Attention as Explanation in Transformer Models
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, and Roger Wattenhofer. 2019 · 1908
Earlier work this paper cites.
Designing and Interpreting Probes with Control Tasks
John Hewitt and Percy Liang. 2019 · 1909
Earlier work this paper cites.
A synopsis of linguistic theory 1930-55
J. R. Firth. 1957 · 1952
Earlier work this paper cites.
Cloze procedure: A new tool for measuring readability
Wilson Taylor. 1953 · 1953
Earlier work this paper cites.
Distributional structure
Zellig Harris. 1954 · 1954
Earlier work this paper cites.
Word And Object
William Van Ormann Quine. 1960 · 1960
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M. Cover. 1965 · 1965
Earlier work this paper cites.
A vector space model for automatic indexing
G. Salton, A. Wong, and C. S. Yang. 1975 · 1975
Earlier work this paper cites.
Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
Peter Rousseeuw. 1987 · 1987
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Language encodes geographical information
Max M. Louwerse and Rolf A. Zwaan. 2009 · 2009
Earlier work this paper cites.
Exemplar-based models for word meaning in context
Katrin Erk and Sebastian Padó. 2010 · 2010
Earlier work this paper cites.
Multi-prototype vector-space models of word meaning
Joseph Reisinger and Raymond Mooney. 2010 · 2010
Cited alongside, same era.
From frequency and to meaning and vector space and models of semantics
Peter D. Turney and Patrick Pantel. 2010 · 2010
Cited alongside, same era.
Dynamic and static prototype vectors for semantic composition
Siva Reddy, Ioannis P. Klapaftis, Diana McCarthy, and Suresh Manandhar. 2011 · 2011
Cited alongside, same era.
On the combination of relative clustering validity criteria
Lucas Vendramin, Pablo A. Jaskowiak, and Ricardo J. G. B. Campello. 2013 · 2013
Cited alongside, same era.
Multimodal distributional semantics
Elia Bruni, Nam-Khanh Tran, and Marco Baroni. 2014 · 2014
Cited alongside, same era.
A sick cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Distributional models of word meaning
Alessandro Lenci. 2018 · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Breaking sticks and ambiguities with adaptive skip-gram
Sergey Bartunov, Dmitry Kondrashkin, Anton Osokin, and Dmitry P. Vetrov. 2015 · 2015
Cited alongside, same era.
Building a shared world: mapping distributional to model-theoretic semantic spaces
Aurélie Herbelot and Eva Maria Vecchi. 2015 · 2015
Cited alongside, same era.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Affixation in semantic space: Modeling morpheme meanings with compositional distributional semantics
Marco Marelli and Marco Baroni. 2015 · 2015
Cited alongside, same era.
Autoextend: Extending word embeddings to embeddings for synsets and lexemes
Sascha Rothe and Hinrich Schütze. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Jamie Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford. 2018 · 2018
Later among the works it cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018 · 2018
Later among the works it cites.
On the dimensionality of word embedding
Zi Yin and Yuanyuan Shen. 2018 · 2018
Later among the works it cites.
What does this word mean? explaining contextualized embeddings with natural language definition
Ting-Yun Chang and Yun-Nung Chen. 2019 · 2019
Closest in time.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Closest in time.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Closest in time.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Closest in time.
Sense Vocabulary Compression through the Semantic Knowledge of WordNet for Neural Word Sense Disambiguation
Loïc Vial, Benjamin Lecouteux, and Didier Schwab. 2019 · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.
Gendered ambiguous pronoun (GAP) shared task at the gender bias in NLP workshop 2019
Kellie Webster, Marta R. Costa-jussà, Christian Hardmeier, and Will Radford. 2019 · 2019
Closest in time.
Don’t blame distributional semantics if it can’t do entailment
Matthijs Westera and Gemma Boleda. 2019 · 2019
Closest in time.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 2019
Closest in time.