Fetching the paper…
Reading the bibliography…
State-of-the-art NLP systems represent inputs with word embeddings, but these are brittle when faced with Out-of-Vocabulary (OOV) words.
Attentive mimicking: Better word embeddings by attending to informative contexts
Timo Schick and Hinrich Schütze. 2019a · 1904
Earlier work this paper cites.
Some methods for classification and analysis of multivariate observations
James MacQueen et al. 1967 · 1967
Earlier work this paper cites.
Self-organizing neural network that discovers surfaces in random-dot stereograms
Suzanna Becker and Geoffrey E Hinton. 1992 · 1992
Earlier work this paper cites.
Signature verification using a “siamese” time delay neural network
Jane Bromley, James W Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah. 1993 · 1993
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Adv-bert: Bert is not robust on misspellings! generating nature adversarial samples on bert
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong. 2020 · 2003
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee. 2005 · 2005
Earlier work this paper cites.
Attributes in lexical acquisition
Abdulrahman Almuhareb. 2006 · 2006
Earlier work this paper cites.
Declutr: Deep contrastive learning for unsupervised textual representations
John M Giorgi, Osvald Nitski, Gary D Bader, and Bo Wang. 2020 · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020 · 2006
Earlier work this paper cites.
Overview of biocreative ii gene mention recognition
Larry Smith, Lorraine K Tanabe, Rie Johnson nee Ando, Cheng-Ju Kuo, I-Fang Chung, Chun-Nan Hsu, Yu-Shi Lin, Roman Klinger, Christoph M Friedrich, Kuzman Ganchev, et al. 2008 · 2008
Earlier work this paper cites.
A study on similarity and relatedness using distributional and wordnet-based approaches
Eneko Agirre, Enrique Alfonseca, Keith Hall, Jana Kravalova, Marius Pasca, and Aitor Soroa. 2009 · 2009
Earlier work this paper cites.
How we blessed distributional semantic evaluation
Marco Baroni and Alessandro Lenci. 2011 · 2011
Earlier work this paper cites.
Large-scale learning of word relatedness with constraints
Guy Halawi, Gideon Dror, Evgeniy Gabrilovich, and Yehuda Koren. 2012 · 2012
Earlier work this paper cites.
Clear: Contrastive learning for sentence representation
Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020 · 2012
Earlier work this paper cites.
Better word representations with recursive neural networks for morphology
Minh-Thang Luong, Richard Socher, and Christopher D Manning. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Multimodal distributional semantics
Elia Bruni, Nam-Khanh Tran, and Marco Baroni. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Earlier work this paper cites.
Bidirectional lstm-crf models for sequence tagging
Zhiheng Huang, Wei Xu, and Kai Yu. 2015 · 2015
Earlier work this paper cites.
Ye Zhang and Byron Wallace. 2015 · 2015
Earlier work this paper cites.
A joint model for word embedding and word morphology
Kris Cao and Marek Rei. 2016 · 2016
Cited alongside, same era.
Intrinsic evaluation of word vectors fails to predict extrinsic performance
Billy Chiu, Anna Korhonen, and Sampo Pyysalo. 2016 · 2016
Cited alongside, same era.
Problems with evaluation of word embeddings using word similarity tasks
Manaal Faruqui, Yulia Tsvetkov, Pushpendre Rastogi, and Chris Dyer. 2016 · 2016
Cited alongside, same era.
Charagram: Embedding words and sentences via character n-grams
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Nlp augmentation
Edward Ma. 2019 · 2019
Later among the works it cites.
Misspelling oblivious word embeddings
Aleksandra Piktus, Necati Bora Edizel, Piotr Bojanowski, Édouard Grave, Rui Ferreira, and Fabrizio Silvestri. 2019 · 2019
Later among the works it cites.
Combating adversarial misspellings with robust word recognition
Danish Pruthi, Bhuwan Dhingra, and Zachary C Lipton. 2019 · 2019
Later among the works it cites.
Subword-based compact reconstruction of word embeddings
Shota Sasaki, Jun Suzuki, and Kentaro Inui. 2019 · 2019
Later among the works it cites.
Evaluating word embedding models: Methods and experimental results
Bin Wang, Angela Wang, Fenxiao Chen, Yuncheng Wang, and C-C Jay Kuo. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
On sampling strategies for neural network-based collaborative filtering
Ting Chen, Yizhou Sun, Yue Shi, and Liangjie Hong. 2017 · 2017
Cited alongside, same era.
Bpemb: Tokenization-free pre-trained subword embeddings in 275 languages
Benjamin Heinzerling and Michael Strube. 2017 · 2017
Cited alongside, same era.
Mimicking word embeddings using subword rnns
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Biowordvec, improving biomedical word embeddings with subword information and mesh
Yijia Zhang, Qingyu Chen, Zhihao Yang, Hongfei Lin, and Zhiyong Lu. 2019 · 2019
Later among the works it cites.
A systematic study of leveraging subword information for learning word representations
Yi Zhu, Ivan Vulić, and Anna Korhonen. 2019 · 2019
Later among the works it cites.
Interpreting pretrained contextualized representations via reductions to static embeddings
Rishi Bommasani, Kelly Davis, and Claire Cardie. 2020 · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Later among the works it cites.
Characterbert: Reconciling elmo and bert for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Later among the works it cites.
Robust backed-off estimation of out-of-vocabulary embeddings
Nobukazu Fukuda, Naoki Yoshinaga, and Masaru Kitsuregawa. 2020 · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Later among the works it cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Later among the works it cites.
Pbos: Probabilistic bag-of-subwords for generalizing word embedding
Zhao Jinman, Shawn Zhong, Xiaomin Zhang, and Yingyu Liang. 2020 · 2020
Later among the works it cites.
Ro{bert}a: A robustly optimized {bert} pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Charbert: Character-aware pre-trained language model
Wentao Ma, Yiming Cui, Chenglei Si, Ting Liu, Shijin Wang, and Guoping Hu. 2020 · 2020
Later among the works it cites.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. 2020 · 2020
Later among the works it cites.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Obtaining better static word embeddings using contextual embedding models
Prakhar Gupta and Martin Jaggi. 2021 · 2021
Later among the works it cites.