Fetching the paper…
Reading the bibliography…
Clustering token-level contextualized word representations produces output that shares many similarities with topic models for English text collections.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Class-based n-gram models of natural language
Peter F. Brown, Peter V. deSouza, Robert L. Mercer, Vincent J. Della Pietra, and Jenifer C. Lai. 1992 · 1992
Earlier work this paper cites.
Database-friendly random projections
Dimitris Achlioptas. 2001 · 2001
Earlier work this paper cites.
Concept decompositions for large sparse text data using clustering
Inderjit S Dhillon and Dharmendra S Modha. 2001 · 2001
Earlier work this paper cites.
Mallet: A machine learning for language toolkit
Andrew Kachites McCallum. 2002 · 2002
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Very sparse random projections
Ping Li, Trevor J. Hastie, and Kenneth Ward Church. 2006 · 2006
Earlier work this paper cites.
Incremental learning for robust visual tracking
David A Ross, Jongwoo Lim, Ruei-Sung Lin, and Ming-Hsuan Yang. 2008 · 2008
Earlier work this paper cites.
The new york times annotated corpus
Evan Sandhaus. 2008 · 2008
Earlier work this paper cites.
Evaluating topic models for digital libraries
David Newman, Youn Noh, Edmund Talley, Sarvnaz Karimi, and Timothy Baldwin. 2010 · 2010
Earlier work this paper cites.
Optimizing semantic coherence in topic models
David Mimno, Hanna Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Summarizing topical content with word frequency and exclusivity
Jonathan Bischof and Edoardo M Airoldi. 2012 · 2012
Earlier work this paper cites.
Exploring topic coherence over many models and many topics
Keith Stevens, Philip Kegelmeyer, David Andrzejewski, and David Buttler. 2012 · 2012
Earlier work this paper cites.
A practical algorithm for topic modeling with provable guarantees
Sanjeev Arora, Rong Ge, Yonatan Halpern, David Mimno, Ankur Moitra, David Sontag, Yichen Wu, and Michael Zhu. 2013 · 2013
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Spherical k-means++ clustering
Yasunori Endo and Sadaaki Miyamoto. 2015 · 2015
Cited alongside, same era.
High-reproducibility and high-accuracy method for automated topic classification
Andrea Lancichinetti, M Irmak Sirer, Jane X Wang, Daniel Acuna, Konrad Körding, and Luís A Nunes Amaral. 2015 · 2015
Cited alongside, same era.
Image-based recommendations on styles and substitutes
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton Van Den Hengel. 2015 · 2015
Cited alongside, same era.
Guided alignment training for topic-aware neural machine translation
Wenhu Chen, Evgeny Matusov, Shahram Khadivi, and Jan-Thorsten Peter. 2016 · 2016
A reinforced topic-aware convolutional sequence-to-sequence model for abstractive text summarization
Li Wang, Junlin Yao, Yunzhe Tao, Li Zhong, Wei Liu, and Qiang Du. 2018 · 2018
Later among the works it cites.
Visualizing and measuring the geometry of BERT
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda B. Viégas, and Martin Wattenberg. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Later among the works it cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Neural variational inference for topic models
Akash Srivastava and Charles Sutton. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Applications of topic models
Jordan Boyd-Graber, Yuening Hu, David Mimno, et al. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Does bert make any sense? interpretable word sense disambiguation with contextualized embeddings
Gregor Wiedemann, Steffen Remus, Avi Chawla, and Chris Biemann. 2019 · 2019
Later among the works it cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg. 2020 · 2020
Closest in time.
Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings
Rishi Bommasani, Kelly Davis, and Claire Cardie. 2020 · 2020
Closest in time.
Topic modeling in embedding spaces
Adji B. Dieng, Francisco J. R. Ruiz, and David M. Blei. 2020 · 2020
Closest in time.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Closest in time.
PhoBERT: Pre-trained language models for Vietnamese
Dat Quoc Nguyen and Anh Tuan Nguyen. 2020 · 2020
Closest in time.
tBERT: Topic models and BERT joining forces for semantic similarity detection
Nicole Peinelt, Dong Nguyen, and Maria Liakata. 2020 · 2020
Closest in time.
Tired of topic models? clusters of pretrained word embeddings make for fast and good topics too!
Suzanna Sia, Ayush Dalmia, and Sabrina J. Mielke. 2020 · 2020
Closest in time.
Probing pretrained language models for lexical semantics
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020 · 2020
Closest in time.