Fetching the paper…
Reading the bibliography…
Topic models can be useful tools to discover latent topics in collections of documents.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang. 1995 · 1995
Earlier work this paper cites.
A probabilistic analysis of the rocchio algorithm with tfidf for text categorization
Thorsten Joachims. 1996 · 1996
Earlier work this paper cites.
The use of mmr, diversity-based reranking for reordering documents and producing summaries
Jaime Carbonell and Jade Goldstein. 1998 · 1998
Earlier work this paper cites.
When is “nearest neighbor” meaningful?
Kevin Beyer, Jonathan Goldstein, Raghu Ramakrishnan, and Uri Shaft. 1999 · 1999
Earlier work this paper cites.
On the surprising behavior of distance metrics in high dimensional space
Charu C Aggarwal, Alexander Hinneburg, and Daniel A Keim. 2001 · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Pre-training is a hot topic: Contextualized document embeddings improve topic coherence
Federico Bianchi, Silvia Terragni, and Dirk Hovy. 2020a · 2004
Earlier work this paper cites.
Cross-lingual contextualized topic models with zero-shot learning
Federico Bianchi, Silvia Terragni, Dirk Hovy, Debora Nozza, and Elisabetta Fersini. 2020b · 2004
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych. 2020 · 2004
Earlier work this paper cites.
Tired of topic models? clusters of pretrained word embeddings make for fast and good topics too!
Suzanna Sia, Ayush Dalmia, and Sabrina J Mielke. 2020 · 2004
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020 · 2004
Earlier work this paper cites.
The challenges of clustering high dimensional data
Michael Steinbach, Levent Ertöz, and Vipin Kumar. 2004 · 2004
Earlier work this paper cites.
Dynamic topic models
David M Blei and John D Lafferty. 2006 · 2006
Earlier work this paper cites.
Practical solutions to the problem of diagonal dominance in kernel document clustering
Derek Greene and Pádraig Cunningham. 2006 · 2006
Cited alongside, same era.
Top2vec: Distributed representations of topics
Dimo Angelov. 2020 · 2008
Cited alongside, same era.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma. 2009 · 2009
Cited alongside, same era.
Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych. 2020 · 2010
Cited alongside, same era.
Topic modeling with contextualized word representation clusters
Laure Thompson and David Mimno. 2020 · 2010
Cited alongside, same era.
Topic modeling over short texts by incorporating word embeddings
Jipeng Qiang, Ping Chen, Tong Wang, and Xindong Wu. 2017 · 2017
Later among the works it cites.
We-lda: a word embeddings augmented lda model for web services clustering
Min Shi, Jianxun Liu, Dong Zhou, Mingdong Tang, and Buqing Cao. 2017 · 2017
Later among the works it cites.
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Algorithms for nonnegative matrix factorization with the β \beta -divergence
Cédric Févotte and Jérôme Idier. 2011 · 2011
Cited alongside, same era.
A neural autoregressive topic model
Hugo Larochelle and Stanislas Lauly. 2012 · 2012
Cited alongside, same era.
Machine reading tea leaves: Automatically evaluating topic coherence and topic model quality
Jey Han Lau, David Newman, and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov. 2014 · 2014
Cited alongside, same era.
A novel neural topic model and its supervised extension
Ziqiang Cao, Sujian Li, Yang Liu, Wenjie Li, and Heng Ji. 2015 · 2015
Cited alongside, same era.
Topical word embeddings
Yang Liu, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. 2015 · 2015
Cited alongside, same era.
Improving topic models with latent feature word representations
Dat Quoc Nguyen, Richard Billingsley, Lan Du, and Mark Johnson. 2015 · 2015
Cited alongside, same era.
L. McInnes, J. Healy, and J. Melville. 2018 · 2018
Later among the works it cites.
Umap: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger. 2018 · 2018
Later among the works it cites.
Systematic review of clustering high-dimensional and large datasets
Divya Pandove, Shivan Goel, and Rinkl Rani. 2018 · 2018
Later among the works it cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Considerably improving clustering algorithms using umap dimensionality reduction technique: A comparative study
Mebarka Allaoui, Mohammed Lamine Kherfi, and Abdelhakim Cheriet. 2020 · 2020
Later among the works it cites.
Topic modeling in embedding spaces
Adji B Dieng, Francisco JR Ruiz, and David M Blei. 2020 · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
Is automated topic model evaluation broken? the incoherence of coherence
Alexander Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov, Jordan Boyd-Graber, and Philip Resnik. 2021 · 2021
Later among the works it cites.
Octis: Comparing and optimizing topic models is simple!
Silvia Terragni, Elisabetta Fersini, Bruno Giovanni Galuzzi, Pietro Tropeano, and Antonio Candelieri. 2021 · 2021
Later among the works it cites.
Topic modelling meets deep neural networks: A survey
He Zhao, Dinh Phung, Viet Huynh, Yuan Jin, Lan Du, and Wray Buntine. 2021 · 2021
Later among the works it cites.