Fetching the paper…
Reading the bibliography…
Transforming unstructured text into structured and meaningful forms, organized by useful category labels, is a fundamental step in text mining for downstream analysis and application.
Nettaxo: Automated topic taxonomy construction from text-rich network. In Proceedings of the Web Conference 2020 . 1908–1919
Jingbo Shang, Xinyang Zhang, Liyuan Liu, Sha Li, and Jiawei Han. 2020 · 1919
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Joseph L Fleiss and Jacob Cohen. 1973 · 1973
Earlier work this paper cites.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
Peter J Rousseeuw. 1987 · 1987
Earlier work this paper cites.
Mixture models: Inference and applications to clustering . Vol. 38
Geoffrey J McLachlan and Kaye E Basford. 1988 · 1988
Earlier work this paper cites.
Online algorithms and stochastic approximations
Léon Bottou. 1998 · 1998
Earlier work this paper cites.
Neural networks: a comprehensive foundation
Simon Haykin. 1998 · 1998
Earlier work this paper cites.
Text mining: The state of the art and the challenges. In Proceedings of the PAKDD 1999 workshop on knowledge discovery from advanced databases , Vol. 8. 65–70
Ah-Hwee Tan et al · 1999
Earlier work this paper cites.
Understanding user goals in web search. In Proceedings of the 13th international conference on World Wide Web . 13–19
Daniel E Rose and Danny Levinson. 2004 · 2004
Earlier work this paper cites.
A brief survey of text mining
Andreas Hotho, Andreas Nürnberger, and Gerhard Paaß. 2005 · 2005
Earlier work this paper cites.
k-means++: The advantages of careful seeding. In Soda , Vol. 7. 1027–1035
David Arthur, Sergei Vassilvitskii, et al · 2007
Earlier work this paper cites.
The new york times annotated corpus
Evan Sandhaus. 2008 · 2008
Earlier work this paper cites.
Reading tea leaves: How humans interpret topic models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei. 2009 · 2009
Cited alongside, same era.
A survey of text clustering algorithms
Charu C Aggarwal and ChengXiang Zhai. 2012 · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing . 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016b · 2016
Enhancing taxonomy completion with concept generation via fusing relational representations. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 2104–2113
Qingkai Zeng, Jinfeng Lin, Wenhao Yu, Jane Cleland-Huang, and Meng Jiang. 2021 · 2021
Later among the works it cites.
One Embedder, Any Task: Instruction-Finetuned Text Embeddings
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022 · 2022
Later among the works it cites.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Later among the works it cites.
Making Large Language Models Better Data Creators. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 15349–15360
Dong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White, and Sujay Jauhar. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016a · 2016
Cited alongside, same era.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Cited alongside, same era.
Taxogen: Unsupervised topic taxonomy construction by adaptive term embedding and clustering. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2701–2709
Chao Zhang, Fangbo Tao, Xiusi Chen, Jiaming Shen, Meng Jiang, Brian Sadler, Michelle Vanni, and Jiawei Han. 2018 · 2018
Cited alongside, same era.
A review of topic modeling methods
Ike Vayansky and Sathish AP Kumar. 2020 · 2020
Cited alongside, same era.
A Taxonomy of Empathetic Response Intents in Human Social Conversations. In Proceedings of the 28th International Conference on Computational Linguistics . 4886–4899
Anuradha Welivita and Pearl Pu. 2020 · 2020
Cited alongside, same era.
An intent taxonomy for questions asked in web search. In Proceedings of the 2021 Conference on Human Information Interaction and Retrieval . 85–94
B Barla Cambazoglu, Leila Tavakoli, Falk Scholer, Mark Sanderson, and Bruce Croft. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023 · 2023
Later among the works it cites.
TopicGPT: A Prompt-based Topic Modeling Framework
Chau Minh Pham, Alexander Hoyle, Simeng Sun, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Automatic Prompt Optimization with “Gradient Descent” and Beam Search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 7957–7968
Reid Pryzant, Dan Iter, Jerry Li, Yin Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Later among the works it cites.
Using large language models to generate, validate, and apply user intent taxonomies
Chirag Shah, Ryen W White, Reid Andersen, Georg Buscher, Scott Counts, Sarkar Snigdha Sarathi Das, Ali Montazer, Sathish Manivannan, Jennifer Neville, Xiaochuan Ni, et al · 2023
Later among the works it cites.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023 · 2023
Later among the works it cites.
Goal-Driven Explainable Clustering via Language Descriptions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 10626–10649
Zihan Wang, Jingbo Shang, and Ruiqi Zhong. 2023 · 2023
Later among the works it cites.
Can Large Language Models Transform Computational Social Science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023 · 2023
Later among the works it cites.