Fetching the paper…
Reading the bibliography…
This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
When is “nearest neighbor” meaningful?
Kevin Beyer, Jonathan Goldstein, Raghu Ramakrishnan, and Uri Shaft. 1999 · 1999
Earlier work this paper cites.
On the surprising behavior of distance metrics in high dimensional space
Charu C. Aggarwal, Alexander Hinneburg, and Daniel A. Keim. 2001 · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Andrew Rosenberg and Julia Hirschberg. 2007 · 2007
Earlier work this paper cites.
Algorithms for nonnegative matrix factorization with the β \beta -divergence
Cédric Févotte and Jérôme Idier. 2011 · 2011
Earlier work this paper cites.
Gottbert: a pure german language model
Raphael Scheible, Fabian Thomczyk, Patric Tippmann, Victor Jaravine, and Martin Boeker. 2020 · 2012
Earlier work this paper cites.
API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Gaël Varoquaux. 2013 · 2013
Earlier work this paper cites.
hdbscan: Hierarchical density based clustering
Leland McInnes, John Healy, and Steve Astels. 2017 · 2017
Earlier work this paper cites.
One million posts: A data set of german online discussions
Dietmar Schabus, Marcin Skowron, and Martin Trapp. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
SentEval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019 · 2019
Cited alongside, same era.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Cited alongside, same era.
Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Germeval 2019 task 1: Hierarchical classification of blurbs
Steffen Remus, Rami Aly, and Chris Biemann. 2019 · 2019
Cited alongside, same era.
River: Machine learning for streaming data in python
Jacob Montiel, Max Halford, Saulo Martiello Mastelini, Geoffrey Bolmier, Raphael Sourty, Robin Vaysse, Adil Zouitine, Heitor Murilo Gomes, Jesse Read, Talel Abdessalem, and Albert Bifet. 2021 · 2021
Later among the works it cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
TSDAE: Using transformer-based sequential denoising auto-encoderfor unsupervised sentence embedding learning
Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Universal sentence representation learning with conditional masked language model
Ziyi Yang, Yinfei Yang, Daniel Cer, Jax Law, and Eric Darve. 2021 · 2021
Later among the works it cites.
Pretrained transformers for text ranking: BERT and beyond
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roee Aharoni and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
The pushshift reddit dataset
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2020
Cited alongside, same era.
German’s next language model
Branden Chan, Stefan Schweter, and Timo Möller. 2020 · 2020
Cited alongside, same era.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Embedding-based retrieval in facebook search
Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Topic modelling meets deep neural networks: A survey
He Zhao, Dinh Phung, Viet Huynh, Yuan Jin, Lan Du, and Wray Buntine. 2021 · 2021
Later among the works it cites.
Bertopic: Neural topic modeling with a class-based tf-idf procedure
Maarten Grootendorst. 2022 · 2022
Later among the works it cites.
Sgpt: Gpt sentence embeddings for semantic search
Niklas Muennighoff. 2022 · 2022
Later among the works it cites.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022 · 2022
Later among the works it cites.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022 · 2022
Later among the works it cites.
KeywordScape: Visual document exploration using contextualized keyword embeddings
Henrik Voigt, Monique Meuschke, Sina Zarrieß, and Kai Lawonn. 2022 · 2022
Later among the works it cites.
Establishing infodemic management in germany: a framework for social listening and integrated analysis to report infodemic insights at the national public health institute
T Sonia Boender, Paula Helene Schneider, Claudia Houareau, Silvan Wehrli, Tina D Purnat, Atsuyoshi Ishizumi, Elisabeth Wilhelm, Christopher Voegeli, Lothar Wieler, and Christina Leuker. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
MTEB: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023 · 2037
Closest in time.