Fetching the paper…
Reading the bibliography…
Over the last few years, multi-vector retrieval methods, spearheaded by ColBERT, have become an increasingly popular approach to Neural IR.
Www’18 open challenge: financial opinion mining and question answering. In Companion proceedings of the the web conference 2018 . 1941–1942
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018 · 1942
Earlier work this paper cites.
A synopsis of linguistic theory, 1930-1955
John Firth. 1957 · 1957
Earlier work this paper cites.
Hierarchical grouping to optimize an objective function
Joe H Ward Jr. 1963 · 1963
Earlier work this paper cites.
Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability , Vol. 1. Oakland, CA, USA, 281–297
James MacQueen et al · 1967
Earlier work this paper cites.
Hierarchical clustering of a Finnish newspaper article collection with graded relevance assessments
Tuomo Korenius, Jorma Laurikkala, Martti Juhola, and Kalervo Järvelin. 2006 · 2006
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010 · 2010
Earlier work this paper cites.
Algorithms for hierarchical clustering: an overview
Fionn Murtagh and Pedro Contreras. 2012 · 2012
Earlier work this paper cites.
Ms marco: A human-generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Yu A Malkov and Dmitry A Yashunin. 2018 · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Overview of Touché 2020: argument retrieval. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 11th International Conference of the CLEF Association, CLEF 2020, Thessaloniki, Greece, September 22–25, 2020, Proceedings 11 . Springer, 384–395
Alexander Bondarenko, Maik Fröbe, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, et al · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Semantic Document Clustering using K-means algorithm and Ward’s Method. In 2020 International Conference on Advanced Science and Engineering (ICOASE) . IEEE, 1–6
Niyaz M Salih and Karwan Jacksi. 2020 · 2020
Cited alongside, same era.
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval
Nandan Thakur, Luiz Bonifacio, Maik Fröbe, Alexander Bondarenko, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, and Jimmy Lin. 2024 · 2020
Cited alongside, same era.
SciPy 1.0: fundamental algorithms for scientific computing in Python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Cited alongside, same era.
Inverted files for text search engines
Justin Zobel and Alistair Moffat. 2006 · 2020
Cited alongside, same era.
Introducing neural bag of whole-words with colberter: Contextualized late interactions using enhanced reduction. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 737–747
Sebastian Hofstätter, Omar Khattab, Sophia Althammer, Mete Sertkan, and Allan Hanbury. 2022 · 2022
Later among the works it cites.
JGLUE: Japanese general language understanding evaluation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference . 2957–2966
Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. 2022 · 2022
Later among the works it cites.
Learned token pruning in contextualized late interaction over BERT (colbert). In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2232–2236
Carlos Lassance, Maroua Maachou, Joohee Park, and Stéphane Clinchant. 2022 · 2022
Later among the works it cites.
PLAID: an efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 1747–1756
Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SPLADE: Sparse lexical and expansion model for first stage ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2288–2292
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021 · 2021
Cited alongside, same era.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021 · 2021
Cited alongside, same era.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Pretrained transformers for text ranking: BERT and beyond. In Proceedings of the 14th ACM International Conference on web search and data mining . 1154–1156
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
Jointly optimizing query encoder and product quantization to improve retrieval performance. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management . 2487–2496
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021 · 2021
Cited alongside, same era.
ranx: A blazing-fast python library for ranking evaluation and comparison. In European Conference on Information Retrieval . Springer, 259–264
Elias Bassani. 2022 · 2022
Cited alongside, same era.
Token merging: Your vit but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2022 · 2022
Cited alongside, same era.
What company do words keep? Revisiting the distributional semantics of JR Firth & Zellig Harris
Mikael Brunila and Jack LaViolette. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Evaluating interpolation and extrapolation performance of neural retrieval models. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 2486–2496
Jingtao Zhan, Xiaohui Xie, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2022 · 2022
Later among the works it cites.
Token merging for fast stable diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4599–4603
Daniel Bolya and Judy Hoffman. 2023 · 2023
Later among the works it cites.
Benjamin Clavié. 2023 · 2023
Later among the works it cites.
Ms-shift: An analysis of ms marco distribution shifts on neural retrieval. In European Conference on Information Retrieval . Springer, 636–652
Simon Lupart, Thibault Formal, and Stéphane Clinchant. 2023 · 2023
Later among the works it cites.
MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2023 · 2023
Later among the works it cites.
Rethinking the role of token retrieval in multi-vector retrieval
Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, and Vincent Zhao. 2024 · 2024
Closest in time.
A Reproducibility Study of PLAID
Sean MacAvaney and Nicola Tonellotto. 2024 · 2024
Closest in time.
Cohere int8 & binary Embeddings - Scale Your Vector Database to Large Datasets
Neils Reimers. 2024 · 2024
Closest in time.