Fetching the paper…
Reading the bibliography…
BERT based ranking models have achieved superior performance on various information retrieval tasks.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020 · 1908
Earlier work this paper cites.
An electronic digital computor using cold cathode counting tubes for storage
RCM Barnes, EH Cooke-Yarborough, and DGA Thomas. 1951 · 1951
Earlier work this paper cites.
Picture coding using pseudo-random noise
L. Roberts. 1962 · 1962
Earlier work this paper cites.
Unified matrix treatment of the fast Walsh-Hadamard transform
Bernard J. Fino and V. Ralph Algazi. 1976 · 1976
Earlier work this paper cites.
Private vs. Common Random Bits in Communication Complexity
Ilan Newman. 1991 · 1991
Earlier work this paper cites.
Vector quantization and signal compression
A. Gersho and R. M. Gray. 1992 · 1992
Earlier work this paper cites.
Dithered quantizers
R.M. Gray and T.G. Stockham. 1993 · 1993
Earlier work this paper cites.
Approximate Nearest Neighbors and the Fast Johnson-Lindenstrauss Transform. In Proceedings of the Thirty-Eighth Annual ACM Symposium on Theory of Computing (Seattle, WA, USA) (STOC ’06) . Association for Computing Machinery, New York, NY, USA, 557–563
Nir Ailon and Bernard Chazelle. 2006 · 2006
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas. 2006 · 2006
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Hadamard Matrices and Their Applications
Kathy J Horadam. 2012 · 2012
Earlier work this paper cites.
Deep Learning with Limited Numerical Precision. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37) , Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1737–1746
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015 · 2015
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.. In CoCo@NIPS
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
QSGD: Communication-Efficient Sgd via Gradient Quantization and Encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Cited alongside, same era.
Dynamic point stochastic rounding algorithm for limited precision arithmetic in Deep Belief Network training. In 2017 8th International IEEE/EMBS Conference on Neural Engineering (NER) . 629–632
Mohaned Essam, Tong Boon Tang, Eric Tatt Wei Ho, and Hsin Chen. 2017 · 2017
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Distributed Mean Estimation With Limited Communication. In International Conference on Machine Learning . PMLR, 3329–3337
Ananda Theertha Suresh, X Yu Felix, Sanjiv Kumar, and H Brendan McMahan. 2017 · 2017
Cited alongside, same era.
Randomized Distributed Mean Estimation: Accuracy vs. Communication
Jakub Konečnỳ and Peter Richtárik. 2018 · 2018
Cited alongside, same era.
Retrieval Augmented Language Model Pre-Training. In ICML 2020: 37th International Conference on Machine Learning , Vol. 1. 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Later among the works it cites.
Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020a · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training deep neural networks with 8-bit floating point numbers. In Proceedings of the 32nd International Conference on Neural Information Processing Systems . 7686–7695
Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan. 2018 · 2018
Cited alongside, same era.
Training and Inference with Integers in Deep Neural Networks. In International Conference on Learning Representations
Shuang Wu, Guoqi Li, Feng Chen, and Luping Shi. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
On the Downstream Performance of Compressed Word Embeddings
Avner May, Jian Zhang, Tri Dao, and Christopher Ré. 2019 · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Cited alongside, same era.
DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 4487–4497
Qingqing Cao, Harsh Trivedi, Aruna Balasubramanian, and Niranjan Balasubramanian. 2020 · 2020
Cited alongside, same era.
DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modeling. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 2925–2937
Jiecao Chen, Liu Yang, Karthik Raman, Michael Bendersky, Jung-Jung Yeh, Yun Zhou, Marc Najork, Danyang Cai, and Ehsan Emadzadeh. 2020 · 2020
Cited alongside, same era.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Later among the works it cites.
TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020 · 2020
Later among the works it cites.
Efficient Document Re-Ranking for Transformers by Precomputing Term Representations. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 49–58
Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder. 2020 · 2020
Later among the works it cites.
DC-BERT: Decoupling Question and Document for Efficient Contextual Encoding. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 1829–1832
Ping Nie, Yuyu Zhang, Xiubo Geng, Arun Ramamurthy, Le Song, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
Simplified TinyBERT: Knowledge Distillation for Document Retrieval. In European Conference on Information Retrieval . 241–248
Xuanang Chen, Ben He, Kai Hui, Le Sun, and Yingfei Sun. 2021 · 2021
Closest in time.
Stochastic Rounding and Its Probabilistic Backward Error Analysis
Michael P. Connolly, Nicholas J. Higham, and Theo Mary. 2021 · 2021
Closest in time.
DRIVE: One-bit Distributed Mean Estimation
Shay Vargaftik, Ran Ben Basat, Amit Portnoy, Gal Mendelson, Yaniv Ben-Itzhak, and Michael Mitzenmacher. 2021 · 2021
Closest in time.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In ICLR 2021: The Ninth International Conference on Learning Representations
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwikj. 2021 · 2021
Closest in time.
Pretrained Transformers for Text Ranking: BERT and Beyond. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Tutorials . Association for Computational Linguistics, Online, 1–4
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Closest in time.