Fetching the paper…
Reading the bibliography…
How much private information do text embeddings reveal about the original text? We investigate the problem of embedding \textit{inversion}, reconstructing the full text represented in dense text embeddings.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A Survey of Text Clustering Algorithms , pages 77–128. Springer US, Boston, MA
Charu C. Aggarwal and ChengXiang Zhai. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc V. Le and Tomas Mikolov. 2014 · 2014
Earlier work this paper cites.
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Understanding deep image representations by inverting them
Aravindh Mahendran and Andrea Vedaldi. 2014 · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Jozefowicz, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Inverting visual representations with convolutional networks
Alexey Dosovitskiy and Thomas Brox. 2016 · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair E.W. Johnson, Tom J. Pollard, Lu Shen, Li-wei H. Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. 2016 · 2016
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018 · 2018
Earlier work this paper cites.
Toward controlled generation of text
Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P. Xing. 2018 · 2018
Earlier work this paper cites.
Disentangled representation learning for non-parallel text style transfer
Vineet John, Lili Mou, Hareesh Bahuleyan, and Olga Vechtomova. 2018 · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
Exploiting unintended feature leakage in collaborative learning
Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov. 2018 · 2018
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Ligeng Zhu, Zhijian Liu, and Song Han. 2019 · 2019
Earlier work this paper cites.
Exploring the privacy-preserving properties of word embeddings: Algorithmic validation study
Mohamed Abdalla, Moustafa Abdalla, Graeme Hirst, and Frank Rudzicz. 2020 · 2020
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Cited alongside, same era.
Vec2face: Unveil human faces from their blackbox features in face recognition
Chi Nhan Duong, Thanh-Dat Truong, Kha Gia Quach, Hung Bui, Kaushik Roy, and Khoa Luu. 2020 · 2020
Cited alongside, same era.
Inverting gradients – how easy is it to break privacy in federated learning?
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, and Michael Moeller. 2020 · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih. 2020 · 2020
Cited alongside, same era.
Toward privacy-preserving text embedding similarity with homomorphic encryption
Donggyu Kim, Garam Lee, and Sungwoo Oh. 2022 · 2022
Later among the works it cites.
Diffusion-lm improves controllable text generation
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto. 2022 · 2022
Later among the works it cites.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nekoul, Girish Sastry, Gretchen Krueger, David Schnurr, Felipe Petroski Such, Kenny Hsu, Madeleine Thompson, Tabarak Khan, Toki Sherbakov, Joanne Jang, Peter Welinder, and Lilian Weng. 2022 · 2022
Later among the works it cites.
Large-scale application of named entity recognition to biomedicine and epidemiology
Shaina Raza, Deepak John Reji, Femi Shajan, and Syed Raza Bashir. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John X. Morris. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Information leakage in embedding models
Congzheng Song and Ananth Raghunathan. 2020 · 2020
Cited alongside, same era.
Adversarial semantic collisions
Congzheng Song, Alexander M. Rush, and Vitaly Shmatikov. 2020 · 2020
Cited alongside, same era.
idlg: Improved deep leakage from gradients
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. 2020 · 2020
Cited alongside, same era.
Does bert pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C. Wallace. 2021 · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H. Bermano. 2021 · 2021
Cited alongside, same era.
Large dual encoders are generalizable retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2021 · 2021
Cited alongside, same era.
Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Retromae: Pre-training retrieval-oriented language models via masked auto-encoder
Shitao Xiao, Zheng Liu, Yingxia Shao, and Zhao Cao. 2022 · 2022
Later among the works it cites.
Sentence embedding encoders are easy to steal but hard to defend
Adam Dziedzic, Franziska Boenisch, Mingjian Jiang, Haonan Duan, and Nicolas Papernot. 2023 · 2023
Closest in time.
Hwchase17/langchain: building applications with llms through composability
LangChain. 2023 · 2023
Closest in time.
Haoran Li, Mingshi Xu, and Yangqiu Song. 2023 · 2023
Closest in time.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023 · 2023
Closest in time.
Pinecone
Pinecone. 2023 · 2023
Closest in time.
Qdrant - vector database
Qdrant. 2023 · 2023
Closest in time.
What are you token about? dense retrieval as distributions over the vocabulary
Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, and Amir Globerson. 2023 · 2023
Closest in time.
Lexmae: Lexicon-bottlenecked pretraining for large-scale retrieval
Tao Shen, Xiubo Geng, Chongyang Tao, Can Xu, Xiaolong Huang, Binxing Jiao, Linjun Yang, and Daxin Jiang. 2023 · 2023
Closest in time.
Vdaas/vald: Vald. a highly scalable distributed vector search engine
Vdaas. 2023 · 2023
Closest in time.
Simlm: Pre-training with representation bottleneck for dense passage retrieval
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2023 · 2023
Closest in time.
Weaviate - vector database
Weaviate. 2023 · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Closest in time.