Fetching the paper…
Reading the bibliography…
This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding model that outperforms both OpenAI Ada-002 and OpenAI text-embedding-3-small on the short-context MTEB benchmark and the long context LoCo benchmark.
BIGPATENT: A large-scale dataset for abstractive and coherent summarization
Eva Sharma, Chen Li, and Lu Wang · 1906
Earlier work this paper cites.
Fourier features let networks learn high frequency functions in low dimensional domains, 2020
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng · 2006
Earlier work this paper cites.
Simple English Wikipedia: A new text simplification task
William Coster and David Kauchak · 2011
Earlier work this paper cites.
Overcoming the lack of parallel data in sentence compression
Katja Filippova and Yasemin Altun · 2013
Earlier work this paper cites.
Open Question Answering Over Curated and Extracted Knowledge Bases
Anthony Fader, Luke Zettlemoyer, and Oren Etzioni · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Identifying causal relations using parallel Wikipedia articles
Christopher Hidey and Kathy McKeown · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Character-level convolutional networks for text classification, 2016
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2016
Earlier work this paper cites.
news-please: A generic news crawler and extractor
Felix Hamborg, Norman Meuschke, Corinna Breitinger, and Bela Gipp · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset, 2018
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset, 2018
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
Mixed precision training, 2018
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Amazonqa: A review-based question answering task, 2019
Mansi Gupta, Nitish Kulkarni, Raghuveer Chanda, Anirudha Rayasam, and Zachary C Lipton · 2019
Earlier work this paper cites.
CodeSearchNet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Jianmo Ni, Jiacheng Li, and Julian McAuley · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Representation learning with contrastive predictive coding, 2019
Sgpt: Gpt sentence embeddings for semantic search, 2022
Niklas Muennighoff · 2022
Later among the works it cites.
Text and code embeddings by contrastive pre-training, 2022
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nekoul, Girish Sastry, Gretchen Krueger, David Schnurr, Felipe Petroski Such, Kenny Hsu, Madeleine Thompson, Tabarak Khan, Toki Sherbakov, Joanne Jang, Peter Welinder, and Lilian Weng · 2022
Later among the works it cites.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang · 2022
Later among the works it cites.
SCROLLS: Standardized CompaRison over long language sequences
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, and Omer Levy · 2022
Later among the works it cites.
Text embeddings by weakly-supervised contrastive pre-training, 2022
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Dense passage retrieval for open-domain question answering, 2020
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih · 2020
Cited alongside, same era.
S2orc: The semantic scholar open research corpus, 2020
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Dan S. Weld · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
Noam Shazeer · 2020
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Cited alongside, same era.
Later among the works it cites.
NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., 2023
bloc97 · 2023
Later among the works it cites.
Extending context window of large language models via positional interpolation, 2023
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian · 2023
Later among the works it cites.
Dynamically scaled rope further increases performance of long context llama with zero fine-tuning, 2023
emozilla · 2023
Later among the works it cites.
Jina embeddings: A novel set of high-performance sentence embedding models, 2023
Michael Günther, Louis Milliken, Jonathan Geuter, Georgios Mastrapas, Bo Wang, and Han Xiao · 2023
Later among the works it cites.
Things I’m learning while training superhot., 2023
kaiokendev · 2023
Later among the works it cites.
Towards general text embeddings with multi-stage contrastive learning, 2023
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang · 2023
Later among the works it cites.
Mteb: Massive text embedding benchmark, 2023
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models, 2023
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Later among the works it cites.
Mosaicbert: A bidirectional encoder optimized for fast pretraining, 2023
Jacob Portes, Alex Trott, Sam Havens, Daniel King, Abhinav Venigalla, Moin Nadeem, Nikhil Sardana, Daya Khudia, and Jonathan Frankle · 2023
Later among the works it cites.
In-context retrieval-augmented language models, 2023
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Introducing embed v3, Nov 2023
Nils Reimers, Elliot Choi, Amr Kayid, Alekhya Nandula, Manoj Govindassamy, and Abdullah Elkady · 2023
Later among the works it cites.
Excited to announce voyage embeddings!, Nov 2023
Voyage · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding, 2023
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff · 2023
Later among the works it cites.
Jina embeddings 2: 8192-token general-purpose text embeddings for long documents, 2024
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, Maximilian Werk, Nan Wang, and Han Xiao · 2024
Closest in time.
Long-context retrieval models with monarch mixer, Jan 2024
Jon Saad-Falcon, Dan Fu, and Simran Arora · 2024
Closest in time.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books, 2015
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2048
Closest in time.