Fetching the paper…
Reading the bibliography…
Retrieval-augmented generation (RAG) has recently emerged as a promising solution for incorporating up-to-date or domain-specific knowledge into large language models (LLMs) and improving LLM factuality, but is predominantly studied in English-only settings.
Fasttext.zip: Compressing text classification models
Joulin, A., Grave, E., Bojanowski, P., Douze, M., Jégou, H., and Mikolov, T. (2016) · 2016
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Nguyen, T., Rosenberg, M., Song, X., Gao, J., Tiwary, S., Majumder, R., and Deng, L. (2016) · 2016
Earlier work this paper cites.
Bag of tricks for efficient text classification
Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T. (2017) · 2017
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al. (2019) · 2019
Earlier work this paper cites.
TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J. (2020) · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020) · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. (2020) · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020) · 2020
Earlier work this paper cites.
XOR QA: Cross-lingual open-retrieval question answering
Asai, A., Kasai, J., Clark, J., Lee, K., Choi, E., and Hajishirzi, H. (2021a) · 2021
Earlier work this paper cites.
MKQA: A linguistically diverse benchmark for multilingual open domain question answering
Longpre, S., Lu, Y., and Daiber, J. (2021) · 2021
Earlier work this paper cites.
MIA 2022 shared task: Evaluating cross-lingual open-retrieval question answering for 16 diverse languages
Asai, A., Longpre, S., Kasai, J., Lee, C.-H., Zhang, R., Hu, J., Yamada, I., Clark, J. H., and Choi, E. (2022) · 2022
Cited alongside, same era.
Cross-lingual open-domain question answering with answer sentence generation
Muller, B., Soldaini, L., Koncel-Kedziorski, R., Lind, E., and Moschitti, A. (2022) · 2022
Cited alongside, same era.
Ask me anything in your native language
Sorokin, N., Abulkhanov, D., Piontkovskaya, I., and Malykh, V. (2022) · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Team, N., Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J. (2022) · 2022
Cited alongside, same era.
Active retrieval augmented generation
Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G. (2023) · 2023
Recomp: Improving retrieval-augmented lms with compression and selective augmentation
Xu, F., Shi, W., and Choi, E. (2023) · 2023
Later among the works it cites.
Language versatilists vs. specialists: An empirical revisiting on multilingual transfer ability
Ye, J., Tao, X., and Kong, L. (2023) · 2023
Later among the works it cites.
Self-RAG: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. (2024) · 2024
Closest in time.
Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation
Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. (2024) · 2024
Closest in time.
Zero-shot cross-lingual transfer in instruction tuning of large language models
Chirkova, N. and Nikoulina, V. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling
Kim, D., Park, C., Kim, S., Lee, W., Song, W., Kim, Y., Kim, H., Kim, Y., Lee, H., Kim, J., Ahn, C., Yang, S., Lee, S., Park, H., Gim, G., Cha, M., Lee, H., and Kim, S. (2023) · 2023
Cited alongside, same era.
Query rewriting in retrieval-augmented large language models
Ma, X., Gong, Y., He, P., Zhao, H., and Duan, N. (2023) · 2023
Cited alongside, same era.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y. (2023) · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. (2023) · 2023
Cited alongside, same era.
Learning to filter context for retrieval-augmented generation
Wang, Z., Araki, J., Jiang, Z., Parvez, M. R., and Neubig, G. (2023) · 2023
Cited alongside, same era.
One question answering model for many languages with cross-lingual dense passage retrieval
Asai, A., Yu, X., Kasai, J., and Hajishirzi, H. (2021b)
Cited in the paper.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2024) · 2024
Closest in time.
Sure: Improving open-domain question answering of LLMs via summarized retrieval
Kim, J., Nam, J., Mo, S., Park, J., Lee, S.-W., Seo, M., Ha, J.-W., and Shin, J. (2024) · 2024
Closest in time.
Bergen: A benchmarking library for retrieval-augmented generation
Rau, D., Déjean, H., Chirkova, N., Formal, T., Wang, S., Nikoulina, V., and Clinchant, S. (2024) · 2024
Closest in time.
Making retrieval-augmented language models robust to irrelevant context
Yoran, O., Wolfson, T., Ram, O., and Berant, J. (2024) · 2024
Closest in time.