Fetching the paper…
Reading the bibliography…
Robustness has become a critical attribute for the deployment of RAG systems in real-world applications.
Ms marco: A human generated machine reading comprehension dataset
Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., et al · 2016
Earlier work this paper cites.
Causality-based feature selection: Methods and evaluations
Yu, K., Guo, X., Liu, L., Li, J., Wang, H., Ling, Z., and Wu, X · 2020
Earlier work this paper cites.
Unsupervised dense information retrieval with contrastive learning
Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J · 2021
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., and Wang, H · 2023
Earlier work this paper cites.
Atlas: Few-shot learning with retrieval augmented language models
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2023
Earlier work this paper cites.
Kang, H., Ni, J., and Yao, H · 2023
Earlier work this paper cites.
Kuhn, L., Gal, Y., and Farquhar, S · 2023
Earlier work this paper cites.
Recall: A benchmark for llms robustness against external counterfactual knowledge
Liu, Y., Huang, L., Li, S., Chen, S., Zhou, H., Meng, F., Zhou, J., and Sun, X · 2023
Earlier work this paper cites.
Spurious features everywhere-large-scale detection of harmful spurious features in imagenet
Neuhaus, Y., Augustin, M., Boreiko, V., and Hein, M · 2023
Cited alongside, same era.
A new benchmark and reverse validation method for passage-level hallucination detection
Yang, S., Sun, R., and Wan, X · 2023
Cited alongside, same era.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Zhang, Y., Zhenqiang Gong, N., et al · 2023
Cited alongside, same era.
Influence of external information on large language models mirrors social cognitive patterns
Bian, N., Lin, H., Liu, P., Lu, Y., Zhang, C., He, B., Han, X., and Sun, L · 2024
Cited alongside, same era.
The power of noise: Redefining retrieval for rag systems
Cuconasu, F., Trappolini, G., Siciliano, F., Filice, S., Campagnano, C., Maarek, Y., Tonellotto, N., and Silvestri, F · 2024
Cited alongside, same era.
Does prompt formatting have any impact on llm performance?
He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F. X., and Hasan, S · 2024
Later among the works it cites.
Better zero-shot reasoning with role-play prompting
Kong, A., Zhao, S., Chen, H., Li, Q., Qin, Y., Sun, R., Zhou, X., Wang, E., and Dong, X · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2024
Later among the works it cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. and Hruschka, E · 2024
Later among the works it cites.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural retrievers are biased towards llm-generated content
Dai, S., Zhou, Y., Pang, L., Liu, W., Hu, X., Liu, Y., Zhang, X., Wang, G., and Xu, J · 2024
Cited alongside, same era.
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H · 2024
Cited alongside, same era.
Detecting hallucinations in large language models using semantic entropy
Farquhar, S., Kossen, J., Kuhn, L., and Gal, Y · 2024
Cited alongside, same era.
Ragged edges: The double-edged sword of retrieval-augmented chatbots
Feldman, P., Foulds, J. R., and Pan, S · 2024
Cited alongside, same era.
Similarity is not all you need: Endowing retrieval augmented generation with multi layered thoughts
Gan, C., Yang, D., Hu, B., Zhang, H., Li, S., Liu, Z., Shen, Y., Ju, L., Zhang, Z., Gu, J., et al · 2024
Cited alongside, same era.
Benchmarking large language models in retrieval-augmented generation
Chen, J., Lin, H., Han, X., and Sun, L
Cited in the paper.
Chen, X., He, B., Lin, H., Han, X., Wang, T., Cao, B., Sun, L., and Sun, Y
Cited in the paper.
Large language models for data annotation and synthesis: A survey
Tan, Z., Li, D., Wang, S., Beigi, A., Jiang, B., Bhattacharjee, A., Karami, M., Li, J., Cheng, L., and Liu, H · 2024
Later among the works it cites.
Optimizing language model’s reasoning abilities with weak supervision
Tong, Y., Wang, S., Li, D., Wang, Y., Han, S., Lin, Z., Huang, C., Huang, J., and Shang, J · 2024
Later among the works it cites.
Retrieval meets long context large language models
Xu, P., Ping, W., Wu, X., McAfee, L., Zhu, C., Liu, Z., Subramanian, S., Bakhturina, E., Shoeybi, M., and Catanzaro, B · 2024
Later among the works it cites.
Trustworthiness in retrieval-augmented generation systems: A survey
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T.-Y., and Yu, P. S · 2024
Later among the works it cites.
Prosa: Assessing and understanding the prompt sensitivity of llms
Zhuo, J., Zhang, S., Fang, X., Duan, H., Lin, D., and Chen, K · 2024
Later among the works it cites.