Fetching the paper…
Reading the bibliography…
Recent advancements in long-context large language models have attracted significant attention, yet their practical applications often suffer from suboptimal context utilization.
Toward Semantics-Based Answer Pinpointing
Hovy, E.; Gerber, L.; Hermjakob, U.; Lin, C.-Y.; and Ravichandran, D. 2001 · 2001
Earlier work this paper cites.
Learning Question Classifiers
Li, X.; and Roth, D. 2002 · 2002
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2009
Earlier work this paper cites.
The Probabilistic Relevance Framework: BM25 and Beyond
Robertson, S. E.; and Zaragoza, H. 2009 · 2009
Earlier work this paper cites.
DBpedia - A large-scale, multilingual knowledge base extracted from Wikipedia
Lehmann, J.; Isele, R.; Jakob, M.; Jentzsch, A.; Kontokostas, D.; Mendes, P. N.; Hellmann, S.; Morsey, M.; van Kleef, P.; Auer, S.; and Bizer, C. 2015 · 2015
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Johnson, J.; Douze, M.; and Jégou, H. 2017 · 2017
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Johnson, J.; Douze, M.; and Jégou, H. 2019 · 2019
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Chen, M.; Tworek, J.; Jun, H.; and et al. 2021 · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Earlier work this paper cites.
A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Dasigi, P.; Lo, K.; Beltagy, I.; Cohan, A.; Smith, N. A.; and Gardner, M. 2021 · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; Phang, J.; Reynolds, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2021 · 2021
Earlier work this paper cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J.; Lu, Y.; Pan, S.; Wen, B.; and Liu, Y. 2021 · 2021
Earlier work this paper cites.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Thakur, N.; Reimers, N.; Rücklé, A.; Srivastava, A.; and Gurevych, I. 2021 · 2021
Earlier work this paper cites.
ProofNet: A Benchmark for Autoformalizing and Formally Proving Undergraduate-Level Mathematics Problems
Azerbayev, Z.; Piotrowski, B.; and Avigad, J. 2022 · 2022
Earlier work this paper cites.
Improving Language Models by Retrieving from Trillions of Tokens
Borgeaud, S.; Mensch, A.; Hoffmann, J.; Cai, T.; Rutherford, E.; Millican, K.; van den Driessche, G.; Lespiau, J.; Damoc, B.; Clark, A.; de Las Casas, D.; Guy, A.; Menick, J.; Ring, R.; Hennigan, T.; Huang, S.; Maggiore, L.; Jones, C.; Cassirer, A.; Brock, A.; Paganini, M.; Irving, G.; Vinyals, O.; Osindero, S.; Simonyan, K.; Rae, J. W.; Elsen, E.; and Sifre, L. 2022 · 2022
Earlier work this paper cites.
Data Distributional Properties Drive Emergent In-Context Learning in Transformers
Chan, S.; Santoro, A.; Lampinen, A. K.; Wang, J.; Singh, A.; Richemond, P. H.; McClelland, J. L.; and Hill, F. 2022 · 2022
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; and et al. 2022 · 2022
Cited alongside, same era.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; Hennigan, T.; Noland, E.; Millican, K.; van den Driessche, G.; Damoc, B.; Guy, A.; Osindero, S.; Simonyan, K.; Elsen, E.; Rae, J. W.; Vinyals, O.; and Sifre, L. 2022 · 2022
Cited alongside, same era.
Unsupervised Dense Information Retrieval with Contrastive Learning
Izacard, G.; Caron, M.; Hosseini, L.; Riedel, S.; Bojanowski, P.; Joulin, A.; and Grave, E. 2022 · 2022
Cited alongside, same era.
The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design
Understanding In-Context Learning via Supportive Pretraining Data
Han, X.; Simig, D.; Mihaylov, T.; Tsvetkov, Y.; Celikyilmaz, A.; and Wang, T. 2023 · 2023
Closest in time.
Needle In A Haystack - Pressure Testing LLMs
Kamradt, G. 2023 · 2023
Closest in time.
Lost in the Middle: How Language Models Use Long Contexts
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2023 · 2023
Closest in time.
Landmark Attention: Random-Access Infinite Context Length for Transformers
Mohtashami, A.; and Jaggi, M. 2023 · 2023
Closest in time.
YaRN: Efficient Context Window Extension of Large Language Models
Peng, B.; Quesnelle, J.; Fan, H.; and Shippole, E. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Levine, Y.; Wies, N.; Jannai, D.; Navon, D.; Hoshen, Y.; and Shashua, A. 2022 · 2022
Cited alongside, same era.
Solving Quantitative Reasoning Problems with Language Models
Lewkowycz, A.; Andreassen, A. J.; Dohan, D.; Dyer, E.; Michalewski, H.; Ramasesh, V. V.; Slone, A.; Anil, C.; Schlag, I.; Gutman-Solo, T.; Wu, Y.; Neyshabur, B.; Gur-Ari, G.; and Misra, V. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022 · 2022
Cited alongside, same era.
SCROLLS: Standardized CompaRison Over Long Language Sequences
Shaham, U.; Segal, E.; Ivgi, M.; Efrat, A.; Yoran, O.; Haviv, A.; Gupta, A.; Xiong, W.; Geva, M.; Berant, J.; and Levy, O. 2022 · 2022
Cited alongside, same era.
Memorizing Transformers
Wu, Y.; Rabe, M. N.; Hutchins, D.; and Szegedy, C. 2022 · 2022
Cited alongside, same era.
SantaCoder: don’t reach for the stars!
Allal, L. B.; Li, R.; Kocetkov, D.; Mou, C.; Akiki, C.; Ferrandis, C. M.; Muennighoff, N.; Mishra, M.; Gu, A.; Dey, M.; Umapathi, L. K.; Anderson, C. J.; Zi, Y.; Poirier, J. L.; Schoelkopf, H.; Troshin, S.; Abulkhanov, D.; Romero, M.; Lappert, M.; Toni, F. D.; del Río, B. G.; Liu, Q.; Bose, S.; Bhattacharyya, U.; Zhuo, T. Y.; Yu, I.; Villegas, P.; Zocca, M.; Mangrulkar, S.; Lansky, D.; Nguyen, H.; Contractor, D.; Villa, L.; Li, J.; Bahdanau, D.; Jernite, Y.; Hughes, S.; Fried, D.; Guha, A.; de Vries, H.; and von Werra, L. 2023 · 2023
Cited alongside, same era.
Model Card and Evaluations for Claude Models
Anthropic. 2023 · 2023
Cited alongside, same era.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; Hui, B.; Ji, L.; Li, M.; Lin, J.; Lin, R.; Liu, D.; Liu, G.; Lu, C.; Lu, K.; Ma, J.; Men, R.; Ren, X.; Ren, X.; Tan, C.; Tan, S.; Tu, J.; Wang, P.; Wang, S.; Wang, W.; Wu, S.; Xu, B.; Xu, J.; Yang, A.; Yang, H.; Yang, J.; Yang, S.; Yao, Y.; Yu, B.; Yuan, H.; Yuan, Z.; Zhang, J.; Zhang, X.; Zhang, Y.; Zhang, Z.; Zhou, C.; Zhou, J.; Zhou, X.; and Zhu, T. 2023 · 2023
Cited alongside, same era.
Closest in time.
Code Llama: Open Foundation Models for Code
Rozière, B.; Gehring, J.; Gloeckle, F.; Sootla, S.; Gat, I.; Tan, X. E.; Adi, Y.; Liu, J.; Remez, T.; Rapin, J.; Kozhevnikov, A.; Evtimov, I.; Bitton, J.; Bhatt, M.; Ferrer, C. C.; Grattafiori, A.; Xiong, W.; Défossez, A.; Copet, J.; Azhar, F.; Touvron, H.; Martin, L.; Usunier, N.; Scialom, T.; and Synnaeve, G. 2023 · 2023
Closest in time.
Large Language Models Can Be Easily Distracted by Irrelevant Context
Shi, F.; Chen, X.; Misra, K.; Scales, N.; Dohan, D.; Chi, E. H.; Schärli, N.; and Zhou, D. 2023 · 2023
Closest in time.
RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset
TogetherComputer. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Focused Transformer: Contrastive Training for Context Scaling
Tworkowski, S.; Staniszewski, K.; Pacek, M.; Wu, Y.; Michalewski, H.; and Milos, P. 2023 · 2023
Closest in time.
Analysing The Impact of Sequence Composition on Language Model Pre-Training
Zhao, Y.; Qu, Y.; Staniszewski, K.; Tworkowski, S.; Liu, W.; Miłoś, P.; Wu, Y.; and Minervini, P. 2024 · 2023
Closest in time.
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Allen-Zhu, Z.; and Li, Y. 2024 · 2024
Closest in time.
Dubey, A.; Jauhri, A.; and et al., A. P. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, G. 2024 · 2024
Closest in time.
OLMo: Accelerating the Science of Language Models
Groeneveld, D.; Beltagy, I.; Walsh, P.; Bhagia, A.; Kinney, R.; Tafjord, O.; Jha, A. H.; Ivison, H.; Magnusson, I.; Wang, Y.; Arora, S.; Atkinson, D.; Authur, R.; Chandu, K. R.; Cohan, A.; Dumas, J.; Elazar, Y.; Gu, Y.; Hessel, J.; Khot, T.; Merrill, W.; Morrison, J.; Muennighoff, N.; Naik, A.; Nam, C.; Peters, M. E.; Pyatkin, V.; Ravichander, A.; Schwenk, D.; Shah, S.; Smith, W.; Strubell, E.; Subramani, N.; Wortsman, M.; Dasigi, P.; Lambert, N.; Richardson, K.; Zettlemoyer, L.; Dodge, J.; Lo, K.; Soldaini, L.; Smith, N. A.; and Hajishirzi, H. 2024 · 2024
Closest in time.
RULER: What’s the Real Context Size of Your Long-Context Language Models?
Hsieh, C.-P.; Sun, S.; Kriman, S.; Acharya, S.; Rekesh, D.; Jia, F.; Zhang, Y.; and Ginsburg, B. 2024 · 2024
Closest in time.
In-Context Pretraining: Language Modeling Beyond Document Boundaries
Shi, W.; Min, S.; Lomeli, M.; Zhou, C.; Li, M.; Lin, X. V.; Smith, N. A.; Zettlemoyer, L.; tau Yih, W.; and Lewis, M. 2024 · 2024
Closest in time.