Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts.
The Role of Understanding in Solving Word Problems
D. D. Cummins, W. Kintsch, K. Reusser, and R. Weimer · 1988
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Open information extraction from the web
M. Banko, M. J. Cafarella, S. Soderland, M. Broadhead, and O. Etzioni · 2007
Earlier work this paper cites.
Open language learning for information extraction
Mausam, M. Schmitz, S. Soderland, R. Bart, and O. Etzioni · 2012
Earlier work this paper cites.
Modeling biological processes for reading comprehension
J. Berant, V. Srikumar, P.-C. Chen, A. Vander Linden, B. Harding, B. Huang, P. Clark, and C. D. Manning · 2014
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
P. Pasupat and P. Liang · 2015
Earlier work this paper cites.
End-to-end neural coreference resolution
K. Lee, L. He, M. Lewis, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Multi-hop knowledge graph reasoning with reward shaping
X. V. Lin, R. Socher, and C. Xiong · 2018
Earlier work this paper cites.
Supervised open information extraction
G. Stanovsky, J. Michael, L. Zettlemoyer, and I. Dagan · 2018
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
BERT for coreference resolution: Baselines and analysis
M. Joshi, O. Levy, L. Zettlemoyer, and D. Weld · 2019
Earlier work this paper cites.
IIRC: A dataset of incomplete information reading comprehension questions
J. Ferguson, M. Gardner, H. Hajishirzi, T. Khot, and P. Dasigi · 2020
Earlier work this paper cites.
Injecting numerical reasoning skills into language models
M. Geva, A. Gupta, and J. Berant · 2020
Earlier work this paper cites.
Span model for open information extraction on accurate corpus
J. Zhan and H. Zhao · 2020
Earlier work this paper cites.
PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization
J. Zhang, Y. Zhao, M. Saleh, and P. Liu · 2020
Earlier work this paper cites.
CDLM: Cross-document language modeling
A. Caciularu, A. Cohan, I. Beltagy, M. Peters, A. Cattan, and I. Dagan · 2021
Earlier work this paper cites.
Finqa: A dataset of numerical reasoning over financial data
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T. Huang, B. R. Routledge, and W. Y. Wang · 2021
Earlier work this paper cites.
iFacetSum: Coreference-based interactive faceted summarization for multi-document exploration
E. Hirsch, A. Eirew, O. Shapira, A. Caciularu, A. Cattan, O. Ernst, R. Pasunuru, H. Ronen, M. Bansal, and I. Dagan · 2021
Earlier work this paper cites.
Coreference resolution without span representations
Y. Kirstain, O. Ram, and O. Levy · 2021
Earlier work this paper cites.
Long context question answering via supervised contrastive learning
A. Caciularu, I. Dagan, J. Goldberger, and A. Cohan · 2022
Earlier work this paper cites.
News Summarization and Evaluation in the Era of GPT-3
T. Goyal, J. J. Li, and G. Durrett · 2022
Earlier work this paper cites.
Open-Vocabulary Argument Role Prediction For Event Extraction
Y. Jiao, S. Li, Y. Xie, M. Zhong, H. Ji, and J. Han · 2022
Earlier work this paper cites.
Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models
J. Ni, G. Hernandez Abrego, N. Constant, J. Ma, K. Hall, D. Cer, and Y. Yang · 2022
Cited alongside, same era.
Talm: Tool augmented language models
A. Parisi, Y. Zhao, and N. Fiedel · 2022
Cited alongside, same era.
Text-to-table: A new way of information extraction
X. Wu, J. Zhang, and H. Li · 2022
Cited alongside, same era.
Turning tables: Generating examples from semi-structured tables for endowing language models with reasoning skills
O. Yoran, A. Talmor, and J. Berant · 2022
Cited alongside, same era.
MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data
Y. Zhao, Y. Li, C. Li, and R. Zhang · 2022
Cited alongside, same era.
Large language model is not a good few-shot information extractor, but a good reranker for hard samples!
Y. Ma, Y. Cao, Y. Hong, and A. Sun · 2023
Later among the works it cites.
ZEROTOP: Zero-shot task-oriented semantic parsing using large language models
D. Mekala, J. Wolfe, and S. Roy · 2023
Later among the works it cites.
Augmented language models: a survey
G. Mialon, R. Dessi, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Roziere, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, E. Grave, Y. LeCun, and T. Scialom · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Later among the works it cites.
Toolformer: Language Models Can Teach Themselves to Use Tools
T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Amouyal, T. Wolfson, O. Rubin, O. Yoran, J. Herzig, and J. Berant · 2023
Cited alongside, same era.
Leveraging Code to Improve In-context Learning for Semantic Parsing
B. Bogin, S. Gupta, P. Clark, and A. Sabharwal · 2023
Cited alongside, same era.
Peek across: Improving multi-document modeling via cross-document question-answering
A. Caciularu, M. Peters, J. Goldberger, I. Dagan, and A. Cohan · 2023
Cited alongside, same era.
Compositional semantic parsing with large language models
A. Drozdov, N. Schärli, E. Akyürek, N. Scales, X. Song, X. Chen, O. Bousquet, and D. Zhou · 2023
Cited alongside, same era.
Why Word Problems are Hard for High School Math Students: Problem Formulation and Disciplinary Literacy
E. C. Elliott · 2023
Cited alongside, same era.
Complexity-based prompting for multi-step reasoning
Y. Fu, H. Peng, A. Sabharwal, P. Clark, and T. Khot · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini-Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Don’t add, don’t miss: Effective content preserving generation from pre-selected text spans
A. Slobodkin, A. Caciularu, E. Hirsch, and I. Dagan · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. V. Le, and E. H. Chi · 2023
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges
J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Anthropic · 2024
Closest in time.
A Survey on Evaluation of Large Language Models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. S. Yu, Q. Yang, and X. Xie · 2024
Closest in time.
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
D. Das, D. Banerjee, S. Aditya, and A. Kulkarni · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma-Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Coverbench: A challenging benchmark for complex claim verification
A. Jacovi, M. Ambar, E. Ben-David, U. Shaham, A. Feder, M. Geva, D. Marcus, and A. Caciularu · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Łukasz Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Łukasz Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. Mély, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. O’Keefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph · 2024
Closest in time.
tinyBenchmarks: evaluating LLMs with fewer examples
F. M. Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
MuSR: Testing the limits of chain-of-thought with multistep soft reasoning
Z. R. Sprague, X. Ye, K. Bostrom, S. Chaudhuri, and G. Durrett · 2024
Closest in time.
FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models
Z. Yuan, K. Wang, S. Zhu, Y. Yuan, J. Zhou, Y. Zhu, and W. Wei · 2024
Closest in time.
UniversalNER: Targeted distillation from large language models for open named entity recognition
W. Zhou, S. Zhang, Y. Gu, M. Chen, and H. Poon · 2024
Closest in time.
FanOutQA: Multi-Hop, Multi-Document Question Answering for Large Language Models
A. Zhu, A. Hwang, L. Dugan, and C. Callison-Burch · 2024
Closest in time.