Fetching the paper…
Reading the bibliography…
Various tasks, such as summarization, multi-hop question answering, or coreference resolution, are naturally phrased over collections of real-world documents.
The cambridge dictionary of statistics
B. S. Everitt and A. Skrondal · 2010
Earlier work this paper cites.
Using a sledgehammer to crack a nut? lexical diversity and event coreference resolution
A. Cybulska and P. Vossen · 2014
Earlier work this paper cites.
Topic Concentration in Query Focused Summarization Datasets
T. Baumel, R. Cohen, and M. Elhadad · 2016
Earlier work this paper cites.
Constructing datasets for multi-hop reading comprehension across documents
J. Welbl, P. Stenetorp, and S. Riedel · 2018
Earlier work this paper cites.
Deep dominance - how to properly compare deep neural models
R. Dror, S. Shlomov, and R. Reichart · 2019
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
A. Fabbri, I. Li, T. She, S. Li, and D. Radev · 2019
Earlier work this paper cites.
What is a replication?, Oct 2019
E. Machery · 2019
Earlier work this paper cites.
Multi-hop reading comprehension across multiple documents by reasoning over heterogeneous graphs
M. Tu, G. Wang, J. Huang, Y. Tang, X. He, and B. Zhou · 2019
Earlier work this paper cites.
Extractive Opinion Summarization in Quantized Transformer Spaces
S. Angelidis, R. K. Amplayo, Y. Suhara, X. Wang, and M. Lapata · 2021
Earlier work this paper cites.
CDLM: Cross-document language modeling
A. Caciularu, A. Cohan, I. Beltagy, M. Peters, A. Cattan, and I. Dagan · 2021
Earlier work this paper cites.
Realistic evaluation principles for cross-document coreference resolution
A. Cattan, A. Eirew, G. Stanovsky, M. Joshi, and I. Dagan · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
WEC: Deriving a large-scale cross-document event coreference dataset from Wikipedia
A. Eirew, A. Cattan, and I. Dagan · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
iFacetSum: Coreference-based interactive faceted summarization for multi-document exploration
E. Hirsch, A. Eirew, O. Shapira, A. Caciularu, A. Cattan, O. Ernst, R. Pasunuru, H. Ronen, M. Bansal, and I. Dagan · 2021
Earlier work this paper cites.
Long context question answering via supervised contrastive learning
A. Caciularu, I. Dagan, J. Goldberger, and A. Cohan · 2022
Cited alongside, same era.
Cross-document event coreference search: Task, dataset and modeling
A. Eirew, A. Caciularu, and I. Dagan · 2022
Cited alongside, same era.
Proposition-level clustering for multi-document summarization
O. Ernst, A. Caciularu, O. Shapira, R. Pasunuru, M. Bansal, J. Goldberger, and I. Dagan · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp · 2022
Cited alongside, same era.
SCROLLS: Standardized CompaRison over long language sequences
U. Shaham, E. Segal, M. Ivgi, A. Efrat, O. Yoran, A. Haviv, A. Gupta, W. Xiong, M. Geva, J. Berant, and O. Levy · 2022
Cited alongside, same era.
MuSiQue: Multihop questions via single-hop question composition
Efficient benchmarking (of language models)
Y. Perlitz, E. Bandel, A. Gera, O. Arviv, L. Ein-Dor, E. Shnarch, N. Slonim, M. Shmueli-Scheuer, and L. Choshen · 2023
Later among the works it cites.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
A. Rao, S. Vashistha, A. Naik, S. Aditya, and M. Choudhury · 2023
Later among the works it cites.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, 2023
M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr · 2023
Later among the works it cites.
ZeroSCROLLS: A zero-shot benchmark for long text understanding
U. Shaham, M. Ivgi, A. Efrat, J. Berant, and O. Levy · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal · 2022
Cited alongside, same era.
PRIMERA: Pyramid-based masked sentence pre-training for multi-document summarization
W. Xiao, I. Beltagy, G. Carenini, and A. Cohan · 2022
Cited alongside, same era.
LinkBERT: Pretraining language models with document links
M. Yasunaga, J. Leskovec, and P. Liang · 2022
Cited alongside, same era.
OpenAsp: A benchmark for multi-document open aspect-based summarization
S. Amar, L. Schiff, O. Ernst, A. Shefer, O. Shapira, and I. Dagan · 2023
Cited alongside, same era.
Peek across: Improving multi-document modeling via cross-document question-answering
A. Caciularu, M. Peters, J. Goldberger, I. Dagan, and A. Cohan · 2023
Cited alongside, same era.
Revisiting sentence union generation as a testbed for text consolidation
E. Hirsch, V. Pyatkin, R. Wolhandler, A. Caciularu, A. Shefer, and I. Dagan · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, et al · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024
M. Levy, A. Jacoby, and Y. Goldberg · 2024
Closest in time.
Lost in the Middle: How Language Models Use Long Contexts
N. F. Liu, K. Lin, J. Hewitt, A. i. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2024
Closest in time.
State of what art? a call for multi-prompt llm evaluation, 2024
M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, and G. Stanovsky · 2024
Closest in time.
Multi-review fusion-in-context, 2024
A. Slobodkin, O. Shapira, R. Levy, and I. Dagan · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
A. Voronov, L. Wolf, and M. Ryabinin · 2024
Closest in time.