Fetching the paper…
Reading the bibliography…
Current benchmarks like Needle-in-a-Haystack (NIAH), Ruler, and Needlebench focus on models' ability to understand long-context input sequences but fail to capture a critical dimension: the generation of high-quality long-form text.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Strategies for structuring story generation
Angela Fan, Mike Lewis, and Yann N. Dauphin · 2019
Earlier work this paper cites.
Huggingface’s Transformers: State-of-the-art natural language processing
Thomas Wolf et al · 2019
Earlier work this paper cites.
PAIR: planning and iterative refinement in pre-trained transformers for long text generation
Xinyu Hua and Lu Wang · 2020
Earlier work this paper cites.
Plan ahead: Self-supervised text planning for paragraph completion task
Dongyeop Kang and Eduard H. Hovy · 2020
Earlier work this paper cites.
MEGATRON-CNTRL: controllable story generation with external knowledge using large-scale language models
Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, and Bryan Catanzaro · 2020
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner · 2021
Earlier work this paper cites.
PLANET: dynamic content planning in autoregressive transformers for long-form text generation
Zhe Hu, Hou Pong Chan, Jiachen Liu, Xinyan Xiao, Hua Wu, and Lifu Huang · 2022
Earlier work this paper cites.
ChatGPT, 2022
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
ASQA: factoid questions meet long-form answers
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang · 2022
Earlier work this paper cites.
Beyond goldfish memory: Long-term open-domain conversation
Jing Xu, Arthur Szlam, and Jason Weston · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai et al · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
David Cheng-Han Chiang and Hung-yi Lee · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Needle In A Haystack - pressure testing LLMs
Gregory Kamradt · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with paged attention
Woosuk Kwon et al · 2023
Cited alongside, same era.
QASA: advanced question answering on scientific articles
Yoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang, Jaehyeon Kim, Hong-In Lee, and Moontae Lee · 2023
Cited alongside, same era.
Prd: Peer rank and discussion improve large language model based evaluations
Introducing claude 2.1, 2024a
Anthropic · 2024
Closest in time.
Introducing the next generation of claude, 2024b
Anthropic · 2024
Closest in time.
Longwriter: Unleashing 10,000+ word generation from long context llms
Yushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng, Siqi Zhu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ruosen Li, Teerth Patel, and Xinya Du · 2023
Cited alongside, same era.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Cited alongside, same era.
La plateforme, 2023
Mistral.AI · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
ZeroSCROLLS: A zero-shot benchmark for long text understanding
Uri Shaham, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy · 2023
Cited alongside, same era.
Large language models are not yet human-level evaluators for abstractive summarization
Chenhui Shen, Liying Cheng, Xuan-Phi Nguyen, Yang You, and Lidong Bing · 2023
Cited alongside, same era.
Albert Q Jiang et al · 2024
Closest in time.
Longlamp: A benchmark for personalized long-form text generation
Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi, Ryan A Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, et al · 2024
Closest in time.
Needlebench: Can llms do retrieval and reasoning in 1 million context window?
Mo Li, Songyang Zhang, Yunxin Liu, and Kai Chen · 2024
Closest in time.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Korhonen, and Nigel Collier · 2024
Closest in time.
Gpt-4o mini: Advancing cost-efficient intelligence, 2024a
OpenAI · 2024
Closest in time.
Hello gpt-4o, 2024b
OpenAI · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, and Linqi Song · 2024
Closest in time.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.