Fetching the paper…
Reading the bibliography…
We introduce Lifelong ICL, a problem setting that challenges long-context language models (LMs) to learn a sequence of language tasks through in-context learning (ICL).
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2004
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee · 2004
Earlier work this paper cites.
SemEval-2019 task 3: EmoContext contextual emotion detection in text
Ankush Chatterjee, Kedhar Nath Narahari, Meghana Joshi, and Puneet Agrawal · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
SemEval-2017 task 7: Detection and interpretation of English puns
Tristan Miller, Christian Hempelmann, and Iryna Gurevych · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
The ACL Anthology reference corpus: A reference dataset for bibliographic research in computational linguistics
Steven Bird, Robert Dale, Bonnie Dorr, Bryan Gibson, Mark Joseph, Min-Yen Kan, Dongwon Lee, Brett Powley, Dragomir Radev, and Yee Fan Tan · 2008
Earlier work this paper cites.
Contributions to the study of sms spam filtering: new collection and results
Tiago A. Almeida, José María G. Hidalgo, and Akebo Yamakami · 2011
Earlier work this paper cites.
The Winograd schema challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Pekka Malo, Ankur Sinha, Pyry Takala, Pekka Korhonen, and Jyrki Wallenius · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli · 2014
Earlier work this paper cites.
Identifying appropriate support for propositions in online user comments
Joonsuk Park and Claire Cardie · 2014
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Stop clickbait: Detecting and preventing clickbaits in online news media
Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly · 2016
Earlier work this paper cites.
Emergent: a novel data-set for stance classification
William Ferreira and Andreas Vlachos · 2016
Earlier work this paper cites.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornel Csernai · 2016
Earlier work this paper cites.
SemEval-2016 task 6: Detecting stance in tweets
Saif Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry · 2016
Earlier work this paper cites.
Creating and characterizing a diverse corpus of sarcasm in dialogue
Shereen Oraby, Vrindavan Harrison, Lena Reed, Ernesto Hernandez, Ellen Riloff, and Marilyn Walker · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
The CogALex-V shared task on the corpus-based identification of semantic relations
Enrico Santus, Anna Gladkova, Stefan Evert, and Alessandro Lenci · 2016
Earlier work this paper cites.
Android apps and user feedback: A dataset for software evolution and quality improvement
Giovanni Grano, Andrea Di Sorbo, Francesco Mercaldo, Corrado A Visaggio, Gerardo Canfora, and Sebastiano Panichella · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Predicting human metaphor paraphrase judgments with deep neural networks
Yuri Bizzoni and Shalom Lappin · 2018
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros · 2018
Earlier work this paper cites.
Quora insincere questions classification, 2018
Alex Ellis, Julia Elliott, Paula Griffin, and William Chen · 2018
Earlier work this paper cites.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Jigsaw unintended bias in toxicity classification, 2019
cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
The commitmentbank: Investigating projection in naturally occurring discourse
Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser · 2019
Earlier work this paper cites.
Episodic memory in lifelong language learning
Cyprien de Masson d'Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama · 2019
Earlier work this paper cites.
A surprisingly robust trick for the Winograd schema challenge
Vid Kocijan, Ana-Maria Cretu, Oana-Maria Camburu, Yordan Yordanov, and Thomas Lukasiewicz · 2019
Earlier work this paper cites.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Cited alongside, same era.
Continual lifelong learning in natural language processing: A survey
Magdalena Biesialska, Katarzyna Biesialska, and Marta R. Costa-jussà · 2020
Cited alongside, same era.
Iitjee neet aiims students questions data
Mrutyunjay Biswal · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
An empirical investigation of the role of pre-training in lifelong learning
Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, and Emma Strubell · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee · 2023
Later among the works it cites.
ZeroSCROLLS: A zero-shot benchmark for long text understanding
Uri Shaham, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, and Alice Xiang et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical pre-training for sequence labelling in spoken dialog
Emile Chapuis, Pierre Colombo, Matteo Manica, Matthieu Labeau, and Chloé Clavel · 2020
Cited alongside, same era.
Climate-fever: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2020
Cited alongside, same era.
A dataset for statutory reasoning in tax law entailment and question answering
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme · 2020
Cited alongside, same era.
Explainable automated fact-checking: A survey
Neema Kotonya and Francesca Toni · 2020
Cited alongside, same era.
“I’d rather just go to bed”: Understanding indirect answers
Annie Louis, Dan Roth, and Filip Radlinski · 2020
Cited alongside, same era.
LiMiT: The literal motion in text dataset
Irene Manotas, Ngoc Phuoc An Vo, and Vadim Sheinin · 2020
Cited alongside, same era.
Effective transfer learning for identifying similar questions: Matching user questions to covid-19 faqs
Clara H. McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain · 2020
Cited alongside, same era.
Later among the works it cites.
Focused transformer: Contrastive training for context scaling
Szymon Tworkowski, Konrad Staniszewski, Mikoł aj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Mił oś · 2023
Later among the works it cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun · 2023
Later among the works it cites.
Yi: Open foundation models by 01.ai, 2024
01.AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
Many-shot in-context learning
Rishabh Agarwal, Avi Singh, Lei M Zhang, Bernd Bohnet, Luis Rosias, Stephanie C.Y. Chan, Biao Zhang, Aleksandra Faust, and Hugo Larochelle · 2024
Closest in time.
Make your llm fully utilize the context
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, and Jian-Guang Lou · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Closest in time.
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Closest in time.
In-context learning with long-context models: An in-depth exploration
Amanda Bertsch, Maor Ivgi, Uri Alon, Jonathan Berant, Matthew R. Gormley, and Graham Neubig · 2024
Closest in time.
How cheap talk in climate disclosures relates to climate initiatives, corporate emissions, and reputation risk
Julia Anna Bingler, Mathias Kraus, Markus Leippold, and Nicolas Webersinke · 2024
Closest in time.
When parts are greater than sums: Individual LLM components can outperform full models
Ting-Yun Chang, Jesse Thomason, and Robin Jia · 2024
Closest in time.
C4ai command-r model card, 2024
Cohere for AI · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Data engineering for scaling language models to 128k context
Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team · 2024
Closest in time.
Omer Goldman, Alon Jacovi, Aviv Slobodkin, Aviya Maimon, Ido Dagan, and Reut Tsarfaty · 2024
Closest in time.
Paul graham essays, 2024
Paul Graham · 2024
Closest in time.
LM-infinite: Zero-shot extreme length generalization for large language models
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang · 2024
Closest in time.
RULER: What’s the real context size of your long-context language models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg · 2024
Closest in time.
Can perplexity reflect large language model’s ability in long text understanding?
Yutong Hu, Quzhe Huang, Mingxu Tao, Chen Zhang, and Yansong Feng · 2024
Closest in time.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Closest in time.
Can long-context language models subsume retrieval, rag, sql, and more?, 2024
Jinhyuk Lee, Anthony Chen, Zhuyun Dai, Dheeru Dua, Devendra Singh Sachan, Michael Boratko, Yi Luan, Sébastien M. R. Arnold, Vincent Perot, Siddharth Dalmia, Hexiang Hu, Xudong Lin, Panupong Pasupat, Aida Amini, Jeremy R. Cole, Sebastian Riedel, Iftekhar Naim, Ming-Wei Chang, and Kelvin Guu · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models
Mosh Levy, Alon Jacoby, and Yoav Goldberg · 2024
Closest in time.
Long-context llms struggle with long in-context learning
Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen · 2024
Closest in time.
Aditya Sharma, Michael Saxon, and William Yang Wang · 2024
Closest in time.
Continual learning of large language models: A comprehensive survey
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, and Hao Wang · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
A benchmark for learning to translate a new language from one grammar book
Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, and Luke Melas-Kyriazi · 2024
Closest in time.
Llama-2-7b-32k model card, 2024
TogetherAI · 2024
Closest in time.
Investigating the effectiveness of task-agnostic prefix prompt for instruction following
Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo · 2024
Closest in time.
Helmet: How to evaluate long-context language models effectively and thoroughly, 2024
Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, and Danqi Chen · 2024
Closest in time.
Stella-en-1.5b-v5 model card, 2024
Dun Zhang · 2024
Closest in time.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Closest in time.
PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee · 2052
Closest in time.
“liar, liar pants on fire”: A new benchmark dataset for fake news detection
William Yang Wang · 2067
Closest in time.
SemEval-2015 task 12: Aspect based sentiment analysis
Maria Pontiki, Dimitris Galanis, Haris Papageorgiou, Suresh Manandhar, and Ion Androutsopoulos · 2082
Closest in time.