Fetching the paper…
Reading the bibliography…
Instruction following is one of the fundamental capabilities of large language models (LLMs).
Abstract meaning representation for sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider · 2013
Earlier work this paper cites.
Sequence to backward and forward sequences: A content-introducing approach to generative short-text conversation
Lili Mou, Yiping Song, Rui Yan, Ge Li, Lu Zhang, and Zhi Jin · 2016
Earlier work this paper cites.
Neural amr: Sequence-to-sequence models for parsing and generation
Ioannis Konstas, Srinivasan Iyer, Mark Yatskar, Yejin Choi, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
MojiTalk: Generating emotional responses at scale
Xianda Zhou and William Yang Wang · 2018
Earlier work this paper cites.
Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer
Sudha Rao and Joel Tetreault · 2018
Earlier work this paper cites.
Adversarially regularized autoencoders
Junbo Jake Zhao, Yoon Kim, Kelly Zhang, Alexander M. Rush, and Yann LeCun · 2018
Earlier work this paper cites.
Emotional chatting machine: Emotional conversation generation with internal and external memory
Hao Zhou, Minlie Huang, Tianyang Zhang, Xiaoyan Zhu, and Bing Liu · 2018
Earlier work this paper cites.
Measuring compositionality in representation learning
Jacob Andreas · 2019
Earlier work this paper cites.
Cogs: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer · 2020
Earlier work this paper cites.
POINTER: constrained progressive text generation via insertion-based generative pre-training
Yizhe Zhang, Guoyin Wang, Chunyuan Li, Zhe Gan, Chris Brockett, and Bill Dolan · 2020
Earlier work this paper cites.
Towards question-answering as an automatic metric for evaluating the content quality of a summary
Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth · 2021
Earlier work this paper cites.
Span-based semantic parsing for compositional generalization
Jonathan Herzig and Jonathan Berant · 2021
Earlier work this paper cites.
On compositional generalization of neural machine translation
Yafu Li, Yongjing Yin, Yulong Chen, and Yue Zhang · 2021
Earlier work this paper cites.
StylePTB: A compositional benchmark for fine-grained controllable text style transfer
Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Improving compositional generalization with self-training for data-to-text generation
Sanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale, Ankur Parikh, and Emma Strubell · 2022
GLM-130B: an open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang · 2023
Later among the works it cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Later among the works it cites.
Baichuan 2
Baichuan-Inc · 2023
Later among the works it cites.
Infobench: Evaluating instruction following ability in large language models
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu · 2024
Closest in time.
Can large language models understand real-world complex instructions?
Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen, Jin Xiao, Qianxi He, Xunzhe Zhou, Jiaqing Liang, and Yanghua Xiao · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Disentangled sequence to sequence learning for compositional generalization
Hao Zheng and Mirella Lapata · 2022
Cited alongside, same era.
Why is constrained neural language generation particularly challenging?
Cristina Garbacea and Qiaozhu Mei · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Cited alongside, same era.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
DecompEval: Evaluating generated texts as unsupervised decomposed question answering
Pei Ke, Fei Huang, Fei Mi, Yasheng Wang, Qun Liu, Xiaoyan Zhu, and Minlie Huang · 2023
Cited alongside, same era.
Benchmarking large language models on controllable generation under diversified instructions
Yihan Chen, Benfeng Xu, Quan Wang, Yi Liu, and Zhendong Mao · 2024
Closest in time.
Chain-of-instructions: Compositional instruction tuning on large language models
Shirley Anugrah Hayati, Taehee Jung, Tristan Bodding-Long, Sudipta Kar, Abhinav Sethy, Joo-Kyung Kim, and Dongyeop Kang · 2024
Closest in time.
Fofo: A benchmark to evaluate llms’ format-following capability
Congying Xia, Chen Xing, Jiangshu Du, Xinyi Yang, Yihao Feng, Ran Xu, Wenpeng Yin, and Caiming Xiong · 2024
Closest in time.
Struc-bench: Are large language models good at generating complex structured tabular data?
Xiangru Tang, Yiming Zong, Jason Phang, Yilun Zhao, Wangchunshu Zhou, Arman Cohan, and Mark Gerstein · 2024
Closest in time.
Benchmarking and improving compositional generalization of multi-aspect controllable text generation
Tianqi Zhong, Zhaoyi Li, Quan Wang, Linqi Song, Ying Wei, Defu Lian, and Zhendong Mao · 2024
Closest in time.
Introducing the next generation of claude, 2024
Anthropic · 2024
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al · 2024
Closest in time.