Fetching the paper…
Reading the bibliography…
Instruction fine-tuning (IFT) elicits instruction following capabilities and steers the behavior of large language models (LLMs) via supervised learning.
Counterfactual fairness
Kusner, M. J., Loftus, J., Russell, C., and Silva, R · 2017
Earlier work this paper cites.
Fair inference on outcomes
Nabi, R. and Shpitser, I · 2018
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Teaching pretrained models with commonsense reasoning: A preliminary kb-based approach
Li, S., Chen, J., and Yu, D · 2019
Earlier work this paper cites.
Counterfactual story reasoning and generation
Qin, L., Bosselut, A., Holtzman, A., Bhagavatula, C., Clark, E., and Choi, Y · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Qin, L., Shwartz, V., West, P., Bhagavatula, C., Hwang, J., Bras, R. L., Bosselut, A., and Choi, Y · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2021
Earlier work this paper cites.
Crass: A novel data set and benchmark to test counterfactual reasoning of large language models
Frohberg, J. and Binder, F · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2021
Earlier work this paper cites.
Counterfactual invariance to spurious correlations in text classification
Veitch, V., D’Amour, A., Yadlowsky, S., and Eisenstein, J · 2021
Earlier work this paper cites.
Promptsource: An integrated development environment and repository for natural language prompts
Bach, S. H., Sanh, V., Yong, Z.-X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., et al · 2022
Earlier work this paper cites.
ConvFinQA: Exploring the chain of numerical reasoning in conversational finance question answering
Chen, Z., Li, S., Smiley, C., Ma, Z., Shah, S., and Wang, W. Y · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Informativeness and invariance: Two perspectives on spurious correlations in natural language
Eisenstein, J · 2022
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
Geiger, A., Wu, Z., Lu, H., Rozner, J., Kreiss, E., Icard, T., Goodman, N., and Potts, C · 2022
Cited alongside, same era.
Unnatural instructions: Tuning language models with (almost) no human labor
Honovich, O., Scialom, T., Levy, O., and Schick, T · 2022
Cited alongside, same era.
Explanations from large language models make small reasoners better
Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks
Lin, B. Y., Fu, Y., Yang, K., Ammanabrolu, P., Brahman, F., Huang, S., Bhagavatula, C., Choi, Y., and Ren, X · 2023
Later among the works it cites.
Training socially aligned language models in simulated human society
Liu, R., Yang, R., Jia, C., Zhang, G., Zhou, D., Dai, A. M., Yang, D., and Vosoughi, S · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, S., Chen, J., Shen, Y., Chen, Z., Zhang, X., Li, Z., Wang, H., Qian, J., Peng, B., Mao, Y., et al · 2022
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Scao, T. L., Bari, M. S., Shen, S., Yong, Z.-X., Schoelkopf, H., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Cited alongside, same era.
awesome-chatgpt-prompts: A curated list of awesome chatgpt prompts
Akın, F. K · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
Limitations of language models in arithmetic and symbolic induction
Qian, J., Wang, H., Li, Z., Li, S., and Yan, X · 2023
Later among the works it cites.
Is chatgpt a general-purpose natural language processing task solver?, 2023
Qin, C., Zhang, A., Zhang, Z., Chen, J., Yasunaga, M., and Yang, D · 2023
Later among the works it cites.
Lamp: When large language models meet personalization
Salemi, A., Mysore, S., Bendersky, M., and Zamani, H · 2023
Later among the works it cites.
Role play with large language models
Shanahan, M., McDonell, K., and Reynolds, L · 2023
Later among the works it cites.
Cognitive architectures for language agents
Sumers, T., Yao, S., Narasimhan, K., and Griffiths, T. L · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al · 2023
Later among the works it cites.
Multi-party chat: Conversational agents in group settings with humans and models
Wei, J., Shuster, K., Szlam, A., Weston, J., Urbanek, J., and Komeili, M · 2023
Later among the works it cites.
Llm-powered autonomous agents
Weng, L · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Alpagasus: Training a better alpaca model with fewer data
Chen, L., Li, S., Yan, J., Wang, H., Gunaratna, K., Yadav, V., Tang, Z., Srinivasan, V., Zhou, T., Huang, H., and Jin, H · 2024
Closest in time.