Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced in various fields and intelligent agent applications.
Brown T, Mann B, Ryder N, et al (2020) Language models are few-shot learners. Advances in neural information processing systems 33:1877–1901
1901
Earlier work this paper cites.
Dewey J (1938) Experience and education: Kappa delta pi. International Honor Society in Education
1938
Earlier work this paper cites.
Searle JR (1986) Minds, brains and science. Harvard university press
1986
Earlier work this paper cites.
Holland JH (1992) Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence. MIT press
1992
Earlier work this paper cites.
Bäck T, Schwefel HP (1993) An overview of evolutionary algorithms for parameter optimization. Evolutionary computation 1(1):1–23
1993
Earlier work this paper cites.
Dennett DC (1993) Consciousness explained. Penguin uk
1993
Earlier work this paper cites.
Chalmers DJ (1997) The conscious mind: In search of a fundamental theory. Oxford Paperbacks
1997
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, et al (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp 311–318
2002
Earlier work this paper cites.
Lin CY (2004) Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out, pp 74–81
2004
Earlier work this paper cites.
Gupta AK, Smith KG, Shalley CE (2006) The interplay between exploration and exploitation. Academy of management journal 49(4):693–706
2006
Earlier work this paper cites.
Boud D, Keogh R, Walker D (2013) Reflection: Turning experience into learning. Routledge
2013
Earlier work this paper cites.
Hinton G, Vinyals O, Dean J (2015) Distilling the knowledge in a neural network. arXiv preprint arXiv:150302531
2015
Earlier work this paper cites.
Silver D, Huang A, Maddison CJ, et al (2016) Mastering the game of go with deep neural networks and tree search. nature 529(7587):484–489
2016
Earlier work this paper cites.
Kirkpatrick J, Pascanu R, Rabinowitz N, et al (2017) Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences 114(13):3521–3526
2017
Earlier work this paper cites.
Schön DA (2017) The reflective practitioner: How professionals think in action. Routledge
2017
Earlier work this paper cites.
Silver D, Hubert T, Schrittwieser J, et al (2017) Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:171201815
2017
Earlier work this paper cites.
Devlin J, Chang MW, Lee K, et al (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:181004805
2018
Earlier work this paper cites.
Hare J (2019) Dealing with sparse rewards in reinforcement learning. arXiv preprint arXiv:191009281
2019
Earlier work this paper cites.
Raffel C, Shazeer N, Roberts A, et al (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21(140):1–67
2020
Earlier work this paper cites.
Liu Y, Sun Y, Xue B, et al (2021) A survey on evolutionary neural architecture search. IEEE transactions on neural networks and learning systems 34(2):550–570
2021
Earlier work this paper cites.
Bai Y, Kadavath S, Kundu S, et al (2022) Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:221208073
2022
Earlier work this paper cites.
Chung HW, Hou L, Longpre S, et al (2022) Scaling instruction-finetuned language models. arXiv preprint arXiv:221011416
2022
Earlier work this paper cites.
Hu EJ, yelong shen, Wallis P, et al (2022) LoRA: Low-rank adaptation of large language models. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=nZeVKeeFYf9
2022
Earlier work this paper cites.
Ouyang L, Wu J, Jiang X, et al (2022) Training language models to follow instructions with human feedback. Advances in neural information processing systems
2022
Earlier work this paper cites.
Perez E, Huang S, Song F, et al (2022) Red teaming language models with language models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp 3419–3448
2022
Earlier work this paper cites.
Saunders W, Yeh C, Wu J, et al (2022) Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:220605802
2022
Earlier work this paper cites.
Wei J, Wang X, Schuurmans D, et al (2022) Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35:24824–24837
2022
Earlier work this paper cites.
Wortsman M, Ilharco G, Gadre SY, et al (2022) Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In: Proceedings of the 39th International Conference on Machine Learning, pp 23965–23998, URL https://proceedings.mlr.press/v162/wortsman22a.html
2022
Earlier work this paper cites.
Yao S, Zhao J, Yu D, et al (2022) React: Synergizing reasoning and acting in language models. In: The Eleventh International Conference on Learning Representations
2022
Earlier work this paper cites.
Zelikman E, Mu J, Goodman ND, et al (2022) Star: Self-taught reasoner bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems (NeurIPS)
2022
Earlier work this paper cites.
Achiam J, Adler S, Agarwal S, et al (2023) Gpt-4 technical report. arXiv preprint arXiv:230308774
2023
Earlier work this paper cites.
Aksitov R, Miryoosefi S, Li Z, et al (2023) Rest meets react: Self-improvement for multi-step reasoning llm agent. arXiv preprint arXiv:231210003
2023
Earlier work this paper cites.
Bai J, Bai S, Chu Y, et al (2023) Qwen technical report. arXiv preprint arXiv:230916609
2023
Earlier work this paper cites.
Bills S, Cammarata N, Mossing D, et al (2023) Language models can explain neurons in language models. URL https://openaipublic blob core windows net/neuron-explainer/paper/index html(Date accessed: 1405 2023)
2023
Earlier work this paper cites.
Bousmalis K, Vezzani G, Rao D, et al (2023) Robocat: A self-improving generalist agent for robotic manipulation. Transactions on Machine Learning Research
2023
Earlier work this paper cites.
Burns C, Izmailov P, Kirchner JH, et al (2023) Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:231209390
2023
Earlier work this paper cites.
Collins KM, Jiang AQ, Frieder S, et al (2023) Evaluating language models for mathematics through interactions. arXiv preprint arXiv:230601694
2023
Earlier work this paper cites.
Cui W, Wang Q (2023) Ada-instruct: Adapting instruction generators for complex reasoning. arXiv preprint arXiv:231004484
2023
Earlier work this paper cites.
Ding N, Chen Y, Xu B, et al (2023) Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:230514233
2023
Cited alongside, same era.
Eloundou T, Manning S, Mishkin P, et al (2023) Gpts are gpts: An early look at the labor market impact potential of large language models. arXiv preprint arXiv:230310130
2023
Cited alongside, same era.
Fernando C, Banarse D, Michalewski H, et al (2023) Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:230916797
2023
Cited alongside, same era.
Ganguli D, Askell A, Schiefer N, et al (2023) The capacity for moral self-correction in large language models. arXiv preprint arXiv:230207459
2023
Cited alongside, same era.
Gero Z, Singh C, Cheng H, et al (2023) Self-verification improves few-shot clinical information extraction. In: ICML 3rd Workshop on Interpretable Machine Learning in Healthcare (IMLH), URL https://openreview.net/forum?id=SBbJICrglS
Besta M, Blach N, Kubicek A, et al (2024) Graph of thoughts: Solving elaborate problems with large language models. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp 17682–17690
2024
Closest in time.
Chan CM, Chen W, Su Y, et al (2024) Chateval: Towards better LLM-based evaluators through multi-agent debate. In: The Twelfth International Conference on Learning Representations
2024
Closest in time.
Chen W, Li B (2024) Grath: Gradual self-truthifying for large language models. arXiv preprint arXiv:240112292
2024
Closest in time.
Chen Z, Deng Y, Yuan H, et al (2024) Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:240101335
2024
Closest in time.
Dettmers T, Pagnoni A, Holtzman A, et al (2024) Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems 36
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Gulcehre C, Paine TL, Srinivasan S, et al (2023) Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:230808998
2023
Cited alongside, same era.
Guo Y, Shang G, Vazirgiannis M, et al (2023) The curious decline of linguistic diversity: Training language models on synthetic text. arXiv preprint arXiv:231109807
2023
Cited alongside, same era.
Honovich O, Scialom T, Levy O, et al (2023) Unnatural instructions: Tuning language models with (almost) no human labor. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 14409–14428
2023
Cited alongside, same era.
Huang J, Gu SS, Hou L, et al (2023b) Large language models can self-improve. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, URL https://aclanthology.org/2023.emnlp-main.67
2023
Cited alongside, same era.
Ilharco G, Ribeiro MT, Wortsman M, et al (2023) Editing models with task arithmetic. In: The Eleventh International Conference on Learning Representations
2023
Cited alongside, same era.
Jiang S, Wang Y, Wang Y (2023) Selfevolve: A code evolution framework via large language models. arXiv preprint arXiv:230602907
2023
Cited alongside, same era.
Li M, Zhang Y, Li Z, et al (2023) From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuning. arXiv preprint arXiv:230812032
2023
Cited alongside, same era.
2024
Closest in time.
Ding N, Chen Y, Cui G, et al (2024) Mastering text, code and math simultaneously via fusing highly specialized language models. arXiv preprint arXiv:240308281
2024
Closest in time.
Dubois Y, Li CX, Taori R, et al (2024) Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Fu S, Zhang S, Wang Y, et al (2024) Towards theoretical understandings of self-consuming generative models. arXiv preprint arXiv:240211778
2024
Closest in time.
Ge Y, Hua W, Mei K, et al (2024) Openagi: When llm meets domain experts. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Gou Z, Shao Z, Gong Y, et al (2024) CRITIC: Large language models can self-correct with tool-interactive critiquing. In: The Twelfth International Conference on Learning Representations
2024
Closest in time.
Hosseini A, Yuan X, Malkin N, et al (2024) V-star: Training verifiers for self-taught reasoners. arXiv preprint arXiv:240206457
2024
Closest in time.
Kim G, Baldi P, McAleer S (2024) Language models can solve computer tasks. Advances in Neural Information Processing Systems
2024
Closest in time.
Koa KJ, Ma Y, Ng R, et al (2024) Learning to generate explainable stock predictions using self-reflective large language models. arXiv preprint arXiv:240203659
2024
Closest in time.
Lee N, Wattanawong T, Kim S, et al (2024) Llm2llm: Boosting llms with novel iterative data enhancement. arXiv preprint arXiv:240315042
2024
Closest in time.
Leike J, Sutskever I (2023) Introducing superalignment. URL https://openai.com/blog/introducing-superalignment , accessed: 2024-04-01
2024
Closest in time.
Lin Y, Lin H, Xiong W, et al (2024) Mitigating the alignment tax of rlhf. In: arXiv, 2309.06256
2024
Closest in time.
Luo Z, Xu C, Zhao P, et al (2024) Wizardcoder: Empowering code large language models with evol-instruct. In: The Twelfth International Conference on Learning Representations, URL https://openreview.net/forum?id=UnUwSIgK5W
2024
Closest in time.
Madaan A, Tandon N, Gupta P, et al (2024) Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems
2024
Closest in time.
Miret S, Krishnan N (2024) Are llms ready for real-world materials discovery? arXiv preprint arXiv:240205200
2024
Closest in time.
Noukhovitch M, Lavoie S, Strub F, et al (2024) Language model alignment with elastic reset. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Pang JC, Wang P, Li K, et al (2024) Language model self-improvement by reinforcement learning contemplation. In: The Twelfth International Conference on Learning Representations, URL https://openreview.net/forum?id=38E4yUbrgr
2024
Closest in time.
Qian C, Liang S, Qin Y, et al (2024) Investigate-consolidate-exploit: A general strategy for inter-task agent self-evolution. arXiv preprint arXiv:240113996
2024
Closest in time.
Qiao S, Zhang N, Fang R, et al (2024) Autoact: Automatic agent learning from scratch via self-planning. arXiv preprint arXiv:240105268
2024
Closest in time.
Ramé A, Vieillard N, Hussenot L, et al (2024) Warm: On the benefits of weight averaged reward models. arXiv preprint arXiv:240112187
2024
Closest in time.
Schoenegger P, Park PS, Karger E, et al (2024) Ai-augmented predictions: Llm assistants improve human forecasting accuracy. arXiv preprint arXiv:240207862
2024
Closest in time.
Shinn N, Cassano F, Gopinath A, et al (2024) Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Singh A, Co-Reyes JD, Agarwal R, et al (2024) Beyond human data: Scaling self-training for problem-solving with language models. Transactions on Machine Learning Research URL https://openreview.net/forum?id=lNAyUngGFK , expert Certification
2024
Closest in time.
Song Y, Yin D, Yue X, et al (2024) Trial and error: Exploration-based trajectory optimization for llm agents. arXiv preprint arXiv:240302502
2024
Closest in time.
Taubenfeld A, Dover Y, Reichart R, et al (2024) Systematic biases in llm simulations of debates. arXiv preprint arXiv:240204049
2024
Closest in time.
Tong Y, Li D, Wang S, et al (2024) Can llms learn from previous mistakes? investigating llms’ errors to boost for reasoning. arXiv preprint arXiv:240320046
2024
Closest in time.
Tu T, Palepu A, Schaekermann M, et al (2024) Towards conversational diagnostic ai. arXiv preprint arXiv:240105654
2024
Closest in time.
Ulmer D, Mansimov E, Lin K, et al (2024) Bootstrapping llm-based task-oriented dialogue agents via self-talk. arXiv preprint arXiv:240105033
2024
Closest in time.
Wan F, Huang X, Cai D, et al (2024) Knowledge fusion of large language models. In: The Twelfth International Conference on Learning Representations, URL https://openreview.net/forum?id=jiDsk12qcz
2024
Closest in time.
Wu T, Luo L, Li YF, et al (2024) Continual learning for large language models: A survey. arXiv preprint arXiv:240201364
2024
Closest in time.
Yao S, Yu D, Zhao J, et al (2024) Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Yuan W, Pang RY, Cho K, et al (2024) Self-rewarding language models. arXiv preprint arXiv:240110020
2024
Closest in time.
Zhu Y, Qiao S, Ou Y, et al (2024) Knowagent: Knowledge-augmented planning for llm-based agents. arXiv preprint arXiv:240303101
2024
Closest in time.