Fetching the paper…
Reading the bibliography…
A Flow is a collection of component models ("Agents") which constructs the solution to a complex problem via iterative communication.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
Pairwise preference learning and ranking
Johannes Fürnkranz and Eyke Hüllermeier · 2003
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu · 2009
Earlier work this paper cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra · 2016
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
MuSiQue: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
A practical guide to multi-objective reinforcement learning and planning
Conor F Hayes, Roxana Rădulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M Zintgraf, Richard Dazeley, Fredrik Heintz, et al · 2022
Earlier work this paper cites.
Bridging rl theory and practice with the effective horizon
Cassidy Laidlaw, Stuart J Russell, and Anca Dragan · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
Tim Dettmers and Luke Zettlemoyer · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Cited alongside, same era.
https://twitter.com/karpathy/status/1748043513156272416?lang=en
Prompt engineering (or rather "Flow engineering") intensifies for code generation · 2024
Cited alongside, same era.
https://microsoft.github.io/autogen/
AutoGen: Enable Next-Gen Large Language Model Applications · 2024
Cited alongside, same era.
https://wensun.github.io/CS4789_data/Imitation_Learning_April_8_annotated.pdf
Wen Sun: Imitation Learning Lecture · 2024
Closest in time.
Large language models as commonsense knowledge for large-scale task planning
Zirui Zhao, Wee Sun Lee, and David Hsu · 2024
Closest in time.
https://github.com/huggingface/transformers
Transformers · 2024
Closest in time.
https://github.com/bremen79/parameterfree
Parameter-Free Optimizers for PyTorch · 2024
Closest in time.
https://github.com/huggingface/peft
State-of-the-art Parameter-Efficient Fine-Tuning (PEFT) methods · 2024
Closest in time.
https://ai.meta.com/blog/meta-llama-3/
Introducing Meta Llama 3: The most capable openly available LLM to date · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Cited alongside, same era.
Direct nash optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Cited alongside, same era.
Learn your reference model for real good alignment
Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov, Nikita Surnachev, Yaroslav Aksenov, Ian Maksimov, Nikita Balagansky, and Daniil Gavrilov · 2024
Cited alongside, same era.
https://github.com/rxlqn/awesome-llm-self-reflection
Awesome LLM Self-Reflection · 2024
Cited alongside, same era.
https://qwenlm.github.io/blog/qwen1.5/
Introducing Qwen1.5 · 2024
Closest in time.
https://azure.microsoft.com/en-us/blog/introducing-phi-3-redefining-whats-possible-with-slms/
Introducing Phi-3: Redefining what’s possible with SLMs · 2024
Closest in time.