Fetching the paper…
Reading the bibliography…
The evolution of machine learning has increasingly prioritized the development of powerful models and more scalable supervision signals.
The curious case of neural text degeneration, 2020
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 1904
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Induction of decision trees
Quinlan, J. R · 1986
Earlier work this paper cites.
Support vector machines
Hearst, M. A., Dumais, S. T., Osuna, E., Platt, J., and Scholkopf, B · 1998
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vapnik, V. N · 2000
Earlier work this paper cites.
Language models are few-shot learners, 2020b
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Finding shortest paths on real road networks: the case for a*
Zeng, W. and Church, R. L · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J · 2014
Earlier work this paper cites.
The lean theorem prover (system description)
de Moura, L. M., Kong, S., Avigad, J., van Doorn, F., and von Raumer, J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G · 2015
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Jordan, M. I. and Mitchell, T. M · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Hierarchical neural story generation, 2018
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
A capacity scaling law for artificial neural networks, 2018
Friedland, G. and Krell, M · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., and Zaremba, W · 2018
Earlier work this paper cites.
Diverse beam search: Decoding diverse solutions from neural sequence models, 2018
Vijayakumar, A. K., Cogswell, M., Selvaraju, R. R., Sun, Q., Lee, S., Crandall, D., and Batra, D · 2018
Earlier work this paper cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Krishna, K., Wieting, J., and Iyyer, M · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., and Vasserman, L · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Goal-conditioned reinforcement learning: Problems and solutions, 2022
Liu, M., Zhu, M., and Zhang, W · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning, 2022
Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Cited alongside, same era.
Compute trends across three eras of machine learning
Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., and Villalobos, P · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision, 2023
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Closest in time.
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., et al · 2024
Closest in time.
Towards scalable automated alignment of llms: A survey, 2024
Cao, B., Lu, K., Lu, X., Chen, J., Ren, M., Xiang, H., Liu, P., Lu, Y., He, B., Han, X., Sun, L., Lin, H., and Yu, B · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J · 2023
Cited alongside, same era.
Code alpaca: An instruction-following llama model for code generation
Chaudhary, S · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models, 2023
Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., and Weston, J · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Cited alongside, same era.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T · 2023
Cited alongside, same era.
Mods: Model-oriented data selection for instruction tuning
Du, Q., Zong, C., and Zhang, J · 2023
Cited alongside, same era.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Closest in time.
Critic: Large language models can self-correct with tool-interactive critiquing, 2024
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W · 2024
Closest in time.
Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting
Guan, X., Liu, Y., Lin, H., Lu, Y., He, B., Han, X., and Sun, L · 2024
Closest in time.
Hu, Z., Liu, C., Feng, X., Zhao, Y., Ng, S.-K., Luu, A. T., He, J., Koh, P. W., and Hooi, B · 2024
Closest in time.
Training language models to self-correct via reinforcement learning, 2024
Kumar, A., Zhuang, V., Agarwal, R., Su, Y., Co-Reyes, J. D., Singh, A., Baumli, K., Iqbal, S., Bishop, C., Roelofs, R., Zhang, L. M., McKinney, K., Shrivastava, D., Paduraru, C., Tucker, G., Precup, D., Behbahani, F., and Faust, A · 2024
Closest in time.
Mario: Math reasoning with code interpreter output – a reproducible pipeline, 2024
Liao, M., Luo, W., Li, C., Wu, J., and Fan, K · 2024
Closest in time.
Lean-star: Learning to interleave thinking and proving, 2024
Lin, H., Sun, Z., Yang, Y., and Welleck, S · 2024
Closest in time.
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning
Liu, W., Zeng, W., He, K., Jiang, Y., and He, J · 2024
Closest in time.
Transferable post-training via inverse value learning, 2024
Lu, X., Wen, X., Lu, Y., Yu, B., Lin, H., Yu, H., Sun, L., Han, X., and Li, Y · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Closest in time.
Rule based rewards for language model safety
Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J., and Weng, L · 2024
Closest in time.
Rex: Rapid exploration and exploitation for ai agents, 2024
Murthy, R., Heinecke, S., Niebles, J. C., Liu, Z., Xue, L., Yao, W., Feng, Y., Chen, Z., Gokul, A., Arpit, D., Xu, R., Mui, P., Wang, H., Xiong, C., and Savarese, S · 2024
Closest in time.
Balancing exploration and exploitation in llm using soft rllf for enhanced negation understanding
Nguyen, H.-T. and Satoh, K · 2024
Closest in time.
Not all contexts are equal: Teaching llms credibility-aware generation, 2024
Pan, R., Cao, B., Lin, H., Han, X., Zheng, J., Wang, S., Cai, X., and Sun, L · 2024
Closest in time.
Dmoerm: Recipes of mixture-of-experts for effective reward modeling
Quan, S · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Closest in time.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms
Tang, Y., Guo, D. Z., Zheng, Z., Calandriello, D., Cao, Y., Tarassov, E., Munos, R., Pires, B. Á., Valko, M., Cheng, Y., et al · 2024
Closest in time.
Chain-of-thought reasoning without prompting
Wang, X. and Zhou, D · 2024
Closest in time.
Codeultrafeedback: An llm-as-a-judge dataset for aligning large language models to coding preferences, 2024
Weyssow, M., Kamanda, A., and Sahraoui, H · 2024
Closest in time.
Wilson, D · 2024
Closest in time.
Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge
Wu, T., Yuan, W., Golovneva, O., Xu, J., Tian, Y., Jiao, J., Weston, J., and Sukhbaatar, S · 2024
Closest in time.
Aligning large language models via self-steering optimization, 2024
Xiang, H., Yu, B., Lin, H., Lu, K., Lu, Y., Han, X., Sun, L., Zhou, J., and Lin, J · 2024
Closest in time.
Monte carlo tree search boosts reasoning via iterative preference learning
Xie, Y., Goyal, A., Zheng, W., Kan, M.-Y., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Closest in time.
Beyond full fine-tuning: Harnessing the power of LoRA for multi-task instruction tuning
Xin, C., Lu, Y., Lin, H., Zhou, S., Zhu, H., Wang, W., Liu, Z., Han, X., and Sun, L · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Closest in time.
OVM, outcome-supervised value models for planning in mathematical reasoning
Yu, F., Gao, A., and Wang, B · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
Zelikman, E., Harik, G., Shao, Y., Jayasiri, V., Haber, N., and Goodman, N. D · 2024
Closest in time.