Fetching the paper…
Reading the bibliography…
Steering the behavior of a strong model pre-trained on internet-scale data can be difficult due to the scarcity of competent supervisors.
Adaptive Mixtures of Local Experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. and Jacobs, R · 1993
Earlier work this paper cites.
A Framework for Behavioural Cloning
Bain, M. and Sammut, C · 1995
Earlier work this paper cites.
Bagging predictors
Breiman, L · 1996
Earlier work this paper cites.
Experiments with a new boosting algorithm
Freund, Y. and Schapire, R. E · 1996
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Schapire, R. E., Freund, Y., Bartlett, P., and Lee, W. S · 1998
Earlier work this paper cites.
Bayesian Model Averaging: A Tutorial
Hoeting, J. A., Madigan, D., Raftery, A. E., and Volinsky, C. T · 1999
Earlier work this paper cites.
Bayesian hierarchical mixtures of experts
Bishop, C. M. and Svenskn, M · 2002
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., and and · 2009
Earlier work this paper cites.
Twenty Years of Mixture of Experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D · 2012
Earlier work this paper cites.
Concrete Problems in AI Safety, arXiv:1606.06565, July 2016
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2016
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Earlier work this paper cites.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M · 2018
Earlier work this paper cites.
Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Cited alongside, same era.
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
Jiang, L., Zhou, Z., Leung, T., Li, L.-J., and Fei-Fei, L · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: A research direction, arXiv:1811.07871, November 2018
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Understanding and Utilizing Deep Neural Networks Trained with Noisy Labels
Chen, P., Liao, B. B., Chen, G., and Zhang, S · 2019
Cited alongside, same era.
Demski, A. and Garrabrant, S · 2019
Cited alongside, same era.
A Review of Sparse Expert Models in Deep Learning, arXiv:2209.01667, September 2022
Fedus, W., Dean, J., and Zoph, B · 2022
Later among the works it cites.
Ensemble deep learning: A review
Ganaie, M. A., Hu, M., Malik, A. K., Tanveer, M., and Suganthan, P. N · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
High-Resolution Image Synthesis With Latent Diffusion Models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Mixture-of-Experts with Expert Choice Routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Chen, Z., Le, Q. V., and Laudon, J · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Moment Matching for Multi-Source Domain Adaptation
Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Cited alongside, same era.
Revisiting Knowledge Distillation via Label Smoothing Regularization
Yuan, L., Tay, F. E., Li, G., Wang, T., and Feng, J · 2020
Cited alongside, same era.
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J · 2023
Later among the works it cites.
Statement on AI risk, 2023
CAIS · 2023
Later among the works it cites.
Evaluating Superhuman Models with Consistency Checks, arXiv:2306.09983, October 2023
Fluri, L., Paleka, D., and Tramèr, F · 2023
Later among the works it cites.
Noisy-label Learning with Sample Selection based on Noise Rate Estimate, arXiv:2305.19486, May 2023
Garg, A., Nguyen, C., Felix, R., Do, T.-T., and Carneiro, G · 2023
Later among the works it cites.
AI Alignment: A Comprehensive Survey, arXiv:2310.19852, November 2023
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., Zeng, F., Ng, K. Y., Dai, J., Pan, X., O’Gara, A., Lei, Y., Xu, H., Tse, B., Fu, J., McAleer, S., Yang, Y., Wang, Y., Zhu, S.-C., Guo, Y., and Gao, W · 2023
Later among the works it cites.
Introducing Superalignment
Leike, J. and Sutskever, I · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models, arXiv:2307.09288, July 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Video generation models as world simulators
Brooks, T., Peebles, B., Homes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C. W. Y., Wang, R., and Ramesh, A · 2024
Closest in time.
Mixtral of Experts, arXiv:2401.04088, January 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
DINOv2: Learning Robust Visual Features without Supervision, arXiv:2304.07193, February 2024
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2024
Closest in time.