Fetching the paper…
Reading the bibliography…
Developing a generalist agent is a longstanding objective in artificial intelligence.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Hypernetworks
Ha, D., Dai, A. M., and Le, Q. V. (2017) · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
Feature importance estimation with self-attention networks
Skrlj, B., Dzeroski, S., Lavrac, N., and Petkovic, M. (2020) · 2020
Earlier work this paper cites.
Feature importance ranking for deep learning
Wojtas, M. and Chen, K. (2020) · 2020
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N. (2021) · 2021
Earlier work this paper cites.
Chatgpt: A large-scale generative model for open-domain chat
OpenAI (2021) · 2021
Earlier work this paper cites.
AdapterFusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I. (2021) · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Earlier work this paper cites.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., et al. (2022) · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D., Zhu, Y., and Anandkumar, A. (2022) · 2022
Cited alongside, same era.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G. (2022) · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Jackson, T., Brown, N., Luu, L., Levine, S., Hausman, K., and Ichter, B. (2022) · 2022
Cited alongside, same era.
Hyperdecoders: Instance-specific decoders for multi-task NLP
Ivison, H. and Peters, M. (2022) · 2022
HiFi: High-information attention heads hold for parameter-efficient model adaptation
Gui, A. and Xiao, H. (2023) · 2023
Later among the works it cites.
Deep reinforcement learning with multitask episodic memory based on task-conditioned hypernetwork
Jin, Y., Wang, C., Xiang, L., Yang, Y., Fu, J., and He, Z. (2023) · 2023
Later among the works it cites.
Otter: A multi-modal model with in-context instruction tuning
Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., and Liu, Z. (2023) · 2023
Later among the works it cites.
Parameter-efficient fine-tuning without introducing new latency
Liao, B., Meng, Y., and Monz, C. (2023) · 2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., and Lee, Y. J. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-game decision transformers
Lee, K., Nachum, O., Yang, M., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., and Mordatch, I. (2022) · 2022
Cited alongside, same era.
Zero-shot reward specification via grounded natural language
Mahmoudieh, P., Pathak, D., and Darrell, T. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. (2022) · 2022
Cited alongside, same era.
A generalist agent
Reed, S. E., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N. (2022) · 2022
Cited alongside, same era.
Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction
Cai, S., Wang, Z., Ma, X., Liu, A., and Liang, Y. (2023a) · 2023
Cited alongside, same era.
Autoagents: A framework for automatic agent generation
Chen, G., Dong, S., Shu, Y., Zhang, G., Sesay, J., Karlsson, B. F., Fu, J., and Shi, Y. (2023) · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P. (2023) · 2023
Cited alongside, same era.
Later among the works it cites.
Conceptual reinforcement learning for language-conditioned tasks
Peng, S., Hu, X., Zhang, R., Guo, J., Yi, Q., Chen, R., Du, Z., Li, L., Guo, Q., and Chen, Y. (2023) · 2023
Later among the works it cites.
Generalization to new sequential decision making tasks with in-context learning
Raparthy, S. C., Hambro, E., Kirk, R., Henaff, M., and Raileanu, R. (2023) · 2023
Later among the works it cites.
Read and reap the rewards: Learning to play atari with the help of instruction manuals
Wu, Y., Fan, Y., Liang, P. P., Azaria, A., Li, Y., and Mitchell, T. M. (2023) · 2023
Later among the works it cites.
Hyper-decision transformer for efficient online policy adaptation
Xu, M., Lu, Y., Shen, Y., Zhang, S., Zhao, D., and Gan, C. (2023a) · 2023
Later among the works it cites.
Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning
Xu, Z., Shen, Y., and Huang, L. (2023b) · 2023
Later among the works it cites.
One network, many masks: Towards more parameter-efficient transfer learning
Zeng, G., Zhang, P., and Lu, W. (2023) · 2023
Later among the works it cites.
Prototype-based HyperAdapter for sample-efficient multi-task tuning
Zhao, H., Fu, J., and He, Z. (2023) · 2023
Later among the works it cites.