Fetching the paper…
Reading the bibliography…
We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language.
Prediction and entropy of printed english
Shannon, C. E · 1951
Earlier work this paper cites.
Development of wabot 1
Kato, I · 1973
Earlier work this paper cites.
Legged robots that balance
Raibert, M. H · 1986
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The development of honda humanoid robot
Hirai, K., Hirose, M., Haikawa, Y., and Takenaka, T · 1998
Earlier work this paper cites.
The 3d linear inverted pendulum mode: A simple modeling for a biped walking pattern generation
Kajita, S., Kanehiro, F., Kaneko, K., Yokoi, K., and Hirukawa, H · 2001
Earlier work this paper cites.
Petman: A humanoid robot for testing chemical protective clothing
Nelson, G., Saunders, A., Neville, N., Swilling, B., Bondaryk, J., Billings, D., Lee, C., Playter, R., and Raibert, M · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
The KIT motion-language dataset
Plappert, M., Mandery, C., and Asfour, T · 2016
Earlier work this paper cites.
Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling
Wu, J., Zhang, C., Xue, T., Freeman, B., and Tenenbaum, J · 2016
Earlier work this paper cites.
Talos: A new humanoid research platform targeted for industrial applications
Stasse, O., Flayols, T., Budhiraja, R., Giraud-Esclasse, K., Carpentier, J., Mirabel, J., Del Prete, A., Souères, P., Mansard, N., Lamiraux, F., et al · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Engel, J., Agrawal, K. K., Chen, S., Gulrajani, I., Donahue, C., and Roberts, A · 2019
Cited alongside, same era.
Robust feedback motion policy design using reinforcement learning on a 3d digit bipedal robot
Castillo, G. A., Weng, B., Zhang, W., and Hereid, A · 2021
Later among the works it cites.
The mit humanoid robot: Design, motion planning, and control for acrobatic behaviors
Chignoli, M., Kim, D., Stanger-Jones, E., and Kim, S · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AMASS: Archive of motion capture as surface shapes
Mahmood, N., Ghorbani, N., Troje, N. F., Pons-Moll, G., and Black, M. J · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Later among the works it cites.
Tracking people by predicting 3d appearance, location and pose
Rajasegaran, J., Pavlakos, G., Kanazawa, A., and Malik, J · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2022
Later among the works it cites.
Robocat: A self-improving foundation agent for robotic manipulation
Bousmalis, K., Vezzani, G., Rao, D., Devin, C., Lee, A. X., Bauza, M., Davchev, T., Zhou, Y., Gupta, A., Raju, A., et al · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Later among the works it cites.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Later among the works it cites.
Smpl: A skinned multi-person linear model
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., and Black, M. J · 2023
Later among the works it cites.