Fetching the paper…
Reading the bibliography…
We study the empirical scaling laws of a family of encoder-decoder autoregressive transformer models on the task of joint motion forecasting and planning in the autonomous driving domain.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
Spatio-temporal graph transformer networks for pedestrian trajectory prediction, 2020
C. Yu, X. Ma, J. Ren, H. Zhao, and S. Yi · 2005
Earlier work this paper cites.
Statistical Methods in Experimental Physics
F. James · 2006
Earlier work this paper cites.
Power-law distributions in empirical data
A. Clauset, C. R. Shalizi, and M. E. J. Newman · 2009
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Multi-head attention for multi-modal joint vehicle motion forecasting
J. Mercat, T. Gilles, N. El Zoghby, G. Sandou, D. Beauvois, and G. P. Gil · 2020
Earlier work this paper cites.
Mp3: A unified model to map, perceive, predict and plan
S. Casas, A. Sadat, and R. Urtasun · 2021
Earlier work this paper cites.
Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset
S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, and D. Anguelov · 2021
Earlier work this paper cites.
Scaling laws for neural machine translation
B. Ghorbani, O. Firat, M. Freitag, A. Bapna, M. Krikun, X. Garcia, C. Chelba, and C. Cherry · 2021
Earlier work this paper cites.
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Earlier work this paper cites.
Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction
B. Varadarajan, A. S. Hefny, A. Srivastava, K. S. Refaat, N. Nayakanti, A. Cornman, K. Chen, B. Douillard, C. P. Lam, D. Anguelov, and B. Sapp · 2021
Earlier work this paper cites.
Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting
Y. Yuan, X. Weng, Y. Ou, and K. M. Kitani · 2021
Earlier work this paper cites.
Latent variable sequential set transformers for joint multi-agent motion prediction, 2022
R. Girgis, F. Golemo, F. Codevilla, M. Weiss, J. A. D’Souza, S. E. Kahou, F. Heide, and C. Pal · 2022
Earlier work this paper cites.
Scaling laws and interpretability of learning from repeated data, 2022
D. Hernandez, T. Brown, T. Conerly, N. DasSarma, D. Drain, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, T. Henighan, T. Hume, S. Johnston, B. Mann, C. Olah, C. Olsson, D. Amodei, N. Joseph, J. Kaplan, and S. McCandlish · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Wayformer: Motion forecasting via simple & efficient attention networks
N. Nayakanti, R. Al-Rfou, A. Zhou, K. Goel, K. S. Refaat, and B. Sapp · 2022
Cited alongside, same era.
Scene transformer: A unified architecture for predicting multiple agent trajectories, 2022
J. Ngiam, B. Caine, V. Vasudevan, Z. Zhang, H.-T. L. Chiang, J. Ling, R. Roelofs, A. Bewley, C. Liu, A. Venugopal, D. Weiss, B. Sapp, Z. Chen, and J. Shlens · 2022
Cited alongside, same era.
Go smol or go home, 2023
H. De Vries · 2023
Cited alongside, same era.
Meshtron: High-fidelity, artist-like 3d mesh generation at scale, 2024
Z. Hao, D. W. Romero, T.-Y. Lin, and M.-Y. Liu · 2024
Later among the works it cites.
Drivegpt: Scaling autoregressive behavior models for driving
X. Huang, E. M. Wolff, P. Vernaza, T. Phan-Minh, H. Chen, D. S. Hayden, M. Edmonds, B. Pierce, X. Chen, P. E. Jacob, et al · 2024
Later among the works it cites.
X. Jia, S. Shi, Z. Chen, L. Jiang, W. Liao, T. He, and J. Yan · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model, 2024
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Cited alongside, same era.
Matformer: Nested transformer for elastic inference
S. Kudugunta, A. Kusupati, T. Dettmers, K. Chen, I. Dhillon, Y. Tsvetkov, H. Hajishirzi, S. Kakade, A. Farhadi, P. Jain, et al · 2023
Cited alongside, same era.
Scaling data-constrained language models, 2023
N. Muennighoff, A. M. Rush, B. Barak, T. L. Scao, A. Piktus, N. Tazi, S. Pyysalo, T. Wolf, and C. Raffel · 2023
Cited alongside, same era.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
N. Sardana and J. Frankle · 2023
Cited alongside, same era.
Motionlm: Multi-agent motion forecasting as language modeling
A. Seff, B. Cera, D. Chen, M. Ng, A. Zhou, N. Nayakanti, K. S. Refaat, R. Al-Rfou, and B. Sapp · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation, 2023
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2023
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling, 2024
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation, 2024
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu · 2024
Cited alongside, same era.
Is ego status all you need for open-loop end-to-end autonomous driving?
Z. Li, Z. Yu, S. Lan, J. Li, J. Kautz, T. Lu, and J. M. Alvarez · 2024
Later among the works it cites.
S. Shi, L. Jiang, D. Dai, and B. Schiele · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation, 2024
D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, Z. Zhao, X. Wu, and H. Meng · 2024
Later among the works it cites.
Latent action pretraining from videos, 2024
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo · 2024
Later among the works it cites.
Y. Zheng, Z. Xia, Q. Zhang, T. Zhang, B. Lu, X. Huo, C. Han, Y. Li, M. Yu, B. Jin, et al · 2024
Later among the works it cites.
Behaviorgpt: Smart agent simulation for autonomous driving with next-patch prediction, 2024
Z. Zhou, H. Hu, X. Chen, J. Wang, N. Guan, K. Wu, Y.-H. Li, Y.-K. Huang, and C. J. Xue · 2024
Later among the works it cites.
Closing the loop: Motion prediction models beyond open-loop benchmarks
M.-K. Bouzidi, C. Schlauch, N. Scheuerer, Y. Yao, N. Klein, D. Göhring, and J. Reichardt · 2025
Closest in time.
Hydra-next: Robust closed-loop driving with open-loop training
Z. Li, S. Wang, S. Lan, Z. Yu, Z. Wu, and J. M. Alvarez · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models, 2025
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
At nuro, we conduct an ai-first approach by using ml everywhere, 2025
B. Yao, A. Ganesh, Z. Li, and A. Petiushko · 2025
Closest in time.