Fetching the paper…
Reading the bibliography…
Self-play has powered breakthroughs in two-player and multi-player games.
A note on two problems in connexion with graphs
Dijkstra, E. W · 1959
Earlier work this paper cites.
Congested traffic states in empirical observations and microscopic simulations
Treiber, M., Hennecke, A., and Helbing, D · 2000
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
General lane-changing model mobil for car-following models
Kesting, A., Treiber, M., and Helbing, D · 2007
Earlier work this paper cites.
Pre-crash scenario typology for crash avoidance research
Najm, W., Smith, J. D., and Yanagisawa, M · 2007
Earlier work this paper cites.
Enhanced intelligent driver model to access the impact of driving strategies on traffic capacity
Kesting, A., Treiber, M., and Helbing, D · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Earlier work this paper cites.
Learning to act by predicting the future
Dosovitskiy, A. and Koltun, V · 2017
Earlier work this paper cites.
CARLA: An open urban driving simulator
Dosovitskiy, A., Ros, G., Codevilla, F., López, A. M., and Koltun, V · 2017
Earlier work this paper cites.
Population based training of neural networks, 2017
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T. P., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Póczos, B., Salakhutdinov, R., and Smola, A. J · 2017
Earlier work this paper cites.
Driving policy transfer via modularity and abstraction
Müller, M., Dosovitskiy, A., Ghanem, B., and Koltun, V · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Earlier work this paper cites.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T · 2019
Earlier work this paper cites.
The trajectron: Probabilistic multi-agent trajectory modeling with dynamic spatiotemporal graphs
Ivanovic, B. and Pavone, M · 2019
Earlier work this paper cites.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castañeda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., Sonnerat, N., Green, T., Deason, L., Leibo, J. Z., Silver, D., Hassabis, D., Kavukcuoglu, K., and Graepel, T · 2019
Earlier work this paper cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Earlier work this paper cites.
nuScenes: A multimodal dataset for autonomous driving
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O · 2020
Earlier work this paper cites.
Learning quadrupedal locomotion over challenging terrain
Lee, J., Hwangbo, J., Wellhausen, L., Koltun, V., and Hutter, M · 2020
Earlier work this paper cites.
DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D · 2020
Earlier work this paper cites.
MP3: A unified model to map, perceive, predict and plan
Casas, S., Sadat, A., and Urtasun, R · 2021
Cited alongside, same era.
Large scale interactive motion forecasting for autonomous driving: The Waymo open motion dataset
Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., Yang, Z., Chouard, A., Sun, P., Ngiam, J., Vasudevan, V., McCauley, A., Shlens, J., and Anguelov, D · 2021
Cited alongside, same era.
Brax - A differentiable physics engine for large scale rigid body simulation
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O · 2021
Cited alongside, same era.
Reimagining an autonomous vehicle, 2021
Hawke, J., E, H., Badrinarayanan, V., and Kendall, A · 2021
Cited alongside, same era.
Autonomy 2.0: Why is self-driving always 5 years away?, 2021
Jain, A., Pero, L. D., Grimmett, H., and Ondruska, P · 2021
Dense reinforcement learning for safety validation of autonomous vehicles
Feng, S., Sun, H., Yan, X., Zhu, H., Zou, Z., Shen, S., and Liu, H. X · 2023
Later among the works it cites.
Establishing a crash rate benchmark using large-scale naturalistic human ridehail data
Flannagan, C., Leslie, A., Kiefer, R., Bogard, S., Chi-Johnston, G., Freeman, L., Huang, R., Walsh, D., and Anthony, J · 2023
Later among the works it cites.
Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research
Gulino, C., Fu, J., Luo, W., Tucker, G., Bronstein, E., Lu, Y., Harb, J., Pan, X., Wang, Y., Chen, X., Co-Reyes, J. D., Agarwal, R., Roelofs, R., Lu, Y., Montali, N., Mougin, P., Yang, Z., White, B., Faust, A., McAllister, R., Anguelov, D., and Sapp, B · 2023
Later among the works it cites.
SceneDM: Scene-level multi-agent trajectory generation with consistent diffusion models, 2023
Guo, Z., Gao, X., Zhou, J., Cai, X., and Shi, B · 2023
Later among the works it cites.
From prediction to planning with goal conditioned lane graph traversals
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Isaac gym: High performance GPU based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., and State, G · 2021
Cited alongside, same era.
Neural scene graphs for dynamic scenes
Ost, J., Mannan, F., Thuerey, N., Knodt, J., and Heide, F · 2021
Cited alongside, same era.
Megaverse: Simulating embodied agents at one million experiences per second
Petrenko, A., Wijmans, E., Shacklett, B., and Koltun, V · 2021
Cited alongside, same era.
Asymmetric self-play for automatic goal discovery in robotic manipulation
Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D’Sa, R., Petron, A., de Oliveira Pinto, H. P., Paino, A., Noh, H., Weng, L., Yuan, Q., Chu, C., and Zaremba, W · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
Learning to walk in minutes using massively parallel deep reinforcement learning
Rudin, N., Hoeller, D., Reist, P., and Hutter, M · 2021
Cited alongside, same era.
Urban driver: Learning to drive from real-world demonstrations using policy gradients
Scheel, O., Bergamini, L., Wolczyk, M., Osinski, B., and Ondruska, P · 2021
Cited alongside, same era.
Hallgarten, M., Stoll, M., and Zell, A · 2023
Later among the works it cites.
GAIA-1: A generative world model for autonomous driving, 2023
Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G · 2023
Later among the works it cites.
Hidden biases of end-to-end driving models
Jaeger, B., Chitta, K., and Geiger, A · 2023
Later among the works it cites.
Champion-level drone racing using deep reinforcement learning
Kaufmann, E., Bauersfeld, L., Loquercio, A., Müller, M., Koltun, V., and Scaramuzza, D · 2023
Later among the works it cites.
Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios
Lu, Y., Fu, J., Tucker, G., Pan, X., Bronstein, E., Roelofs, R., Sapp, B., White, B., Faust, A., Whiteson, S., Anguelov, D., and Levine, S · 2023
Later among the works it cites.
The Waymo open sim agents challenge
Montali, N., Lambert, J., Mougin, P., Kuefler, A., Rhinehart, N., Li, M., Gulino, C., Emrich, T., Yang, Z. Z., Whiteson, S., White, B., and Anguelov, D · 2023
Later among the works it cites.
Wayformer: Motion forecasting via simple & efficient attention networks
Nayakanti, N., Al-Rfou, R., Zhou, A., Goel, K., Refaat, K. S., and Sapp, B · 2023
Later among the works it cites.
PDexPBT: Scaling up dexterous manipulation for hand-arm systems with population based training
Petrenko, A., Allshire, A., State, G., Handa, A., and Makoviychuk, V · 2023
Later among the works it cites.
Trajeglish: Learning the language of driving scenarios
Philion, J., Peng, X. B., and Fidler, S · 2023
Later among the works it cites.
A simple yet effective method for simulating realistic multi-agent behaviors
Qian, C., Xiu, D., and Tian, M · 2023
Later among the works it cites.
An extensible, data-oriented architecture for high-performance, many-world simulation
Shacklett, B., Rosenzweig, L. G., Xie, Z., Sarkar, B., Szot, A., Wijmans, E., Koltun, V., Batra, D., and Fatahalian, K · 2023
Later among the works it cites.
Overview of motor vehicle traffic crashes in 2021
Stewart, T · 2023
Later among the works it cites.
Joint-multipath++ for simulation agents
Wang, W. and Zhen, H · 2023
Later among the works it cites.
Multiverse transformer: 1st place solution for Waymo open sim agents challenge 2023, 2023b
Wang, Y., Zhao, T., and Yi, F · 2023
Later among the works it cites.
Bits: Bi-level imitation for traffic simulation
Xu, D., Chen, Y., Ivanovic, B., and Pavone, M · 2023
Later among the works it cites.
UniSim: A neural closed-loop sensor simulator
Yang, Z., Chen, Y., Wang, J., Manivasagam, S., Ma, W.-C., Yang, A. J., and Urtasun, R · 2023
Later among the works it cites.
Learning realistic traffic agents in closed-loop
Zhang, C., Tu, J., Zhang, L., Wong, K., Suo, S., and Urtasun, R · 2023
Later among the works it cites.
PyTorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., et al · 2024
Later among the works it cites.
The Llama 3 herd of models, 2024
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Versatile scene-consistent traffic scenario generation as optimization with diffusion, 2024
Huang, Z., Zhang, Z., Vaidya, A., Chen, Y., Lv, C., and Fisac, J. F · 2024
Later among the works it cites.
Yang, B., Su, H., Gkanatsios, N., Ke, T.-W., Jain, A., Schneider, J., and Fragkiadaki, K · 2024
Later among the works it cites.