Fetching the paper…
Reading the bibliography…
We introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense.
Application of a particle-in-cell method to solid mechanics
Sulsky, D., Zhou, S.-J., and Schreyer, H. L · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
A history of the unity game engine
Haas, J. K · 2014
Earlier work this paper cites.
Unified particle physics for real-time applications
Macklin, M., Müller, M., Chentanez, N., and Kim, T.-Y · 2014
Earlier work this paper cites.
The affine particle-in-cell method
Jiang, C., Schroeder, C., Selle, A., Teran, J., and Stomakhin, A · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y., Song, Y., Yu, Y., Kim, Y., and Kim, G · 2017
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Lei, J., Yu, L., Bansal, M., and Berg, T. L · 2018
Earlier work this paper cites.
Learning particle dynamics for manipulating rigid bodies, deformable objects, and fluids
Li, Y., Wu, J., Tedrake, R., Tenenbaum, J. B., and Torralba, A · 2018
Earlier work this paper cites.
Intphys: A framework and benchmark for visual intuitive physics reasoning
Riochet, R., Castro, M. Y., Bernard, M., Lerer, A., Fergus, R., Izard, V., and Dupoux, E · 2018
Earlier work this paper cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J · 2018
Earlier work this paper cites.
Propagation networks for model-based control under partial observation
Li, Y., Wu, J., Zhu, J.-Y., Tenenbaum, J. B., Torralba, A., and Tedrake, R · 2019
Earlier work this paper cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.-F., and Wang, W. Y · 2019
Earlier work this paper cites.
Detectron2
Wu, Y., Kirillov, A., Massa, F., Lo, W.-Y., and Girshick, R · 2019
Cited alongside, same era.
Social-iq: A question answering benchmark for artificial social intelligence
Zadeh, A., Chan, M., Liang, P. P., Tong, E., and Morency, L.-P · 2019
Cited alongside, same era.
Craft: A benchmark for causal reasoning about forces and interactions
Ates, T., Atesoglu, M. S., Yigit, C., Kesen, I., Kobas, M., Erdem, E., Erdem, A., Goksun, T., and Yuret, D · 2020
Cited alongside, same era.
Cophy: Counterfactual learning of physical dynamics
Baradel, F., Neverova, N., Mille, J., Mori, G., and Wolf, C · 2020
Cited alongside, same era.
Threedworld: A platform for interactive multi-modal physical simulation
Gan, C., Schwartz, J., Alter, S., Mrowca, D., Schrimpf, M., Traer, J., De Freitas, J., Kubilius, J., Bhandwaldar, A., Haber, N., et al · 2020
Cited alongside, same era.
Star: A benchmark for situated reasoning in real-world videos
Wu, B., Yu, S., Chen, Z., Tenenbaum, J. B., and Gan, C · 2021
Later among the works it cites.
Comphy: Compositional physical reasoning of objects and events from videos
Chen, Z., Yi, K., Torralba, A., Tenenbaum, J., and Gan, C · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
CRIPP-VQA: Counterfactual reasoning about implicit physical properties via video question answering
Patel, M., Gokhale, T., Baral, C., and Yang, Y · 2022
Later among the works it cites.
Ulip: Learning unified representation of language, image and point cloud for 3d understanding
Xue, L., Gao, M., Xing, C., Martín-Martín, R., Wu, J., Xiong, C., Xu, R., Niebles, J. C., and Savarese, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Girdhar, R. and Ramanan, D · 2020
Cited alongside, same era.
Mental mechanics: How humans reason through a physical world
Kill, C. and Kim, O · 2020
Cited alongside, same era.
Hierarchical conditional relation networks for video question answering
Le, T. M., Le, V., Venkatesh, S., and Tran, T · 2020
Cited alongside, same era.
Disentangling physical dynamics from unknown factors for unsupervised video prediction
Le Guen, V. and Thome, N · 2020
Cited alongside, same era.
Esprit: explaining solutions to physical reasoning tasks
Rajani, N. F., Zhang, R., Tan, Y. C., Zheng, S., Weiss, J., Vyas, A., Gupta, A., Xiong, C., Socher, R., and Radev, D · 2020
Cited alongside, same era.
Sapien: A simulated part-based interactive environment
Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al · 2020
Cited alongside, same era.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2020
Cited alongside, same era.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., and Lu, J · 2022
Later among the works it cites.
Gpt-4 technical report
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
3d concept learning and reasoning from multi-view images
Hong, Y., Lin, C., Du, Y., Chen, Z., Tenenbaum, J. B., and Gan, C · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Tung, H.-Y., Ding, M., Chen, Z., Bear, D., Gan, C., Tenenbaum, J. B., Yamins, D. L., Fan, J. E., and Smith, K. A · 2023
Later among the works it cites.
ACQUIRED: A dataset for answering counterfactual questions in real-life videos
Wu, T.-L., Dou, Z.-Y., Hu, Q., Hou, Y., Chandra, N., Freedman, M., Weischedel, R., and Peng, N · 2023
Later among the works it cites.
Fluidlab: A differentiable environment for benchmarking complex fluid manipulation
Xian, Z., Zhu, B., Xu, Z., Tung, H.-Y., Torralba, A., Fragkiadaki, K., and Gan, C · 2023
Later among the works it cites.
Ulip-2: Towards scalable multimodal pre-training for 3d understanding, 2023
Xue, L., Yu, N., Zhang, S., Li, J., Martín-Martín, R., Wu, J., Xiong, C., Xu, R., Niebles, J. C., and Savarese, S · 2023
Later among the works it cites.
Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge
Wang, A., Wu, B., Chen, S., Chen, Z., Guan, H., Lee, W.-N., Li, L. E., Tenenbaum, J. B., and Gan, C · 2024
Closest in time.