Fetching the paper…
Reading the bibliography…
Visual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others.
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam · 1905
Earlier work this paper cites.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
La Représentation de L’espace Chez L’enfant. The Child’s Conception of Space… Translated… by FJ Langdon & JL Lunzer. With Illustrations
Jean Piaget, Bärbel Inhelder, Frederick John Langdon, and J L Lunzer · 1956
Earlier work this paper cites.
RANDOMIZATION TESTS
E S Edgington · 1964
Earlier work this paper cites.
The visual perception of 3D shape
James T Todd · 2004
Earlier work this paper cites.
Do visual perspective tasks need theory of mind?
Markus Aichhorn, Josef Perner, Martin Kronbichler, Wolfgang Staffen, and Gunther Ladurner · 2006
Earlier work this paper cites.
Two kinds of visual perspective taking
Pascale Michelon and Jeffrey M Zacks · 2006
Earlier work this paper cites.
3D Shape: Its Unique Place in Visual Perception
Zygmunt Pizlo · 2010
Earlier work this paper cites.
Picturing perspectives: development of perspective-taking abilities in 4- to 8-year-olds
Andrea Frick, Wenke Möhring, and Nora S Newcombe · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Daniel L K Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Earlier work this paper cites.
Using goal-driven deep learning models to understand sensory cortex
Daniel L K Yamins and James J DiCarlo · 2016
Earlier work this paper cites.
Understanding gaps in research networks: using “spatial reasoning” as a window into the importance of networked educational research
Catherine D Bruce, Brent Davis, Nathalie Sinclair, Lynn McGarvey, David Hallowell, Michelle Drefs, Krista Francis, Zachary Hawes, Joan Moss, Joanne Mulligan, Yukari Okamoto, Walter Whiteley, and Geoff Woolcott · 2017
Earlier work this paper cites.
Dual path networks, 2017
Yunpeng Chen, Jianan Li, Huaxin Xiao, Xiaojie Jin, Shuicheng Yan, and Jiashi Feng · 2017
Earlier work this paper cites.
Superhuman accuracy on the SNEMI3D connectomics challenge
Kisuk Lee, Jonathan Zung, Peter Li, Viren Jain, and H Sebastian Seung · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
The neural correlates of visual perspective taking: a critical review
Henryk Bukowski · 2018
Earlier work this paper cites.
Densely connected convolutional networks, 2018
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2018
Earlier work this paper cites.
Unity: A general platform for intelligent agents
Arthur Juliani, Vincent-Pierre Berges, Ervin Teng, Andrew Cohen, Jonathan Harper, Chris Elion, Chris Goy, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange · 2018
Earlier work this paper cites.
Res2net: A new multi-scale backbone architecture
Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip Torr · 2019
Earlier work this paper cites.
Visual perspective taking in young and older adults
Andrew K Martin, Garon Perceval, Islay Davies, Peter Su, Jasmine Huang, and Marcus Meinzer · 2019
Earlier work this paper cites.
PyTorch: An imperative style, High-Performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Deep learning: The good, the bad, and the ugly
Thomas Serre · 2019
Earlier work this paper cites.
Single-path nas: Designing hardware-efficient convnets in less than 4 hours, 2019
Dimitrios Stamoulis, Ruizhou Ding, Di Wang, Dimitrios Lymberopoulos, Bodhi Priyantha, Jie Liu, and Diana Marculescu · 2019
Earlier work this paper cites.
High-resolution representations for labeling pixels and regions, 2019
Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang · 2019
Earlier work this paper cites.
Mixconv: Mixed depthwise convolutional kernels, 2019
Mingxing Tan and Quoc V. Le · 2019
Earlier work this paper cites.
Mnasnet: Platform-aware neural architecture search for mobile, 2019
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le · 2019
Earlier work this paper cites.
Pytorch image models
Ross Wightman · 2019
Earlier work this paper cites.
Deep layer aggregation, 2019
Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell · 2019
Cited alongside, same era.
Controversial stimuli: Pitting neural networks against each other as models of human cognition
Tal Golan, Prashant C Raju, and Nikolaus Kriegeskorte · 2020
Cited alongside, same era.
Disentangling neural mechanisms for perceptual grouping
Junkyung Kim*, Drew Linsley*, Kalpit Thakkar, and Thomas Serre · 2020
Cited alongside, same era.
Beyond the feedforward sweep: feedback computations in the visual cortex
Gabriel Kreiman and Thomas Serre · 2020
Cited alongside, same era.
Recurrent neural circuits for contour detection
Drew Linsley, Junkyung Kim, Alekh Ashok, and Thomas Serre · 2020
Cited alongside, same era.
Designing network design spaces, 2020
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
A convnet for the 2020s, 2022
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Edgenext: Efficiently amalgamated cnn-transformer architecture for mobile vision applications
Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muhammad Anwer, and Fahad Shahbaz Khan · 2022
Later among the works it cites.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
Sachin Mehta and Mohammad Rastegari · 2022
Later among the works it cites.
Problem Solving: Cognitive Mechanisms and Formal Models
Zygmunt Pizlo · 2022
Later among the works it cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
Mingxing Tan and Quoc V. Le · 2020
Cited alongside, same era.
Resnest: Split-attention networks, 2020
Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Haibin Lin, Zhi Zhang, Yue Sun, Tong He, Jonas Mueller, R. Manmatha, Mu Li, and Alexander Smola · 2020
Cited alongside, same era.
Deep ViT features as dense visual descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel · 2021
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Cited alongside, same era.
Pp-lcnet: A lightweight cpu convolutional neural network, 2021
Cheng Cui, Tingquan Gao, Shengyu Wei, Yuning Du, Ruoyu Guo, Shuilong Dong, Bin Lu, Ying Zhou, Xueying Lv, Qiwen Liu, Xiaoguang Hu, Dianhai Yu, and Yanjun Ma · 2021
Cited alongside, same era.
Coatnet: Marrying convolution and attention for all data sizes
Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan · 2021
Cited alongside, same era.
Rene Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2022
Later among the works it cites.
Patches are all you need?, 2022
Asher Trockman and J. Zico Kolter · 2022
Later among the works it cites.
Maxvit: Multi-axis vision transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li · 2022
Later among the works it cites.
Pvtv2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2022
Later among the works it cites.
Focal modulation networks, 2022
Jianwei Yang, Chunyuan Li, Xiyang Dai, and Jianfeng Gao · 2022
Later among the works it cites.
Volo: Vision outlooker for visual recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
GPT-4 technical report
Openai Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, S Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, B Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, L Fedus, Niko Felix, Sim’on Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, C Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Raphael Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, S Gray, Ryan Greene, Joshua Gross, S Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Lukasz Kaiser, Ali Kamali, I Kanitscheider, N Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Hendrik Kirchner, J Kiros, Matthew Knight, Daniel Kokotajlo, Lukasz Kondraciuk, A Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, J Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, A Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, S McKinney, C McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel P Mossing, Tong Mu, Mira Murati, O Murk, David M’ely, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Ouyang Long, Cullen O’Keefe, J Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alexandre Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Pondé de Oliveira Pinto, Michael Pokorny, Michelle Pokrass, Vitchyr H Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, M Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, T Sherbakov, Jessica Shieh, S Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin D Sokolowsky, Yang Song, Natalie Staudacher, F Such, Natalie Summers, I Sutskever, Jie Tang, N Tezak, Madeleine Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cer’on Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll L Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, C J Weinmann, Akila Welihinda, P Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph · 2023
Later among the works it cites.
StyleGAN knows normal, depth, albedo, and more
Anand Bhattad, Daniel McKee, Derek Hoiem, and David Forsyth · 2023
Later among the works it cites.
Beyond surface statistics: Scene representations in a latent diffusion model
Yida Chen, Fernanda Viégas, and Martin Wattenberg · 2023
Later among the works it cites.
Eva-02: A visual representation for neon genesis, 2023
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
3D gaussian splatting for real-time radiance field rendering
B Kerbl, Georgios Kopanas, Thomas Leimkuehler, and G Drettakis · 2023
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, and Others · 2023
Later among the works it cites.
Your diffusion model is secretly a zero-shot classifier
Alexander C Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak · 2023
Later among the works it cites.
Zero-1-to-3: Zero-shot one image to 3D object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, and Others · 2023
Later among the works it cites.
Unity gaussian splatting
Aras Pranckevicius · 2023
Later among the works it cites.
Zero-Shot metric depth with a Field-of-View conditioned diffusion model
Saurabh Saxena, Junhwa Hur, Charles Herrmann, Deqing Sun, and David J Fleet · 2023
Later among the works it cites.
Emergent correspondence from image diffusion
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, and Others · 2023
Later among the works it cites.
Claude, 2024
Anthropic · 2024
Closest in time.
Probing the 3D Awareness of Visual Foundation Models
Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas Guibas, Justin Johnson, and Varun Jampani · 2024
Closest in time.
Shape guides visual pretense
Peng Qian and Tomer D Ullman · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Closest in time.
ReCapture: Generative video camera controls for user-provided videos using masked video fine-tuning
David Junhao Zhang, Roni Paiss, Shiran Zada, Nikhil Karnad, David E Jacobs, Yael Pritch, Inbar Mosseri, Mike Zheng Shou, Neal Wadhwa, and Nataniel Ruiz · 2024
Closest in time.