Fetching the paper…
Reading the bibliography…
Natural language processing and 2D vision models have attained remarkable proficiency on many tasks primarily by escalating the scale of training data.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
D3: Data-driven documents
M. Bostock, V. Ogievetsky, and J. Heer · 2011
Earlier work this paper cites.
Parsing ikea objects: Fine pose estimation
J. J. Lim, H. Pirsiavash, and A. Torralba · 2013
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Large-scale data for multiple-view stereopsis
H. Aanæs, R. R. Jensen, G. Vogiatzis, E. Tola, and A. B. Dahl · 2016
Earlier work this paper cites.
3d-r2n2: A unified approach for single and multi-view 3d object reconstruction
C. B. Choy, D. Xu, J. Gwak, K. Chen, and S. Savarese · 2016
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Neural 3d mesh renderer
H. Kato, Y. Ushiku, and T. Harada · 2018
Earlier work this paper cites.
Photoshape: Photorealistic materials for large-scale shape collections
K. Park, K. Rematas, A. Farhadi, and S. M. Seitz · 2018
Earlier work this paper cites.
Pixel2mesh: Generating 3d mesh models from single rgb images
N. Wang, Y. Zhang, Z. Li, Y. Fu, W. Liu, and Y.-G. Jiang · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Learning to predict 3d objects with an interpolation-based differentiable renderer
W. Chen, H. Ling, J. Gao, E. Smith, J. Lehtinen, A. Jacobson, and S. Fidler · 2019
Earlier work this paper cites.
PyTorch Lightning, Mar. 2019
W. Falcon and The PyTorch Lightning team · 2019
Earlier work this paper cites.
Mesh r-cnn
G. Gkioxari, J. Malik, and J. Johnson · 2019
Earlier work this paper cites.
Occupancy networks: Learning 3d reconstruction in function space
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
Array programming with NumPy
C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. F. del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Sheppard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2020
Cited alongside, same era.
Egad! an evolved grasping analysis dataset for diversity and reproducibility in robotic manipulation
D. Morrison, P. Corke, and J. Leitner · 2020
Shadows shed light on 3d objects
R. Liu, S. Menon, C. Mao, D. Park, S. Stent, and C. Vondrick · 2022
Later among the works it cites.
Unified-io: A unified model for vision, language, and multi-modal tasks
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi · 2022
Later among the works it cites.
Point-e: A system for generating 3d point clouds from complex prompts
A. Nichol, H. Jun, P. Dhariwal, P. Mishkin, and M. Chen · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
pandas-dev/pandas: Pandas, Feb. 2020
T. pandas development team · 2020
Cited alongside, same era.
Accelerating 3d deep learning with pytorch3d
N. Ravi, J. Reizenstein, D. Novotny, T. Gordon, W.-Y. Lo, J. Johnson, and G. Gkioxari · 2020
Cited alongside, same era.
3d-future: 3d furniture shape with texture
H. Fu, R. Jia, L. Gao, M. Gong, B. Zhao, S. Maybank, and D. Tao · 2021
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford · 2021
Cited alongside, same era.
Putting nerf on a diet: Semantically consistent few-shot view synthesis
A. Jain, M. Tancik, and P. Abbeel · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Ibrnet: Learning multi-view image-based rendering
Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser · 2021
Cited alongside, same era.
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Later among the works it cites.
Stable-dreamfusion: Text-to-3d with stable-diffusion, 2022
J. Tang · 2022
Later among the works it cites.
Blender - a 3d modelling and rendering package
Blender Online Community · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2023
Closest in time.
Shap-e: Generating conditional 3d implicit functions
H. Jun and A. Nichol · 2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Closest in time.
Magic3d: High-resolution text-to-3d content creation
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin · 2023
Closest in time.
Humans as light bulbs: 3d human reconstruction from thermal reflection
R. Liu and C. Vondrick · 2023
Closest in time.
Zero-1-to-3: Zero-shot one image to 3d object, 2023
R. Liu, R. Wu, B. V. Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation
H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich · 2023
Closest in time.
Lima: Less is more for alignment
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, et al · 2023
Closest in time.
Thingi10k: A dataset of 10,000 3d-printing models
Q. Zhou and A. Jacobson · 2048
Closest in time.