Fetching the paper…
Reading the bibliography…
The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views.
Selective search for object recognition
Uijlings J. R., Van De Sande K. E., Gevers T., Smeulders A. W · 2013
Earlier work this paper cites.
Automatic editing of footage from multiple social cameras
Arev I., Park H. S., Sheikh Y., Hodgins J., Shamir A · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
Zitnick C. L., Dollár P · 2014
Earlier work this paper cites.
Panoptic studio: A massively multiview system for social motion capture
Joo H., Liu H., Tan L., Gui L., Nabbe B., Matthews I., Kanade T., Nobuhara S., Sheikh Y · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren S., He K., Girshick R., Sun J · 2015
Earlier work this paper cites.
Scale-aware pixelwise object proposal networks
Jie Z., Liang X., Feng J., Lu W. F., Tay E. H. F., Yan S · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks L., Wang O., Shechtman E., Sivic J., Darrell T., Russell B · 2017
Earlier work this paper cites.
Time slice video synthesis by robust video alignment
Cui Z., Wang O., Tan P., Wang J · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao J., Sun C., Yang Z., Nevatia R · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
Hu R., Rohrbach M., Andreas J., Darrell T., Saenko K · 2017
Earlier work this paper cites.
Computational video editing for dialogue-driven scenes
Leake M., Davis A., Truong A., Agrawala M · 2017
Earlier work this paper cites.
Spatio-temporal person retrieval via natural language queries
Yamaguchi M., Saito K., Ushiku Y., Harada T · 2017
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments
Gordon D., Kembhavi A., Rastegari M., Redmon J., Fox D., Farhadi A · 2018
Earlier work this paper cites.
Object referring in videos with language and human gaze
Vasudevan A. B., Dai D., Van Gool L · 2018
Earlier work this paper cites.
Weakly-supervised video object grounding from text by loss weighting and object interaction
Zhou L., Louis N., Corso J. J · 2018
Earlier work this paper cites.
Weakly-supervised spatio-temporally grounding natural sentence in video
Chen Z., Ma L., Luo W., Wong K.-Y. K · 2019
Earlier work this paper cites.
Compositing light field video using multiplane images
DuVall M., Flynn J., Broxton M., Debevec P · 2019
Earlier work this paper cites.
Immersive light field video with a layered mesh representation
Broxton M., Flynn J., Overbeck R., Erickson D., Hedman P., DuVall M., Dourgarian J., Busch J., Whalen M., Debevec P · 2020
Earlier work this paper cites.
ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
Chen D. Z., Chang A. X., Nießner M · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho J., Jain A., Abbeel P · 2020
Earlier work this paper cites.
Layered neural rendering for retiming people in video
Lu E., Cole F., Dekel T., Xie W., Zisserman A., Salesin D., Freeman W., Rubinstein M · 2020
Earlier work this paper cites.
Multi-view neural human rendering
Wu M., Wang Y., Hu Q., Yu J · 2020
Earlier work this paper cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhang Z., Zhao Z., Zhao Y., Wang Q., Liu H., Gao L · 2020
Earlier work this paper cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhang Z., Zhao Z., Zhao Y., Wang Q., Liu H., Gao L · 2020
Earlier work this paper cites.
Who’s waldo? linking people across text and images
Cui Y., Khandelwal A., Artzi Y., Snavely N., Averbuch-Elor H · 2021
Earlier work this paper cites.
Dynamic view synthesis from dynamic monocular video
Gao C., Saraf A., Kopf J., Huang J.-B · 2021
Earlier work this paper cites.
Layered neural atlases for consistent video editing
Kasten Y., Ofri D., Wang O., Dekel T · 2021
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall B., Srinivasan P. P., Tancik M., Barron J. T., Ramamoorthi R., Ng R · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol A., Dhariwal P., Ramesh A., Shyam P., Mishkin P., McGrew B., Sutskever I., Chen M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford A., Kim J. W., Hallacy C., Ramesh A., Goh G., Agarwal S., Sastry G., Askell A., Mishkin P., Clark J., et al · 2021
Cited alongside, same era.
Human-centric spatio-temporal video grounding with visual transformers
Tang Z., Liao Y., Liu S., Li G., Jin X., Jiang H., Yu Q., Xu D · 2021
Cited alongside, same era.
Lerf: Language embedded radiance fields
Kerr J., Kim C. M., Goldberg K., Kanazawa A., Tancik M · 2023
Later among the works it cites.
3d gaussian splatting for real-time radiance field rendering
Kerbl B., Kopanas G., Leimkühler T., Drettakis G · 2023
Later among the works it cites.
Kirillov A., Mintun E., Ravi N., Mao H., Rolland C., Gustafson L., Xiao T., Whitehead S., Berg A. C., Lo W.-Y., Dollár P., Girshick R · 2023
Later among the works it cites.
Spacetime gaussian feature splatting for real-time dynamic view synthesis
Li Z., Chen Z., Li Z., Xu Y · 2023
Later among the works it cites.
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis
Luiten J., Kopanas G., Leibe B., Ramanan D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scanqa: 3d question answering for spatial scene understanding
Azuma D., Miyanishi T., Kurita S., Kawanabe M · 2022
Cited alongside, same era.
Text2live: Text-driven layered image and video editing
Bar-Tal O., Ofri-Amar D., Fridman R., Kasten Y., Dekel T · 2022
Cited alongside, same era.
Simvqa: Exploring simulated environments for visual question answering
Cascante-Bonilla P., Wu H., Wang L., Feris R. S., Ordonez V · 2022
Cited alongside, same era.
Ham: Hierarchical attention model with high performance for 3d visual grounding
Chen J., Luo W., Wei X., Ma L., Zhang W · 2022
Cited alongside, same era.
Voxel-informed language grounding
Corona R., Zhu S., Klein D., Darrell T · 2022
Cited alongside, same era.
Multi-view transformer for 3d visual grounding
Huang S., Chen Y., Jia J., Wang L · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Hertz A., Mokady R., Tenenbaum J., Aberman K., Pritch Y., Cohen-Or D · 2022
Cited alongside, same era.
Later among the works it cites.
Video-p2p: Video editing with cross-attention control
Liu S., Zhang Y., Li W., Lin Z., Jia J · 2023
Later among the works it cites.
Localizing object-level shape variations with text-to-image diffusion models
Patashnik O., Garibi D., Azuri I., Averbuch-Elor H., Cohen-Or D · 2023
Later among the works it cites.
Langsplat: 3d language gaussian splatting
Qin M., Li W., Zhou J., Wang H., Pfister H · 2023
Later among the works it cites.
Language embedded 3d gaussians for open-vocabulary scene understanding
Shi J.-C., Wang M., Duan H.-B., Guan S.-H · 2023
Later among the works it cites.
Tracking everything everywhere all at once
Wang Q., Chang Y.-Y., Cai R., Li Z., Hariharan B., Holynski A., Snavely N · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu J. Z., Ge Y., Wang X., Lei S. W., Gu Y., Shi Y., Hsu W., Shan Y., Qie X., Shou M. Z · 2023
Later among the works it cites.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Wang Y., He Y., Li Y., Li K., Yu J., Ma X., Li X., Chen G., Chen X., Wang Y., et al · 2023
Later among the works it cites.
4d gaussian splatting for real-time dynamic scene rendering
Wu G., Yi T., Fang J., Xie L., Zhang X., Wei W., Liu W., Tian Q., Wang X · 2023
Later among the works it cites.
Gaussian grouping: Segment and edit anything in 3d scenes
Ye M., Danelljan M., Yu F., Ke L · 2023
Later among the works it cites.
Track anything: Segment anything meets videos, 2023
Yang J., Gao M., Li Z., Gao S., Wang F., Zheng F · 2023
Later among the works it cites.
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction
Yang Z., Gao X., Zhou W., Jiao S., Zhang Y., Jin X · 2023
Later among the works it cites.
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Zhou S., Chang H., Jiang S., Fan Z., Zhu Z., Xu D., Chari P., You S., Wang Z., Kadambi A · 2023
Later among the works it cites.
Halo-nerf: Learning geometry-guided semantics for exploring unconstrained photo collections
Dudai C., Alper M., Bezalel H., Hanocka R., Lang I., Averbuch-Elor H · 2024
Closest in time.
Context-guided spatio-temporal video grounding
Gu X., Fan H., Huang Y., Luo T., Zhang L · 2024
Closest in time.
Semantic anything in 3d gaussians
Hu X., Wang Y., Fan L., Fan J., Peng J., Lei Z., Li Q., Zhang Z · 2024
Closest in time.
Garfield: Group anything with radiance fields
Kim C. M., Wu M., Kerr J., Goldberg K., Tancik M., Kanazawa A · 2024
Closest in time.
Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle
Lin Y., Dai Z., Zhu S., Yao Y · 2024
Closest in time.
3d geometry-aware deformable gaussian splatting for dynamic view synthesis
Lu Z., Guo X., Hui L., Chen T., Yang M., Tang X., Zhu F., Dai Y · 2024
Closest in time.
Dgd: Dynamic 3d gaussians distillation
Labe I., Issachar N., Lang I., Benaim S · 2024
Closest in time.
3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos
Sun J., Jiao H., Li G., Zhang Z., Zhao L., Xing W · 2024
Closest in time.