Fetching the paper…
Reading the bibliography…
Multi-class multi-instance segmentation is the task of identifying masks for multiple object classes and multiple instances of the same class within an image.
“Ai2-thor: An interactive 3d environment for visual ai,”
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al., · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Segan: Segmenting and generating the invisible,”
Kiana Ehsani, Roozbeh Mottaghi, and Ali Farhadi, · 2018
Earlier work this paper cites.
“Chalet: Cornell house agent learning environment,”
Claudia Yan, Dipendra Misra, Andrew Bennnett, Aaron Walsman, Yonatan Bisk, and Yoav Artzi, · 2018
Earlier work this paper cites.
“Fast online object tracking and segmentation: A unifying approach,”
Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip HS Torr, · 2019
Earlier work this paper cites.
“Habitat: A platform for embodied ai research,”
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al., · 2019
Earlier work this paper cites.
“Vrkitchen: an interactive 3d virtual environment for task-oriented learning,”
Xiaofeng Gao, Ran Gong, Tianmin Shu, Xu Xie, Shu Wang, and Song-Chun Zhu, · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Earlier work this paper cites.
“Threedworld: A platform for interactive multi-modal physical simulation,”
Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, et al., · 2020
Earlier work this paper cites.
“Real–sim–real transfer for real-world robot control policy learning with deep reinforcement learning,”
Naijun Liu, Yinghao Cai, Tao Lu, Rui Wang, and Shuo Wang, · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al., · 2020
Earlier work this paper cites.
“igibson 2.0: Object-centric simulation for robot learning of everyday household tasks,”
Chengshu Li, Fei Xia, Roberto Martín-Martín, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, et al., · 2021
Earlier work this paper cites.
“Visual room rearrangement,”
Luca Weihs, Matt Deitke, Aniruddha Kembhavi, and Roozbeh Mottaghi, · 2021
Cited alongside, same era.
“Segment anything in medical images,”
Jun Ma and Bo Wang, · 2023
Cited alongside, same era.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al., · 2023
Cited alongside, same era.
“Llama: Open and efficient foundation language models,”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al., · 2023
Cited alongside, same era.
“Caption anything: Interactive image description with diverse multimodal controls,”
Teng Wang, Jinrui Zhang, Junjie Fei, Yixiao Ge, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao, Ying Shan, et al., · 2023
Later among the works it cites.
“Av-sam: Segment anything model meets audio-visual localization and segmentation,”
Shentong Mo and Yapeng Tian, · 2023
Later among the works it cites.
“Anything-3d: Towards single-view anything reconstruction in the wild,”
Qiuhong Shen, Xingyi Yang, and Xinchao Wang, · 2023
Later among the works it cites.
“Sam-path: A segment anything model for semantic segmentation in digital pathology,”
Jingwei Zhang, Ke Ma, Saarthak Kapse, Joel Saltz, Maria Vakalopoulou, Prateek Prasanna, and Dimitris Samaras, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee, · 2023
Cited alongside, same era.
“Seggpt: Segmenting everything in context,”
Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang, · 2023
Cited alongside, same era.
Wei Ji, Jingjing Li, Qi Bi, Wenbo Li, and Li Cheng, · 2023
Cited alongside, same era.
“Instruct2act: Mapping multi-modality instructions to robotic actions with large language model,”
Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, and Hongsheng Li, · 2023
Cited alongside, same era.
“Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion,”
Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang, · 2023
Cited alongside, same era.
“Edit everything: A text-guided generative system for images editing,”
Defeng Xie, Ruichen Wang, Jian Ma, Chen Chen, Haonan Lu, Dong Yang, Fobo Shi, and Xiaodong Lin, · 2023
Cited alongside, same era.
“Can sam segment anything? when sam meets camouflaged object detection,”
Lv Tang, Haoke Xiao, and Bo Li, · 2023
Cited alongside, same era.
Dongsheng Han, Chaoning Zhang, Yu Qiao, Maryam Qamar, Yuna Jung, SeungKyu Lee, Sung-Ho Bae, and Choong Seon Hong, · 2023
Cited alongside, same era.
Shurong Chai, Rahul Kumar Jain, Shiyu Teng, Jiaqing Liu, Yinhao Li, Tomoko Tateyama, and Yen-wei Chen, · 2023
Later among the works it cites.
“Adaptive low rank adaptation of segment anything to salient object detection,”
Ruikai Cui, Siyuan He, and Shi Qiu, · 2023
Later among the works it cites.
“Aquasam: Underwater image foreground segmentation,”
Muduo Xu, Jianhao Su, and Yutao Liu, · 2023
Later among the works it cites.
“Personalize segment anything model with one shot,”
Renrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan, Junting Pan, Hao Dong, Peng Gao, and Hongsheng Li, · 2023
Later among the works it cites.
“Semantic-sam: Segment and recognize anything at any granularity,”
Feng Li, Hao Zhang, Peize Sun, Xueyan Zou, Shilong Liu, Jianwei Yang, Chunyuan Li, Lei Zhang, and Jianfeng Gao, · 2023
Later among the works it cites.
“Online robot navigation and and manipulation with distilled vision-language models,”
Kangcheng Liu, Xinhu Zheng, Chaoqun Wang, Hesheng Wang, Ming Liu, and Kai Tang, · 2024
Closest in time.
“Open-vocabulary sam: Segment and recognize twenty-thousand classes interactively,”
Haobo Yuan, Xiangtai Li, Chong Zhou, Yining Li, Kai Chen, and Chen Change Loy, · 2024
Closest in time.