Fetching the paper…
Reading the bibliography…
3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence for describing each located object.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J.; Karpathy, A.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
Enet: A deep neural network architecture for real-time semantic segmentation
Paszke, A.; Chaurasia, A.; Kim, S.; and Culurciello, E. 2016 · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J.; Karpathy, A.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
Enet: A deep neural network architecture for real-time semantic segmentation
Paszke, A.; Chaurasia, A.; Kim, S.; and Culurciello, E. 2016 · 2016
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Lu, J.; Xiong, C.; Parikh, D.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Dense captioning with joint inference and visual context
Yang, L.; Tang, K.; Yang, J.; and Li, L.-J. 2017 · 2017
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Lu, J.; Xiong, C.; Parikh, D.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Dense captioning with joint inference and visual context
Yang, L.; Tang, K.; Yang, J.; and Li, L.-J. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Align2ground: Weakly supervised phrase grounding guided by image-caption alignment
Datta, S.; Sikka, K.; Roy, A.; Ahuja, K.; Parikh, D.; and Divakaran, A. 2019 · 2019
Earlier work this paper cites.
Fast, diverse and accurate image captioning guided by part-of-speech
Deshpande, A.; Aneja, J.; Wang, L.; Schwing, A. G.; and Forsyth, D. 2019 · 2019
Earlier work this paper cites.
Dense relational captioning: Triple-stream networks for relationship-based captioning
Kim, D.-J.; Choi, J.; Oh, T.-H.; and Kweon, I. S. 2019 · 2019
Earlier work this paper cites.
Entangled transformer for image captioning
Li, G.; Zhu, L.; Liu, P.; and Yang, Y. 2019 · 2019
Cited alongside, same era.
Deep hough voting for 3d object detection in point clouds
Qi, C. R.; Litany, O.; He, K.; and Guibas, L. J. 2019 · 2019
Cited alongside, same era.
MUCH: Mutual Coupling Enhancement of Scene Recognition and Dense Captioning
Song, X.; Wang, B.; Chen, G.; and Jiang, S. 2019 · 2019
Cited alongside, same era.
Hierarchical attention network for image captioning
Wang, W.; Chen, Z.; and Hu, H. 2019 · 2019
Cited alongside, same era.
Context and attribute grounded dense captioning
Yin, G.; Sheng, L.; Liu, B.; Yu, N.; Wang, X.; and Shao, J. 2019 · 2019
Cited alongside, same era.
Align2ground: Weakly supervised phrase grounding guided by image-caption alignment
Datta, S.; Sikka, K.; Roy, A.; Ahuja, K.; Parikh, D.; and Divakaran, A. 2019 · 2019
X-linear attention networks for image captioning
Pan, Y.; Yao, T.; Li, Y.; and Mei, T. 2020 · 2020
Later among the works it cites.
Comprehensive Image Captioning via Scene Graph Decomposition
Zhong, Y.; Wang, L.; Chen, J.; Yu, D.; and Li, Y. 2020 · 2020
Later among the works it cites.
Group-free 3d object detection via transformers
Liu, Z.; Zhang, Z.; Cao, Y.; Hu, H.; and Tong, X. 2021 · 2021
Later among the works it cites.
Dual-level collaborative transformer for image captioning
Luo, Y.; Ji, J.; Sun, X.; Cao, L.; Wu, Y.; Huang, F.; Lin, C.-W.; and Ji, R. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning with adaptive attention on visual and non-visual words
Zhang, X.; Sun, X.; Luo, Y.; Ji, J.; Zhou, Y.; Wu, Y.; Huang, F.; and Ji, R. 2021 · 2021
Later among the works it cites.
Group-free 3d object detection via transformers
Liu, Z.; Zhang, Z.; Cao, Y.; Hu, H.; and Tong, X. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast, diverse and accurate image captioning guided by part-of-speech
Deshpande, A.; Aneja, J.; Wang, L.; Schwing, A. G.; and Forsyth, D. 2019 · 2019
Cited alongside, same era.
Dense relational captioning: Triple-stream networks for relationship-based captioning
Kim, D.-J.; Choi, J.; Oh, T.-H.; and Kweon, I. S. 2019 · 2019
Cited alongside, same era.
Entangled transformer for image captioning
Li, G.; Zhu, L.; Liu, P.; and Yang, Y. 2019 · 2019
Cited alongside, same era.
Deep hough voting for 3d object detection in point clouds
Qi, C. R.; Litany, O.; He, K.; and Guibas, L. J. 2019 · 2019
Cited alongside, same era.
MUCH: Mutual Coupling Enhancement of Scene Recognition and Dense Captioning
Song, X.; Wang, B.; Chen, G.; and Jiang, S. 2019 · 2019
Cited alongside, same era.
Hierarchical attention network for image captioning
Wang, W.; Chen, Z.; and Hu, H. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Dual-level collaborative transformer for image captioning
Luo, Y.; Ji, J.; Sun, X.; Cao, L.; Wu, Y.; Huang, F.; Lin, C.-W.; and Ji, R. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning with adaptive attention on visual and non-visual words
Zhang, X.; Sun, X.; Luo, Y.; Ji, J.; Zhou, Y.; Wu, Y.; Huang, F.; and Ji, R. 2021 · 2021
Later among the works it cites.
3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds
Cai, D.; Zhao, L.; Zhang, J.; Sheng, L.; and Xu, D. 2022 · 2022
Closest in time.
MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes
Jiao, Y.; Chen, S.; Jie, Z.; Chen, J.; Ma, L.; and Jiang, Y.-G. 2022 · 2022
Closest in time.
Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds
Wang, H.; Zhang, C.; Yu, J.; and Cai, W. 2022 · 2022
Closest in time.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Yuan, Z.; Yan, X.; Liao, Y.; Guo, Y.; Li, G.; Cui, S.; and Li, Z. 2022 · 2022
Closest in time.
3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds
Cai, D.; Zhao, L.; Zhang, J.; Sheng, L.; and Xu, D. 2022 · 2022
Closest in time.
MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes
Jiao, Y.; Chen, S.; Jie, Z.; Chen, J.; Ma, L.; and Jiang, Y.-G. 2022 · 2022
Closest in time.
Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds
Wang, H.; Zhang, C.; Yu, J.; and Cai, W. 2022 · 2022
Closest in time.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Yuan, Z.; Yan, X.; Liao, Y.; Guo, Y.; Li, G.; Cui, S.; and Li, Z. 2022 · 2022
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2057
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2057
Closest in time.