Fetching the paper…
Reading the bibliography…
In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing.
1904
Earlier work this paper cites.
1906
Earlier work this paper cites.
H. Suzuki, K. Shimomura, T. Hirakawa, T. Yamashita, H. Fujiyoshi, S. Okubo, N. Takuya, and W. Siyuan, “Human-like guidance by generating navigation using spatial-temporal scene graph,” in 2024 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 2024, pp. 1988–1995
1995
Earlier work this paper cites.
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
X. Ma, Z. Song, Y. Li, and G. R. Arce, “Block-based mask optimization for optical lithography,” Applied optics , vol. 52, no. 14, pp. 3351 – 3363, 2013
2013
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio, Deep learning . MIT press Cambridge, 2016, vol. 1, no. 2
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Kaduk, The Semantic Link: Action & Language. An Investigation of Relations Between Different Cognitive Domains in Early Development . Lancaster University (United Kingdom), 2017
2017
Earlier work this paper cites.
Z. Li, J. Zhang, K. Zhang, and Z. Li, “Visual tracking with weighted adaptive local sparse appearance model via spatio-temporal context learning,” IEEE Transactions on Image Processing , vol. 27, no. 9, pp. 4478–4489, 2018
2018
Earlier work this paper cites.
T. Song, Y. Song, Y. Wang, and X. Huang, “Learning efficient residual networks through dense connecting,” in 2018 37th Chinese Control Conference (CCC) , 2018, pp. 9181–9185
2018
Earlier work this paper cites.
A. Zadeh, M. Chan, P. P. Liang et al. , “Social-iq: A question answering benchmark for artificial social intelligence,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8807–8817
2019
Earlier work this paper cites.
A. Balakrishnan and J. V. Deshmukh, “Structured reward shaping using signal temporal logic specifications,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 3481–3486
2019
Earlier work this paper cites.
H. Yang, Z. Xiong, J. Zhao, D. Niyato, L. Xiao, and Q. Wu, “Deep reinforcement learning-based intelligent reflecting surface for secure wireless communications,” IEEE Transactions on Wireless Communications , vol. 20, no. 1, p. 375–388, Jan. 2021. [Online]. Available: http://dx.doi.org/10.1109/TWC.2020.3024860
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. K. Shaydyuk and E. B. John, “Fpga implementation of mobilenetv2 cnn model using semi-streaming architecture for low power inference applications,” 2020 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA/BDCloud/SocialCom/SustainCom) , pp. 160–167, 2020
2020
Earlier work this paper cites.
K. Shi, G. Jia, Y. Li, Y. Yin, C. Jiang, and L. Zhou, “Quantitative analysis of real-time performance and hardware requirements for edge computing platform,” Data Science and Informetrics , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
N. Tawfik, H. A. Elnemr, M. Fakhr, M. I. Dessouky, and F. E. A. El-Samie, “Multimodal medical image fusion using stacked auto-encoder in nsct domain,” Journal of Digital Imaging , vol. 35, no. 5, pp. 1308–1325, 2022
2022
Cited alongside, same era.
Z.-X. Wang, P.-N. Shao, and C. Deng, “Caffe inference acceleration method on heterogeneous parallel platform,” Computer Systems & Applications , vol. 31, no. 2, pp. 220–226, 2022
2022
Cited alongside, same era.
M.-H. Guo, T.-X. Xu, J.-J. Liu, Z.-N. Liu, P.-T. Jiang, T.-J. Mu, S.-H. Zhang, R. R. Martin, M.-M. Cheng, and S.-M. Hu, “Attention mechanisms in computer vision: A survey,” Computational Visual Media , vol. 8, no. 3, p. 331–368, Mar. 2022. [Online]. Available: http://dx.doi.org/10.1007/s41095-022-0271-y
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
X. Wang and P. Peng, “Open-r1-video,” https://github.com/Wang-Xiaodong1899/Open-R1-Video, 2025
2025
Closest in time.