Fetching the paper…
Reading the bibliography…
Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data.
Ratnasingham, S., Hebert, P.D.: Bold: The barcode of life data system (http://www. barcodinglife. org). Molecular ecology notes 7
2007
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255 (2009). Ieee
2009
Earlier work this paper cites.
Samanta, R., Ghosh, I.: Tea insect pests classification based on artificial neural networks. International Journal of Computer Engineering Science (IJCES) 2
2012
Earlier work this paper cites.
Wang, J., Lin, C., Ji, L., Liang, A.: A new automatic identification system of insect images at the order level. Knowledge-Based Systems 33
2012
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25
2012
Earlier work this paper cites.
Venugoban, K., Ramanan, A.: Image classification of paddy field insect pests using gradient-based features. International Journal of Machine Learning and Computing 4
2014
Earlier work this paper cites.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755 (2014). Springer
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Xie, C., Zhang, J., Li, R., Li, J., Hong, P., Xia, J., Chen, P.: Automatic classification for field crop insects via multiple-task sparse representation and multiple-kernel learning. Computers and Electronics in Agriculture 119
2015
Earlier work this paper cites.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–9 (2015)
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28
2015
Earlier work this paper cites.
Liu, Z., Gao, J., Yang, G., Zhang, H., He, Y.: Localization and classification of paddy field pests using a saliency map and deep convolutional neural network. Scientific reports 6
2016
Earlier work this paper cites.
Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: European Conference on Computer Vision, pp. 69–84 (2016). Springer
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)
2016
Earlier work this paper cites.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., Berg, A.C.: Ssd: Single shot multibox detector. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, pp. 21–37 (2016). Springer
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117–2125 (2017)
2017
Earlier work this paper cites.
Xie, C., Wang, R., Zhang, J., Chen, P., Dong, W., Li, R., Chen, T., Chen, H.: Multi-level learning features for automatic classification of field crop pests. Computers and Electronics in Agriculture 152
2018
Earlier work this paper cites.
Deng, L., Wang, Y., Han, Z., Yu, R.: Research on insect pest image detection and recognition based on bio-inspired methods. Biosystems Engineering 169
2018
Earlier work this paper cites.
Alfarisy, A.A., Chen, Q., Guo, M.: Deep learning based classification for paddy pests & diseases recognition. In: Proceedings of 2018 International Conference on Mathematics and Artificial Intelligence, pp. 21–25 (2018)
2018
Earlier work this paper cites.
Stork, N.E.: How many species of insects and other terrestrial arthropods are there on earth? Annual review of entomology 63
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Zhang, S., Wen, L., Bian, X., Lei, Z., Li, S.Z.: Single-shot refinement neural network for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4203–4212 (2018)
2018
Earlier work this paper cites.
Redmon, J., Farhadi, A.: Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 (2018)
2018
Earlier work this paper cites.
Liu, L., Wang, R., Xie, C., Yang, P., Wang, F., Sudirman, S., Liu, W.: Pestnet: An end-to-end deep learning approach for large-scale multi-class pest detection and classification. Ieee Access 7
2019
Earlier work this paper cites.
Wu, X., Zhan, C., Lai, Y.-K., Cheng, M.-M., Yang, J.: Ip102: A large-scale benchmark dataset for insect pest recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8787–8796 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729–9738 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: International Conference on Machine Learning, pp. 1597–1607 (2020). PMLR
2020
Earlier work this paper cites.
Alves, A.N., Souza, W.S., Borges, D.L.: Cotton pests classification in field-based images using deep residual networks. Computers and Electronics in Agriculture 174
2020
Earlier work this paper cites.
Bollis, E., Pedrini, H., Avila, S.: Weakly supervised learning guided by activation mapping applied to a novel citrus pest benchmark. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 70–71 (2020)
2020
Earlier work this paper cites.
Nguyen, X.-B., Lee, G.S., Kim, S.H., Yang, H.J.: Self-supervised learning based on spatial awareness for medical image analysis. IEEE Access 8
2020
Cited alongside, same era.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Nanni, L., Maguolo, G., Pancino, F.: Insect pest image detection and recognition based on bio-inspired methods. Ecological Informatics 57
2020
Cited alongside, same era.
2023
Later among the works it cites.
Chen, Y., Shen, X., Liu, Y., Tao, Q., Suykens, J.A.: Jigsaw-vit: Learning jigsaw puzzles in vision transformer. Pattern Recognition Letters 166
2023
Later among the works it cites.
Nguyen, X.-B., Duong, C.N., Li, X., Gauch, S., Seo, H.-S., Luu, K.: Micron-bert: Bert-based facial micro-expression recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1482–1492 (2023)
2023
Later among the works it cites.
Truong, T.-D., Le, N., Raj, B., Cothren, J., Luu, K.: Fredom: Fairness domain adaptation approach to semantic scene understanding. In: IEEE/CVF Computer Vision and Pattern Recognition (CVPR) (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ayan, E., Erbay, H., Varçın, F.: Crop pest classification with a genetic algorithm-based weighted ensemble of deep convolutional neural networks. Computers and Electronics in Agriculture 179
2020
Cited alongside, same era.
Fan, D.-P., Ji, G.-P., Sun, G., Cheng, M.-M., Shen, J., Shao, L.: Camouflaged object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2777–2787 (2020)
2020
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Wang, R., Liu, L., Xie, C., Yang, P., Li, R., Zhou, M.: Agripest: A large-scale domain-specific benchmark dataset for practical agricultural pest detection in the wild. Sensors 21
2021
Cited alongside, same era.
Badirli, S., Akata, Z., Mohler, G., Picard, C., Dundar, M.M.: Fine-grained zero-shot learning with dna as side information. Advances in Neural Information Processing Systems 34
2021
Cited alongside, same era.
Van Horn, G., Cole, E., Beery, S., Wilber, K., Belongie, S., Mac Aodha, O.: Benchmarking representation learning for natural world image collections. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12884–12893 (2021)
2021
Cited alongside, same era.
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9650–9660 (2021)
2021
Cited alongside, same era.
Truong, T.-D., Duong, C.N., Quach, K.G., Le, N., Bui, T.D., Luu, K.: Liaad: Lightweight attentive angular distillation for large-scale age-invariant face recognition. Neurocomputing 543
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Pham, H., Dai, Z., Ghiasi, G., Kawaguchi, K., Liu, H., Yu, A.W., Yu, J., Chen, Y.-T., Luong, M.-T., Wu, Y., et al
2023
Later among the works it cites.
2023
Later among the works it cites.
Luo, Z., Zhao, P., Xu, C., Geng, X., Shen, T., Tao, C., Ma, J., Lin, Q., Jiang, D.: Lexlip: Lexicon-bottlenecked language-image pre-training for large-scale image-text sparse retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11206–11217 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhao, Y., Misra, I., Krähenbühl, P., Girdhar, R.: Learning video representations from large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6586–6597 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., Stoica, I., Xing, E.P.: Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality (2023). https://lmsys.org/blog/2023-03-30-vicuna/
2023
Later among the works it cites.
He, C., Li, K., Zhang, Y., Tang, L., Zhang, Y., Guo, Z., Li, X.: Camouflaged object detection with feature decomposition and edge reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22046–22055 (2023)
2023
Later among the works it cites.
Li, C., Gan, Z., Yang, Z., Yang, J., Li, L., Wang, L., Gao, J., et al
2024
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems 36
2024
Later among the works it cites.
Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26296–26306 (2024)
2024
Later among the works it cites.
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., Gao, J.: Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
Nguyen, H.-Q., Truong, T.-D., Nguyen, X.B., Dowling, A., Li, X., Luu, K.: Insect-foundation: A foundation model and large-scale 1m dataset for visual insect understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21945–21955 (2024)
2024
Later among the works it cites.
Truong, T.-D., Nguyen, H.-Q., Raj, B., Luu, K.: Fairness continual learning approach to semantic scene understanding in open-world environments. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
Chen, J., Zhang, A.: FedMBridge: Bridgeable multimodal federated learning. In: Forty-first International Conference on Machine Learning (2024). https://openreview.net/forum?id=jrHUbftLd6
2024
Later among the works it cites.
Ren, Z., Huang, Z., Wei, Y., Zhao, Y., Fu, D., Feng, J., Jin, X.: Pixellm: Pixel reasoning with large multimodal model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26374–26383 (2024)
2024
Later among the works it cites.
Yuan, Y., Li, W., Liu, J., Tang, D., Luo, X., Qin, C., Zhang, L., Zhu, J.: Osprey: Pixel understanding with visual instruction tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 28202–28211 (2024)
2024
Later among the works it cites.
Zhang, T., Li, X., Fei, H., Yuan, H., Wu, S., Ji, S., Loy, C.C., YAN, S.: OMG-LLaVA: Bridging image-level, object-level, pixel-level reasoning and understanding. In: The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024). https://openreview.net/forum?id=WeoNd6PRqS
2024
Later among the works it cites.
Fei, H., Wu, S., Zhang, H., Chua, T.-S., Yan, S.: VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing. CoRR (2024)
2024
Later among the works it cites.
Wu, S., Fei, H., Qu, L., Ji, W., Chua, T.-S.: NExT-GPT: Any-to-any multimodal LLM. In: Proceedings of the International Conference on Machine Learning, pp. 53366–53397 (2024)
2024
Later among the works it cites.
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., Chen, W.: Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9556–9567 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
He, C., Li, K., Zhang, Y., Xu, G., Tang, L., Zhang, Y., Guo, Z., Li, X.: Weakly-supervised concealed object segmentation with sam-based pseudo labeling and multi-scale feature grouping. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.