Fetching the paper…
Reading the bibliography…
In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research.
J. Chen and C. M. Ho, “Mm-vit: Multi-modal video transformer for compressed video action recognition,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 1910–1921
1921
Earlier work this paper cites.
D. L. Robins, D. Fein, and M. Barton, “M-chat-r/f: The modified checklist for autism in toddlers, revised with follow-up,” Online, 2009, available at: https://www.mchatscreen.com/
2009
Earlier work this paper cites.
J.-C. Lin, C.-H. Wu, and W.-L. Wei, “Error weighted semi-coupled hidden markov model for audio-visual emotion recognition,” IEEE Transactions on Multimedia , vol. 14, no. 1, pp. 142–156, 2011
2011
Earlier work this paper cites.
W. Jones and A. Klin, “Attention to eyes is present but in decline in 2–6-month-old infants later diagnosed with autism,” Nature , vol. 504, no. 7480, pp. 427–431, 2013
2013
Earlier work this paper cites.
J. Rehg, G. Abowd, A. Rozga, M. Romero, M. Clements, S. Sclaroff, I. Essa, O. Ousley, Y. Li, C. Kim et al. , “Decoding children’s social behavior,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2013, pp. 3414–3421
2013
Earlier work this paper cites.
S. Rajagopalan, A. Dhall, and R. Goecke, “Self-stimulatory behaviours in the wild for autism diagnosis,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2013, pp. 755–761
2013
Earlier work this paper cites.
S. S. Rajagopalan and R. Goecke, “Detecting self-stimulatory behaviours for autism diagnosis,” in 2014 IEEE International Conference on Image Processing (ICIP) . IEEE, 2014, pp. 1470–1474
2014
Earlier work this paper cites.
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
M. Del Coco, M. Leo, P. Carcagnì, F. Fama, L. Spadaro, L. Ruta, G. Pioggia, and C. Distante, “Study of mechanisms of social interaction stimulation in autism spectrum disorder by assisted humanoid robot,” IEEE Transactions on Cognitive and Developmental Systems , vol. 10, no. 4, pp. 993–1004, 2017
2017
Earlier work this paper cites.
A. Zunino, P. Morerio, A. Cavallo, C. Ansuini, J. Podda, F. Battaglia, E. Veneselli, C. Becchio, and V. Murino, “Video gesture analysis for autism spectrum disorder detection,” in 2018 24th international conference on pattern recognition (ICPR) . IEEE, 2018, pp. 3421–3426
2018
Earlier work this paper cites.
G. Dawson, K. Campbell, J. Hashemi, S. J. Lippmann, V. Smith, K. Carpenter, H. Egger, S. Espinosa, S. Vermeer, J. Baker et al. , “Atypical postural control can be detected via computer vision analysis in toddlers with autism spectrum disorder,” Scientific reports , vol. 8, no. 1, p. 17008, 2018
2018
Earlier work this paper cites.
K. B. Martin, Z. Hammal, G. Ren, J. F. Cohn, J. Cassell, M. Ogihara, J. C. Britton, A. Gutierrez, and D. S. Messinger, “Objective measurement of head movement differences in children with and without autism spectrum disorder,” Molecular autism , vol. 9, pp. 1–10, 2018
2018
Earlier work this paper cites.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 247–263
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 570–586
2018
Earlier work this paper cites.
A. S. Nahmias, M. Pellecchia, A. C. Stahmer, and D. S. Mandell, “Effectiveness of community-based early intervention for children with autism spectrum disorder: A meta-analysis,” Journal of Child Psychology and Psychiatry , vol. 60, no. 11, pp. 1200–1209, 2019
2019
Earlier work this paper cites.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5492–5501
2019
Earlier work this paper cites.
H. Alamri, V. Cartillier, A. Das, J. Wang, A. Cherian, I. Essa, D. Batra, T. K. Marks, C. Hori, P. Anderson et al. , “Audio visual scene-aware dialog,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 7558–7567
2019
Earlier work this paper cites.
E. A. Fuller and A. P. Kaiser, “The effects of early intervention on social communication outcomes for children with autism spectrum disorder: A meta-analysis,” Journal of autism and developmental disorders , vol. 50, pp. 1683–1700, 2020
2020
Earlier work this paper cites.
E. Billing, T. Belpaeme, H. Cai, H.-L. Cao, A. Ciocan, C. Costescu, D. David, R. Homewood, D. Hernandez Garcia, P. Gómez Esteban et al. , “The dream dataset: Supporting a data-driven study of autism spectrum disorder and robot enhanced therapy,” PloS one , vol. 15, no. 8, p. e0236939, 2020
2020
Earlier work this paper cites.
G. Riva, E. Riva et al. , “De-enigma: Multimodal human-robot interaction for teaching and expanding social imagination in autistic children,” Cyberpsychology, behavior and social networking , vol. 23, no. 11, pp. 806–807, 2020
2020
Earlier work this paper cites.
P. Pandey, A. Prathosh, M. Kohli, and J. Pritchard, “Guided weak supervision for action recognition with scarce data to assess skills of children with autism,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2020, pp. 463–470
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani, “Listen to look: Action recognition by previewing audio,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 457–10 467
2020
Cited alongside, same era.
F. Tao and C. Busso, “End-to-end audiovisual speech recognition system with multitask learning,” IEEE Transactions on Multimedia , vol. 23, pp. 1–11, 2020
2020
Cited alongside, same era.
Y. Tian, D. Li, and C. Xu, “Unified multisensory perception: Weakly-supervised audio-visual video parsing,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 . Springer, 2020, pp. 436–454
2020
Cited alongside, same era.
Y. Zhu, Y. Wu, Y. Yang, and Y. Yan, “Describing unseen videos via multi-modal cooperative dialog agents,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16 . Springer, 2020, pp. 153–169
2020
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong, “Audio–visual segmentation,” in European Conference on Computer Vision . Springer, 2022, pp. 386–403
2022
Later among the works it cites.
G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 19 108–19 118
2022
Later among the works it cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 730–27 744, 2022
2022
Later among the works it cites.
OpenAI, “Gpt-4 technical report,” 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 736–740
2020
Cited alongside, same era.
K. Bottema-Beutel, S. K. Kapp, J. N. Lester, N. J. Sasson, and B. N. Hand, “Avoiding ableist language: Suggestions for autism researchers,” Autism in adulthood , vol. 3, no. 1, pp. 18–29, 2021
2021
Cited alongside, same era.
M. van’t Hof, C. Tisseur, I. van Berckelear-Onnes, A. van Nieuwenhuyzen, A. M. Daniels, M. Deen, H. W. Hoek, and W. A. Ester, “Age at autism spectrum disorder diagnosis: A systematic review and meta-analysis from 2012 to 2019,” Autism , vol. 25, no. 4, pp. 862–873, 2021
2021
Cited alongside, same era.
F. Cilia, R. Carette, M. Elbattah, G. Dequen, J.-L. Guérin, J. Bosche, L. Vandromme, B. Le Driant et al. , “Computer-aided screening of autism spectrum disorder: Eye-tracking study using data visualization and deep learning,” JMIR human factors , vol. 8, no. 4, p. e27706, 2021
2021
Cited alongside, same era.
F. Negin, B. Ozyer, S. Agahian, S. Kacdioglu, and G. T. Ozyer, “Vision-assisted recognition of stereotype behaviors for early diagnosis of autism spectrum disorders,” Neurocomputing , vol. 446, pp. 145–155, 2021
2021
Cited alongside, same era.
C. Xue, X. Zhong, M. Cai, H. Chen, and W. Wang, “Audio-visual event localization by learning spatial and semantic co-attention,” IEEE Transactions on Multimedia , vol. 25, pp. 418–429, 2021
2021
Cited alongside, same era.
R. Gao and K. Grauman, “Visualvoice: Audio-visual speech separation with cross-modal consistency,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2021, pp. 15 490–15 500
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. O. Ribeiro, M. Grellert, and J. T. Carvalho, “Stimming behavior dataset-unifying stereotype behavior dataset in the wild,” in 2023 IEEE 36th International Symposium on Computer-Based Medical Systems (CBMS) . IEEE, 2023, pp. 225–230
2023
Later among the works it cites.
C. Huang, Y. Tian, A. Kumar, and C. Xu, “Egocentric audio-visual object localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 910–22 921
2023
Later among the works it cites.
S. Mo and Y. Tian, “Audio-visual grouping network for sound localization from mixtures,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 565–10 574
2023
Later among the works it cites.
Y. Jiang, J. Yin, and Y. Dang, “Leveraging the video-level semantic consistency of event for audio-visual event localization,” IEEE Transactions on Multimedia , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 180–15 190
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Later among the works it cites.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Gao, X. Jiang, Y. Yang, D. Li, and L. Qiu, “Unsupervised video anomaly detection for stereotypical behaviours in autism,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.