Fetching the paper…
Reading the bibliography…
Surprising videos, such as funny clips, creative performances, or visual illusions, attract significant attention.
Gruner, C.R.: Understanding laughter: The workings of wit & humor. Burnham Incorporated Pub (1978)
1978
Earlier work this paper cites.
Martin, M.W.: Humour and aesthetic enjoyment of incongruities. The British Journal of Aesthetics 23
1983
Earlier work this paper cites.
Kant, I.: Critique of judgment. Hackett Publishing (1987)
1987
Earlier work this paper cites.
Latta, R.L.: The basic humor process: A cognitive-shift theory and the case against incongruity. De Gruyter Mouton (1999)
1999
Earlier work this paper cites.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics. pp. 311–318 (2002)
2002
Earlier work this paper cites.
Boyd, B.: Laughter and literature: A play theory of humor. Philosophy and literature 28
2004
Earlier work this paper cites.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out. pp. 74–81 (2004)
2004
Earlier work this paper cites.
Billig, M.: Laughter and ridicule: Towards a social critique of humour. Sage (2005)
2005
Earlier work this paper cites.
Lamont, P., Wiseman, R.: Magic in theory: An introduction to the theoretical and psychological elements of conjuring. Univ of Hertfordshire Press (2005)
2005
Earlier work this paper cites.
Attardo, S.: A primer for the linguistics of humor. The primer of humor research 8
2008
Earlier work this paper cites.
Nelms, H.: Magic and showmanship: A handbook for conjurers. Courier Corporation (2012)
2012
Earlier work this paper cites.
Runco, M.A., Jaeger, G.J.: The standard definition of creativity. CREATIVITY RESEARCH JOURNAL 24
2012
Earlier work this paper cites.
Noordewier, M.K., Breugelmans, S.M.: On the valence of surprise. Cognition & emotion 27
2013
Earlier work this paper cites.
Malinowski, M., Fritz, M.: A multi-world approach to question answering about real-world scenes based on uncertain input. Advances in neural information processing systems 27
2014
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: Cider: Consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4566–4575 (2015)
2015
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M.S., Li, F.F.: Visual genome: Connecting language and vision using crowdsourced dense image annotations (2016)
2016
Earlier work this paper cites.
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: Movieqa: Understanding stories in movies through question-answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4631–4640 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering (2017)
2017
Earlier work this paper cites.
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: The kinetics human action video dataset (2017)
2017
Earlier work this paper cites.
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J.: Dense-captioning events in videos. In: Proceedings of the IEEE international conference on computer vision. pp. 706–715 (2017)
2017
Earlier work this paper cites.
Mun, J., Hongsuck Seo, P., Jung, I., Han, B.: Marioqa: Answering questions by watching gameplay videos. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2867–2875 (2017)
2017
Earlier work this paper cites.
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., Zhuang, Y.: Video question answering via gradually refined attention over appearance and motion. In: Proceedings of the 25th ACM international conference on Multimedia. pp. 1645–1653 (2017)
2017
Earlier work this paper cites.
Ye, Y., Zhao, Z., Li, Y., Chen, L., Xiao, J., Zhuang, Y.: Video question answering via attribute-augmented attention network learning. In: Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR ’17, ACM (Aug 2017). https://doi.org/10.1145/3077136.3080655, http://dx.doi.org/10.1145/3077136.3080655
2017
Earlier work this paper cites.
Lei, J., Yu, L., Bansal, M., Berg, T.L.: Tvqa: Localized, compositional video question answering. In: EMNLP (2018)
2018
Cited alongside, same era.
Zhao, Z., Zhang, Z., Xiao, S., Yu, Z., Yu, J., Cai, D., Wu, F., Zhuang, Y.: Open-ended long-form video question answering via adaptive hierarchical reinforced networks. In: IJCAI. vol. 2, p. 8 (2018)
2018
Cited alongside, same era.
Zhou, L., Xu, C., Corso, J.: Towards automatic learning of procedures from web instructional videos. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 32 (2018)
2018
Cited alongside, same era.
Castro, S., Hazarika, D., Pérez-Rosas, V., Zimmermann, R., Mihalcea, R., Poria, S.: Towards multimodal sarcasm detection (an _obviously_ perfect paper). In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Florence, Italy (7 2019)
2019
Cited alongside, same era.
Xu, L., Huang, H., Liu, J.: Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9878–9888 (2021)
2021
Later among the works it cites.
Castro, S., Wang, R., Huang, P., Stewart, I., Ignat, O., Liu, N., Stroud, J.C., Mihalcea, R.: Fiber: Fill-in-the-blanks as a challenging video understanding evaluation framework (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Hessel, J., Marasović, A., Hwang, J.D., Lee, L., Da, J., Zellers, R., Mankoff, R., Choi, Y.: Do androids laugh at electric sheep? humor "understanding" benchmarks from the new yorker caption contest (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hasan, M.K., Rahman, W., Bagher Zadeh, A., Zhong, J., Tanveer, M.I., Morency, L.P., Hoque, M.E.: UR-FUNNY: A multimodal language dataset for understanding humor. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). pp. 2046–2056. Association for Computational Linguistics, Hong Kong, China (Nov 2019). https://doi.org/10.18653/v1/D19-1211, https://aclanthology.org/D19-1211
2019
Cited alongside, same era.
Jang, Y., Song, Y., Kim, C.D., Yu, Y., Kim, Y., Kim, G.: Video Question Answering with Spatio-Temporal Reasoning. IJCV (2019)
2019
Cited alongside, same era.
Lei, J., Yu, L., Berg, T.L., Bansal, M.: Tvqa+: Spatio-temporal grounding for video question answering. In: Tech Report, arXiv (2019)
2019
Cited alongside, same era.
Li, X., Song, J., Gao, L., Liu, X., Huang, W., He, X., Gan, C.: Beyond rnns: Positional self-attention with co-attention for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33, pp. 8658–8665 (2019)
2019
Cited alongside, same era.
Miech, A., Zhukov, D., Alayrac, J.B., Tapaswi, M., Laptev, I., Sivic, J.: Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2630–2640 (2019)
2019
Cited alongside, same era.
Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., Tao, D.: Activitynet-qa: A dataset for understanding complex web videos via question answering (2019)
2019
Cited alongside, same era.
Zadeh, A., Chan, M., Liang, P.P., Tong, E., Morency, L.P.: Social-iq: A question answering benchmark for artificial social intelligence. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8807–8817 (2019)
2019
Cited alongside, same era.
Epstein, D., Chen, B., Vondrick, C.: Oops! predicting unintentional action in video. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (June 2020)
2020
Cited alongside, same era.
Li, C., Xu, H., Tian, J., Wang, W., Yan, M., Bi, B., Ye, J., Chen, H., Xu, G., Cao, Z., Zhang, J., Huang, S., Huang, F., Zhou, J., Si, L.: mplug: Effective and efficient vision-language learning by cross-modal skip-connections (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., Guo, B.: Advancing high-resolution video-language representation with large-scale video transcriptions. In: International Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Later among the works it cites.
Yang, P., Wang, X., Duan, X., Chen, H., Hou, R., Jin, C., Zhu, W.: Avqa: A dataset for audio-visual question answering on videos. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 3480–3491 (2022)
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
Corporation, N.T.N.: Kasou taishou. https://www.ntv.co.jp/kasoh/index.html , [Accessed 23-Apr-2023]
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Kougia, V., Fetzel, S., Kirchmair, T., Çano, E., Baharlou, S.M., Sharifzadeh, S., Roth, B.: Memegraphs: Linking memes to knowledge graphs (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Morreall, J.: Philosophy of Humor. In: Zalta, E.N., Nodelman, U. (eds.) The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Summer 2023 edn. (2023)
2023
Closest in time.
Muhammad Maaz, Hanoona Rasheed, S.K., Khan, F.: Video-chatgpt. https://github.com/mbzuai-oryx/Video-ChatGPT (2023)
2023
Closest in time.
OpenAI: Gpt-4 technical report (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Yang, J., Peng, W., Li, X., Guo, Z., Chen, L., Li, B., Ma, Z., Zhou, K., Zhang, W., Loy, C.C., et al.: Panoptic video scene graph generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18675–18685 (2023)
2023
Closest in time.
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Jiang, C., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., Huang, F.: mplug-owl: Modularization empowers large language models with multimodality (2023)
2023
Closest in time.