Fetching the paper…
Reading the bibliography…
While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood.
Generalized inverses and ranks of block matrices
Meyer, Jr, C. D · 1973
Earlier work this paper cites.
A coupled HMM approach to video-realistic speech animation
Xie, L. and Liu, Z.-Q · 2006
Earlier work this paper cites.
Analyzing the Behavior of Visual Question Answering Models
Agrawal, A., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
On implementing 2D rectangular assignment algorithms
Crouse, D. F · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G. A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A · 2016
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., and Fidler, S · 2016
Earlier work this paper cites.
TALL: temporal activity localization via language query
Gao, J., Sun, C., Yang, Z., and Nevatia, R · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
An Analysis of Visual Question Answering Algorithms
Kafle, K. and Kanan, C · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A · 2018
Earlier work this paper cites.
Being Negative but Constructively: Lessons Learnt from Creating Better Visual Question Answering Datasets
Chao, W.-L., Hu, H., and Sha, F · 2018
Earlier work this paper cites.
TVQA: Localized, compositional video question answering
Lei, J., Yu, L., Bansal, M., and Berg, T · 2018
Earlier work this paper cites.
Approximating cnns with bag-of-local-features models works surprisingly well on imagenet
Brendel, W. and Bethge, M · 2019
Earlier work this paper cites.
RUBi: Reducing Unimodal Biases for Visual Question Answering
Cadène, R., Dancette, C., Ben-younes, H., Cord, M., and Parikh, D · 2019
Earlier work this paper cites.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Clark, C., Yatskar, M., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Explicit Bias Discovery in Visual Question Answering Models
Manjunatha, V., Saini, N., and Davis, L. S · 2019
Earlier work this paper cites.
Analyzing compositionality in visual question answering
Subramanian, S., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., and Tao, D · 2019
Earlier work this paper cites.
Social-IQ: A question answering benchmark for artificial social intelligence
Zadeh, A., Chan, M., Liang, P. P., Tong, E., and Morency, L · 2019
Earlier work this paper cites.
Low-Rank Bottleneck in Multi-head Attention Models
Bhojanapalli, S., Yun, C., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
Does my multimodal model learn cross-modal interactions? It’s harder to tell than you might think!
Hessel, J. and Lee, L · 2020
Earlier work this paper cites.
Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Otani, M., Nakashima, Y., Rahtu, E., and Heikkilä, J · 2020
Earlier work this paper cites.
On Modality Bias in the TVQA dataset
Winterbottom, T., Xiao, S., McLean, A., and Moubayed, N. A · 2020
Cited alongside, same era.
CLEVRER: collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2020
Cited alongside, same era.
Putting Visual Object Recognition in Context
Zhang, M., Tseng, C., and Kreiman, G · 2020
Cited alongside, same era.
Large image datasets: A pyrrhic win for computer vision?
Birhane, A. and Prabhu, V. U · 2021
Cited alongside, same era.
When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes
Bomatter, P., Zhang, M., Karev, D., Madan, S., Tseng, C., and Kreiman, G · 2021
Cited alongside, same era.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Chefer, H., Gur, S., and Wolf, L · 2021
VQuAD: Video Question Answering Diagnostic Dataset
Gupta, V., Patro, B. N., Parihar, H., and Namboodiri, V. P · 2022
Later among the works it cites.
Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding
Hao, J., Sun, H., Ren, P., Wang, J., Qi, Q., and Liao, J · 2022
Later among the works it cites.
Gender and Racial Bias in Visual Question Answering Datasets
Hirota, Y., Nakashima, Y., and Garcia, N · 2022
Later among the works it cites.
Voxel-wise intermodal coupling analysis of two or more modalities using local covariance decomposition
Hu, F., Weinstein, S. M., Baller, E. B., Valcarcel, A. M., Adebimpe, A., Raznahan, A., Roalf, D. R., Robert-Fitzgerald, T. E., Gonzenbach, V., Gur, R. C., Gur, R. E., Vandekar, S., Detre, J. A., Linn, K. A., Alexander-Bloch, A., Satterthwaite, T. D., and Shinohara, R. T · 2022
Later among the works it cites.
EgoTaskQA: Understanding Human Tasks in Egocentric Videos
Jia, B., Lei, T., Zhu, S.-C., and Huang, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering
Dancette, C., Cadène, R., Teney, D., and Cord, M · 2021
Cited alongside, same era.
Perceptual Score: What Data Modalities Does Your Model Perceive?
Gat, I., Schwartz, I., and Schwing, A. G · 2021
Cited alongside, same era.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., III, H. D., and Crawford, K · 2021
Cited alongside, same era.
Loss Re-Scaling VQA: Revisiting the Language Prior Problem From a Class-Imbalance View
Guo, Y., Nie, L., Cheng, Z., Tian, Q., and Zhang, M · 2021
Cited alongside, same era.
What Makes Multi-Modal Learning Better than Single (Provably)
Huang, Y., Du, C., Xue, Z., Chen, X., Zhao, H., and Huang, L · 2021
Cited alongside, same era.
Mitigating dataset harms requires stewardship: Lessons from 1000 papers
Peng, K., Mathur, A., and Narayanan, A · 2021
Cited alongside, same era.
A deeper dive into what deep spatiotemporal networks encode: Quantifying static vs. dynamic information
Kowal, M., Siam, M., Islam, M. A., Bruce, N. D. B., Wildes, R. P., and Derpanis, K. G · 2022
Later among the works it cites.
When classifying grammatical role, BERT doesn’t care about word order… except when it matters
Papadimitriou, I., Futrell, R., and Mahowald, K · 2022
Later among the works it cites.
Beyond Additive Fusion: Learning Non-Additive Multimodal Interactions
Wörtwein, T., Sheeber, L., Allen, N., Cohn, J., and Morency, L.-P · 2022
Later among the works it cites.
Video Graph Transformer for Video Question Answering
Xiao, J., Zhou, P., Chua, T.-S., and Yan, S · 2022
Later among the works it cites.
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2022
Later among the works it cites.
Video Question Answering: Datasets, Algorithms and Challenges
Zhong, Y., Ji, W., Xiao, J., Li, Y., Deng, W., and Chua, T.-S · 2022
Later among the works it cites.
Test of Time: Instilling Video-Language Models with a Sense of Time
Bagad, P., Tapaswi, M., and Snoek, C. G · 2023
Closest in time.
Token Merging: Your ViT But Faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Closest in time.
HiFi: High-information attention heads hold for parameter-efficient model adaptation
Gui, A. and Xiao, H · 2023
Closest in time.
Watching the news: Towards videoqa models that can read
Jahagirdar, S., Mathew, M., Karatzas, D., and Jawahar, C · 2023
Closest in time.
Revealing Single Frame Bias for Video-and-Language Learning
Lei, J., Berg, T., and Bansal, M · 2023
Closest in time.
Modality Coupling for Privacy Image Classification
Liu, Y., Huang, Y., Wang, S., Lu, W., and Wu, H · 2023
Closest in time.
Verbs in action: Improving verb understanding in video-language models
Momeni, L., Caron, M., Nagrani, A., Zisserman, A., and Schmid, C · 2023
Closest in time.
Beyond Distribution Shift: Spurious Features Through the Lens of Training Dynamics
Murali, N., Puli, A. M., Yu, K., Ranganath, R., and kayhan Batmanghelich · 2023
Closest in time.
Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Pătrăucean, V., Smaira, L., Gupta, A., Continente, A. R., Markeeva, L., Banarse, D., Koppula, S., Heyward, J., Malinowski, M., Yang, Y., Doersch, C., Matejovicova, T., Sulsky, Y., Miech, A., Frechette, A., Klimczak, H., Koster, R., Zhang, J., Winkler, S., Aytar, Y., Osindero, S., Damen, D., Zisserman, A., and Carreira, J · 2023
Closest in time.
Paxion: Patching Action Knowledge in Video-Language Foundation Models
Wang, Z., Blume, A., Li, S., Liu, G., Cho, J., Tang, Z., Bansal, M., and Ji, H · 2023
Closest in time.
When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?
Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J · 2023
Closest in time.
Foundations & Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions
Liang, P. P., Zadeh, A., and Morency, L.-P · 2024
Closest in time.