Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos.
Facial action coding system
Paul Ekman and Wallace V Friesen · 1978
Earlier work this paper cites.
A novel algorithm for remote photoplethysmography: Spatial subspace rotation
Wenjin Wang, Sander Stuijk, and Gerard De Haan · 1984
Earlier work this paper cites.
Disfa: A spontaneous facial action intensity database
S Mohammad Mavadati, Mohammad H Mahoor, Kevin Bartlett, Philip Trinh, and Jeffrey F Cohn · 2013
Earlier work this paper cites.
Compound facial expressions of emotion
Shichuan Du, Yong Tao, and Aleix M Martinez · 2014
Earlier work this paper cites.
Lbp with six intersection points: Reducing redundant information in lbp-top for micro-expression recognition
Yandan Wang, John See, Raphael C-W Phan, and Yee-Hui Oh · 2014
Earlier work this paper cites.
Casme ii: An improved spontaneous micro-expression database and the baseline evaluation
Wen-Jing Yan, Xiaobai Li, Su-Jing Wang, Guoying Zhao, Yong-Jin Liu, Yu-Hsin Chen, and Xiaolan Fu · 2014
Earlier work this paper cites.
Algorithmic principles of remote ppg
Wenjin Wang, Albertus C Den Brinker, Sander Stuijk, and Gerard De Haan · 2016
Earlier work this paper cites.
Deep region and multi-label learning for facial action unit detection
Kaili Zhao, Wen-Sheng Chu, and Honggang Zhang · 2016
Earlier work this paper cites.
A review of affective computing: From unimodal analysis to multimodal fusion
Soujanya Poria, Erik Cambria, Rajiv Bajpai, and Amir Hussain · 2017
Earlier work this paper cites.
Deepphys: Video-based physiological measurement using convolutional attention networks
Weixuan Chen and Daniel McDuff · 2018
Earlier work this paper cites.
Deep structure inference network for facial action unit recognition
Ciprian Corneanu, Meysam Madadi, and Sergio Escalera · 2018
Earlier work this paper cites.
Eac-net: Deep nets with enhancing and cropping for facial action unit detection
Wei Li, Farnaz Abtahi, Zhigang Zhu, and Lijun Yin · 2018
Earlier work this paper cites.
Reliable crowdsourcing and deep locality-preserving learning for unconstrained facial expression recognition
Li Shan and Weihong Deng · 2018
Earlier work this paper cites.
Deep adaptive attention for joint facial action unit detection and face alignment
Zhiwen Shao, Zhilei Liu, Jianfei Cai, and Lizhuang Ma · 2018
Earlier work this paper cites.
Semantic relationships guided representation learning for facial action unit recognition
Guanbin Li, Xin Zhu, Yirui Zeng, Qing Wang, and Liang Lin · 2019
Earlier work this paper cites.
Facial action unit detection using attention and relation learning
Zhiwen Shao, Zhilei Liu, Jianfei Cai, Yunsheng Wu, and Lizhuang Ma · 2019
Cited alongside, same era.
Multimodal deception detection using real-life trial data
M Umut Şen, Veronica Perez-Rosas, Berrin Yanikoglu, Mohamed Abouelenien, Mihai Burzo, and Rada Mihalcea · 2020
Cited alongside, same era.
Indonlu: Benchmark and resources for evaluating indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, et al · 2020
Cited alongside, same era.
Clue: A chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, et al · 2020
Cited alongside, same era.
Facial action unit detection with transformers
Geethu Miriam Jacob and Bjorn Stenger · 2021
Audio-visual deception detection: Dolos dataset and parameter-efficient crossmodal learning
Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong, Bingquan Shen, and Alex Kot · 2023
Later among the works it cites.
Challenges and prospects of visual contactless physiological monitoring in clinical study
Bin Huang, Shen Hu, Zimeng Liu, Chun-Liang Lin, Junfeng Su, Changchen Zhao, Li Wang, and Wenjin Wang · 2023
Later among the works it cites.
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao · 2023
Later among the works it cites.
Multi-scale promoted self-adjusting correlation learning for facial action unit detection
Xin Liu, Kaishen Yuan, Xuesong Niu, Jingang Shi, Zitong Yu, Huanjing Yue, and Jingyu Yang · 2023
Later among the works it cites.
Neuron structure modeling for generalizable remote physiological measurement
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Micro-expression action unit detection with dual-view attentive similarity-preserving knowledge distillation
Yante Li, Wei Peng, and Guoying Zhao · 2021
Cited alongside, same era.
imigue: An identity-free video dataset for micro-gesture understanding and emotion analysis
Xin Liu, Henglin Shi, Haoyu Chen, Zitong Yu, Xiaobai Li, and Guoying Zhao · 2021
Cited alongside, same era.
Dual-gan: Joint bvp and noise modeling for remote physiological measurement
Hao Lu, Hu Han, and S Kevin Zhou · 2021
Cited alongside, same era.
Piap-df: Pixel-interested and anti person-specific facial action unit detection net with discrete feedback learning
Yang Tang, Wangding Zeng, Dafei Zhao, and Honggang Zhang · 2021
Cited alongside, same era.
Facial-video-based physiological signal measurement: Recent advances and affective applications
Zitong Yu, Xiaobai Li, and Guoying Zhao · 2021
Cited alongside, same era.
Learning multi-dimensional edge feature-based au relation graph for facial action unit recognition
Cheng Luo, Siyang Song, Weicheng Xie, Linlin Shen, and Hatice Gunes · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Hao Lu, Zitong Yu, Xuesong Niu, and Ying-Cong Chen · 2023
Later among the works it cites.
Valley: Video assistant with large language model enhanced ability
Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Minghui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving
Licheng Wen, Xuemeng Yang, Daocheng Fu, Xiaofeng Wang, Pinlong Cai, Xin Li, Tao Ma, Yingxuan Li, Linran Xu, Dengke Shang, et al · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al · 2023
Later among the works it cites.
Physformer++: Facial video-based physiological measurement with slowfast temporal difference transformer
Zitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao, Yawen Cui, Jiehua Zhang, Philip Torr, and Guoying Zhao · 2023
Later among the works it cites.
Facial micro-expressions: An overview
Guoying Zhao, Xiaobai Li, Yante Li, and Matti Pietikäinen · 2023
Later among the works it cites.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2024
Closest in time.
rppg-mae: Self-supervised pretraining with masked autoencoders for remote physiological measurements
Xin Liu, Yuting Zhang, Zitong Yu, Hao Lu, Huanjing Yue, and Jingyu Yang · 2024
Closest in time.
Video anomaly detection and explanation via large language models
Hui Lv and Qianru Sun · 2024
Closest in time.