Fetching the paper…
Reading the bibliography…
The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Detecting activities of daily living in first-person camera views”
Hamed Pirsiavash and Deva Ramanan · 2012
Earlier work this paper cites.
“Delving into egocentric actions”
Yin Li, Zhefan Ye and James Rehg · 2015
Earlier work this paper cites.
“Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning”
William Lotter, Gabriel Kreiman and David Cox · 2016
Earlier work this paper cites.
“Next-active-object prediction from egocentric videos”
Antonino Furnari, Sebastiano Battiato, Kristen Grauman and Giovanni Farinella · 2017
Earlier work this paper cites.
“Decomposing motion and content for natural video sequence prediction”
Ruben Villegas et al · 2017
Earlier work this paper cites.
“Charades-ego: A large-scale dataset of paired third and first person videos”
Gunnar Sigurdsson et al · 2018
Earlier work this paper cites.
“EgoVQA-an egocentric video question answering benchmark dataset”
Chenyou Fan · 2019
Earlier work this paper cites.
“Zero-shot anticipation for instructional activities”
Fadime Sener and Angela Yao · 2019
Earlier work this paper cites.
“Cross-task weakly supervised learning from instructional videos”
Dimitri Zhukov et al · 2019
Earlier work this paper cites.
“Coin: A large-scale dataset for comprehensive instructional video analysis”
Yansong Tang et al · 2019
Earlier work this paper cites.
“BERTScore: Evaluating Text Generation with BERT”
Tianyi Zhang et al · 2019
Earlier work this paper cites.
“Procedure planning in instructional videos”
Chien-Yi Chang et al · 2020
Earlier work this paper cites.
“Long-term anticipation of activities with cycle consistency”
Yazan Abu, Qiuhong Ke, Bernt Schiele and Juergen Gall · 2021
Earlier work this paper cites.
“Truthfulqa: Measuring how models mimic human falsehoods”
Stephanie Lin, Jacob Hilton and Owain Evans · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“Egotaskqa: Understanding human tasks in egocentric videos”
Baoxiong Jia, Ting Lei, Song-Chun Zhu and Siyuan Huang · 2022
Earlier work this paper cites.
“Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100”
Dima Damen et al · 2022
Earlier work this paper cites.
“Ego4d: Around the world in 3,000 hours of egocentric video”
Kristen Grauman et al · 2022
Earlier work this paper cites.
“Scaling instruction-finetuned language models”
Hyung Chung et al · 2022
Earlier work this paper cites.
“Joint hand motion and interaction hotspots prediction from egocentric videos”
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar and Xiaolong Wang · 2022
Earlier work this paper cites.
“Egocentric video-language pretraining”
Kevin Lin et al · 2022
Earlier work this paper cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Earlier work this paper cites.
“Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality”
Wei-Lin Chiang et al · 2023
Earlier work this paper cites.
“Artificial general intelligence is already here”
Blaiseüera y Arcas and Peter Norvig · 2023
Cited alongside, same era.
“Sparks of artificial general intelligence: Early experiments with gpt-4”
Sébastien Bubeck et al · 2023
Cited alongside, same era.
“Levels of AGI: Operationalizing Progress on the Path to AGI”
Meredith Morris et al · 2023
Cited alongside, same era.
“Seed-bench-2: Benchmarking multimodal large language models”
Bohao Li et al · 2023
Cited alongside, same era.
URL: https://api.semanticscholar.org/CorpusID:263218031
“GPT-4V(ision) System Card”, 2023 · 2023
Cited alongside, same era.
“Structured world models from human videos”
Russell Mendonca, Shikhar Bahl and Deepak Pathak · 2023
Closest in time.
“Pretrained language models as visual planners for human assistance”
Dhruvesh Patel et al · 2023
Closest in time.
“LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning”
Bolin Lai et al · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Do as i can, not as i say: Grounding language in robotic affordances”
Anthony Brohan et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gemini Team et al · 2023
Cited alongside, same era.
“Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models”
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Cited alongside, same era.
“Instructblip: Towards general-purpose vision-language models with instruction tuning”
Wenliang Dai et al · 2023
Cited alongside, same era.
“MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”
Deyao Zhu et al · 2023
Cited alongside, same era.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Cited alongside, same era.
“mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration”, 2023
Qinghao Ye et al · 2023
Cited alongside, same era.
Ao Zhang et al · 2023
Closest in time.
“MultiModal-GPT: A Vision and Language Model for Dialogue with Humans”, 2023
Tao Gong et al · 2023
Closest in time.
“Otter: A Multi-Modal Model with In-Context Instruction Tuning”
Bo Li et al · 2023
Closest in time.
“OpenFlamingo”, https://github.com/mlfoundations/open_flamingo, 2023
ml · 2023
Closest in time.
“LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model”
Peng Gao et al · 2023
Closest in time.
“What Makes for Good Visual Tokenizers for Large Language Models?”
Guangzhi Wang et al · 2023
Closest in time.
“mplug-owl: Modularization empowers large language models with multimodality”
Qinghao Ye et al · 2023
Closest in time.
“Kosmos-2: Grounding Multimodal Large Language Models to the World”
Zhiliang Peng et al · 2023
Closest in time.
“Qwen-vl: A frontier large vision-language model with versatile abilities”
Jinze Bai et al · 2023
Closest in time.
“Valley: Video Assistant with Large Language model Enhanced abilitY”
Ruipu Luo et al · 2023
Closest in time.
“Cogvlm: Visual expert for pretrained language models”
Weihan Wang et al · 2023
Closest in time.
Pan Zhang et al · 2023
Closest in time.
“RRHF: Rank Responses to Align Language Models with Human Feedback”
Hongyi Yuan et al · 2023
Closest in time.
“Deepseek-vl: towards real-world vision-language understanding”
Haoyu Lu et al · 2024
Closest in time.
“Yi: Open foundation models by 01. ai”
Alex Young et al · 2024
Closest in time.
“Seed-x: Multimodal models with unified multi-granularity comprehension and generation”
Yuying Ge et al · 2024
Closest in time.
“Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark”
Zhenfei Yin et al · 2024
Closest in time.
Kaining Ying et al · 2024
Closest in time.
“Ego4d goal-step: Toward hierarchical understanding of procedural activities”
Yale Song et al · 2024
Closest in time.