Fetching the paper…
Reading the bibliography…
Visually Impaired Assistance (VIA) aims to automatically help the visually impaired (VI) handle daily activities.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, and et al · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, and et al · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, and et al · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
BrowseWithMe: An online clothes shopping assistant for people with visual impairments
Abigale J. Stangl, Esha Kothari, and et al · 2018
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, and et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, and et al · 2019
Earlier work this paper cites.
ReCog: Supporting blind people in recognizing personal objects
Dragan Ahmetovic, Daisuke Sato, and et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, and et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, and et al · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, and et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, and et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, and et al · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae, Sebastian Borgeaud, and et al · 2021
Earlier work this paper cites.
Deformable DETR: deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, and et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, and et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, and et al · 2022
Earlier work this paper cites.
GLaM: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, and et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, and et al · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, and et al · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, and et al · 2022
Earlier work this paper cites.
Do as I can, not as I say: Grounding language in robotic affordances
Brian Ichter, nthony Brohan, and et al · 2022
Earlier work this paper cites.
OPT-IML: scaling language model instruction meta learning through the lens of generalization
Srinivasan Iyer, Xi Victoria Lin, and et al · 2022
Earlier work this paper cites.
Cognitive map formation through tactile map navigation in visually impaired and sighted persons
Loes Ottink, Bram van Raalte, and et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, and et al · 2022
Earlier work this paper cites.
BLOOM: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, and et al · 2022
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, and et al · 2022
Cited alongside, same era.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, and et al · 2022
Cited alongside, same era.
LaMDA: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, and et al · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, and et al · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, and et al · 2022
Cited alongside, same era.
One shot learning as instruction data prospector for large language models, 2023
Yunshui Li, Binyuan Hui, and et al · 2023
Later among the works it cites.
Code as Policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, and et al · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, and et al · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, and et al · 2023
Later among the works it cites.
An overview of bard: an early experiment with generative ai
James Manyika · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, and et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, and et al · 2022
Cited alongside, same era.
OPT: open pre-trained transformer language models
Susan Zhang, Stephen Roller, and et al · 2022
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, and et al · 2023
Cited alongside, same era.
PaLM 2 technical report
Rohan Anil, Andrew M. Dai, and et al · 2023
Cited alongside, same era.
Qwen-VL: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, and et al · 2023
Cited alongside, same era.
RT-2: vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, and et al · 2023
Cited alongside, same era.
NLG evaluation metrics beyond correlation analysis: An empirical metric preference checklist
Iftitahu Ni’mah, Meng Fang, and et al · 2023
Later among the works it cites.
GPT-4 technical report
OpenAI · 2023
Later among the works it cites.
RWKV: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, and et al · 2023
Later among the works it cites.
ProgPrompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, and et al · 2023
Later among the works it cites.
EVA-CLIP: improved training techniques for CLIP at scale
Quan Sun, Yuxin Fang, and et al · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, and et al · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, and et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, and et al · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, and et al · 2023
Later among the works it cites.
CogVLM: Visual expert for pretrained language models
Weihan Wang, Qingsong Lv, and et al · 2023
Later among the works it cites.
VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks
Wenhai Wang, Zhe Chen, and et al · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, and et al · 2023
Later among the works it cites.
ReAct: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, and et al · 2023
Later among the works it cites.
GLM-130B: an open bilingual pre-trained model
Aohan Zeng, Xiao Liu, and et al · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, and et al · 2023
Later among the works it cites.
Chat with the environment: Interactive multimodal perception using large language models
Xufeng Zhao, Mengdi Li, and et al · 2023
Later among the works it cites.
Ghost in the Minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory
Xizhou Zhu, Yuntao Chen, and et al · 2023
Later among the works it cites.
Autort: Embodied foundation models for large scale orchestration of robotic agents, 2024
Michael Ahn, Debidatta Dwibedi, and et al · 2024
Closest in time.
BLIVA: A simple multimodal LLM for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, and et al · 2024
Closest in time.
Mixtral of experts
Albert Q. Jiang, Alexandre Sablayrolles, and et al · 2024
Closest in time.