Fetching the paper…
Reading the bibliography…
This paper explores the potential of Large Language Models(LLMs) in zero-shot anomaly detection for safe visual navigation.
Abnormal crowd behavior detection using social force model
Ramin Mehran, Alexis Oyama, and Mubarak Shah · 2009
Earlier work this paper cites.
The user as a sensor: navigating users with visual impairments in indoor spaces using tactile landmarks
Navid Fallah, Ilias Apostolopoulos, Kostas Bekris, and Eelke Folmer · 2012
Earlier work this paper cites.
Incremental learning of 3d-dct compact representations for robust visual tracking
Xi Li, Anthony Dick, Chunhua Shen, Anton Van Den Hengel, and Hanzi Wang · 2012
Earlier work this paper cites.
Mobile assistive technologies for the visually impaired
Lilit Hakobyan, Jo Lumsden, Dympna O’Sullivan, and Hannah Bartlett · 2013
Earlier work this paper cites.
Object recognition in a mobile phone application for visually impaired users
Karol Matusiak, Piotr Skulimowski, and P Strurniłło · 2013
Earlier work this paper cites.
After access: Inclusion, development, and a more mobile Internet
Jonathan Donner · 2015
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
People with visual impairment training personal object recognizers: Feasibility and challenges
Hernisa Kacorri, Kris M Kitani, Jeffrey P Bigham, and Chieko Asakawa · 2017
Earlier work this paper cites.
Vision-based mobile indoor assistive navigation aid for blind people
Bing Li, Juan Pablo Munoz, Xuejian Rong, Qingtian Chen, Jizhong Xiao, Yingli Tian, Aries Arditi, and Mohammed Yousuf · 2018
Earlier work this paper cites.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Earlier work this paper cites.
Wearable smart system for visually impaired people
Ali Jasim Ramadhan · 2018
Earlier work this paper cites.
Zero-shot text classification with generative language models
Raul Puri and Bryan Catanzaro · 2019
Earlier work this paper cites.
Smart technologies for visually impaired: Assisting and conquering infirmity of blind people using ai technologies
Fatma Al-Muqbali, Noura Al-Tourshi, Khuloud Al-Kiyumi, and Faizal Hajmohideen · 2020
Earlier work this paper cites.
An ai-based visual aid with integrated reading assistant for the completely blind
Muiz Ahmed Khan, Pias Paul, Mahmudur Rashid, Mainul Hossain, and Md Atiqur Rahman Ahad · 2020
Earlier work this paper cites.
An evaluation of retinanet on indoor object detection for blind and visually impaired persons assistance navigation
Mouna Afif, Riadh Ayachi, Yahia Said, Edwige Pissaloux, and Mohamed Atri · 2020
Earlier work this paper cites.
Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection
Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
On extractive and abstractive neural document summarization with transformer language models
Jonathan Pilault, Raymond Li, Sandeep Subramanian, and Christopher Pal · 2020
Earlier work this paper cites.
Efficient multi-object detection and smart navigation using artificial intelligence for visually impaired people
Rakesh Chandra Joshi, Saumya Yadav, Malay Kishore Dutta, and Carlos M Travieso-Gonzalez · 2020
Earlier work this paper cites.
Improved visual-semantic alignment for zero-shot object detection
Shafin Rahman, Salman Khan, and Nick Barnes · 2020
Earlier work this paper cites.
Object detection and recognition: using deep learning to assist the visually impaired
Abinash Bhandari, PWC Prasad, Abeer Alsadoon, and Angelika Maag · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Open-vocabulary object detection using captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang · 2021
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
Zsd-yolo: Zero-shot yolo detection using vision-language knowledgedistillation
Johnathan Xie and Shuai Zheng · 2021
Cited alongside, same era.
Cnn-based object recognition and tracking system to assist visually impaired people
Fahad Ashiq, Muhammad Asif, Maaz Bin Ahmad, Sadia Zafar, Khalid Masood, Toqeer Mahmood, Muhammad Tariq Mahmood, and Ik Hyun Lee · 2022
Cited alongside, same era.
Image segmentation using text and image prompts
Timo Lüddecke and Alexander Ecker · 2022
Cited alongside, same era.
Regionclip: Region-based language-image pretraining
Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al · 2022
Cited alongside, same era.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra · 2022
Cited alongside, same era.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, et al · 2023
Later among the works it cites.
Langnav: Language as a perceptual representation for navigation
Bowen Pan, Rameswar Panda, SouYoung Jin, Rogerio Feris, Aude Oliva, Phillip Isola, and Yoon Kim · 2023
Later among the works it cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, Sergey Levine, et al · 2023
Later among the works it cites.
Aligning bag of regions for open-vocabulary object detection
Size Wu, Wenwei Zhang, Sheng Jin, Wentao Liu, and Chen Change Loy · 2023
Later among the works it cites.
Video owl-vit: Temporally-consistent open-world localization in video
Georg Heigold, Matthias Minderer, Alexey Gritsenko, Alex Bewley, Daniel Keysers, Mario Lučić, Fisher Yu, and Thomas Kipf · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grounded language-image pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al · 2022
Cited alongside, same era.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Measuring and narrowing the compositionality gap in language models
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis · 2022
Cited alongside, same era.
Anomaly detection for visually impaired people using a 360 degree wearable camera
Dong-in Kim and Jangwon Lee · 2022
Cited alongside, same era.
Multi-functional glasses for the blind and visually impaired: Design and development
Ashish Bastola, Md Atik Enam, Ananta Bastola, Aaron Gluck, and Julian Brinkley · 2023
Cited alongside, same era.
Chatgpt for visually impaired and blind
Askat Kuzdeuov, Shakhizat Nurgaliyev, and Hüseyin Atakan Varol · 2023
Cited alongside, same era.
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Later among the works it cites.
A systematic survey of prompt engineering on vision-language foundation models
Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr · 2023
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Later among the works it cites.
Prompt engineering for healthcare: Methodologies and applications
Jiaqi Wang, Enze Shi, Sigang Yu, Zihao Wu, Chong Ma, Haixing Dai, Qiushi Yang, Yanqing Kang, Jinru Wu, Huawen Hu, et al · 2023
Later among the works it cites.
Ashish Bastola, Hao Wang, Judsen Hembree, Pooja Yadav, Nathan McNeese, and Abolfazl Razi · 2023
Later among the works it cites.
Better zero-shot reasoning with self-adaptive prompting
Xingchen Wan, Ruoxi Sun, Hanjun Dai, Sercan O Arik, and Tomas Pfister · 2023
Later among the works it cites.
Large language models in the workplace: A case study on prompt engineering for job type classification
Benjamin Clavié, Alexandru Ciceu, Frederick Naylor, Guillaume Soulié, and Thomas Brightwell · 2023
Later among the works it cites.
A dataset for the recognition of obstacles on blind sidewalk
Wu Tang, De-er Liu, Xiaoli Zhao, Zenghui Chen, and Chen Zhao · 2023
Later among the works it cites.
Driving towards inclusion: Revisiting in-vehicle interaction in autonomous vehicles
Ashish Bastola, Julian Brinkley, Hao Wang, and Abolfazl Razi · 2024
Closest in time.
Vialm: A survey and benchmark of visually impaired assistance with large models
Yi Zhao, Yilin Zhang, Rong Xiang, Jing Li, and Hillming Li · 2024
Closest in time.
Vision-language navigation: A survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, and Yue Hu · 2024
Closest in time.
Supervised abnormal event detection based on chatgpt attention mechanism
Feng Tian, Yuanyuan Lu, Fang Liu, Guibao Ma, Neili Zong, Xin Wang, Chao Liu, Ningbin Wei, and Kaiguang Cao · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Gpt-4 enhanced multimodal grounding for autonomous driving: Leveraging cross-modal attention with large language models
Haicheng Liao, Huanming Shen, Zhenning Li, Chengyue Wang, Guofa Li, Yiming Bie, and Chengzhong Xu · 2024
Closest in time.
Mapgpt: Map-guided prompting for unified vision-and-language navigation
Jiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai, Xiaodan Liang, and Kwan-Yee K Wong · 2024
Closest in time.
Navhint: Vision and language navigation agent with a hint generator
Yue Zhang, Quan Guo, and Parisa Kordjamshidi · 2024
Closest in time.
Is it safe to cross? interpretable risk assessment with gpt-4v for safety-aware street crossing
Hochul Hwang, Sunjae Kwon, Yekyung Kim, and Donghyun Kim · 2024
Closest in time.
Yolo-world: Real-time open-vocabulary object detection
Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan · 2024
Closest in time.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang · 2024
Closest in time.