Fetching the paper…
Reading the bibliography…
We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Multispectral pedestrian detection: Benchmark dataset and baseline
Soonmin Hwang, Jaesik Park, Namil Kim, Yukyung Choi, and In So Kweon · 2015
Earlier work this paper cites.
Converting static image datasets to spiking neuromorphic datasets using saccades
Garrick Orchard, Ajinkya Jayawant, Gregory K Cohen, and Nitish Thakor · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Earlier work this paper cites.
Semantic-aware scene recognition
Alejandro López-Cifuentes, Marcos Escudero-Vinolo, Jesús Bescós, and Álvaro García-Martín · 2020
Earlier work this paper cites.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Earlier work this paper cites.
Kat: A knowledge augmented transformer for vision-and-language
Liangke Gui, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao · 2021
Earlier work this paper cites.
Knowledgeable prompt-tuning: Incorporating knowledge into prompt verbalizer for text classification
Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, and Maosong Sun · 2021
Earlier work this paper cites.
Llvip: A visible-infrared paired dataset for low-light vision
Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou · 2021
Cited alongside, same era.
N-imagenet: Towards robust, fine-grained object recognition with event cameras
Junho Kim, Jaehyeok Bae, Gangin Park, Dongsu Zhang, and Young Min Kim · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Cited alongside, same era.
A review on language models as knowledge bases
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad · 2022
Cited alongside, same era.
Learning visual representations via language-guided sampling
Mohamed El Banani, Karan Desai, and Justin Johnson · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al · 2023
Later among the works it cites.
Long Lian, Boyi Li, Adam Yala, and Trevor Darrell · 2023
Later among the works it cites.
Query rewriting for retrieval-augmented large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge base question answering by case-based reasoning over subgraphs
Rajarshi Das, Ameya Godbole, Ankita Naik, Elliot Tower, Manzil Zaheer, Hannaneh Hajishirzi, Robin Jia, and Andrew McCallum · 2022
Cited alongside, same era.
Multimodal sensors and ml-based data fusion for advanced robots
Shengshun Duan, Qiongfeng Shi, and Jun Wu · 2022
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Cited alongside, same era.
Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval
Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu, and Ge Yu · 2022
Cited alongside, same era.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Cited alongside, same era.
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan · 2023
Later among the works it cites.
Ave-clip: Audioclip-based multi-window temporal transformer for audio visual event localization
Tanvir Mahmud and Diana Marculescu · 2023
Later among the works it cites.
Open vocabulary semantic segmentation with patch aligned contrastive learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip HS Torr, and Ser-Nam Lim · 2023
Later among the works it cites.
I2mvformer: Large language model generated multi-view document supervision for zero-shot image classification
Muhammad Ferjad Naeem, Muhammad Gul Zain Ali Khan, Yongqin Xian, Muhammad Zeshan Afzal, Didier Stricker, Luc Van Gool, and Federico Tombari · 2023
Later among the works it cites.
Llm2kb: Constructing knowledge bases using instruction tuned context aware large language models
Anmol Nayak and Hari Prasad Timmapathini · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Pre-training multi-modal dense retrievers for outside-knowledge visual question answering
Alireza Salemi, Mahta Rafiee, and Hamed Zamani · 2023
Later among the works it cites.
Teler: A general taxonomy of llm prompts for benchmarking complex tasks
Shubhra Kanti Karmaker Santu and Dongji Feng · 2023
Later among the works it cites.
Recent advancements in multimodal human–robot interaction
Hang Su, Wen Qi, Jiahao Chen, Chenguang Yang, Juan Sandoval, and Med Amine Laribi · 2023
Later among the works it cites.
Learning audio-visual source localization via false negative aware contrastive learning
Weixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu, Yiran Zhong, Tianpeng Feng, Yandong Guo, Yanhao Zhang, and Nick Barnes · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
iclip: Bridging image classification and contrastive language-image pre-training for visual recognition
Yixuan Wei, Yue Cao, Zheng Zhang, Houwen Peng, Zhuliang Yao, Zhenda Xie, Han Hu, and Baining Guo · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
E-clip: Towards label-efficient event-based open-world understanding by clip
Jiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, and Lin Wang · 2023
Later among the works it cites.
Image anything: Towards reasoning-coherent and training-free multi-modal image generation
Yuanhuiyi Lyu, Xu Zheng, and Lin Wang · 2024
Closest in time.