Fetching the paper…
Reading the bibliography…
Advances in ML have motivated the design of video analytics systems that allow for structured queries over video datasets.
The use of MMR, diversity-based reranking for reordering documents and producing summaries
Jaime Carbonell and Jade Goldstein. 1998 · 1998
Earlier work this paper cites.
Evaluating top-k queries over web-accessible databases
Amélie Marian, Nicolas Bruno, and Luis Gravano. 2004 · 2004
Earlier work this paper cites.
Learning pairwise similarity for data clustering. In 18th International Conference on Pattern Recognition (ICPR’06) . IEEE
Ana LN Fred and Anil K Jain. 2006 · 2006
Earlier work this paper cites.
Analyzing who and what appears in a decade of US cable TV news
James Hong, Will Crichton, Haotian Zhang, Daniel Y Fu, Jacob Ritchie, Jeremy Barenholtz, Ben Hannel, Xinwei Yao, Michaela Murray, Geraldine Moriba, et al · 2008
Earlier work this paper cites.
A survey of top-k query processing techniques in relational database systems
Ihab F Ilyas, George Beskales, and Mohamed A Soliman. 2008 · 2008
Earlier work this paper cites.
TopX: efficient and versatile top-k query processing for semistructured data
Martin Theobald, Holger Bast, Debapriyo Majumdar, Ralf Schenkel, and Gerhard Weikum. 2008 · 2008
Earlier work this paper cites.
Spark: Cluster computing with working sets.. In HotCloud
Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica. 2010 · 2010
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV . Springer
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
NoScope: optimizing neural network queries over video at scale. In PVLDB
Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Recognition in Terra Incognita
Sara Beery, Grant Van Horn, and Pietro Perona. 2018 · 2018
Earlier work this paper cites.
Accelerating Machine Learning Inference with Probabilistic Predicates. In SIGMOD
Yao Lu, Aakanksha Chowdhery, Srikanth Kandula, and Surajit Chaudhuri. 2018 · 2018
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick. 2019 · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Earlier work this paper cites.
BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics. In PVLDB
Daniel Kang, Peter Bailis, and Matei Zaharia. 2019 · 2019
Earlier work this paper cites.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. 2020 · 2020
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. 2021 · 2021
Earlier work this paper cites.
Analysis of Faces in a Decade of US Cable TV News. In SIGKDD
James Hong, Will Crichton, Haotian Zhang, Daniel Y. Fu, Jacob Ritchie, Jeremy Barenholtz, Ben Hannel, Xinwei Yao, Michaela Murray, Geraldine Moriba, Maneesh Agrawala, and Kayvon Fatahalian. 2021 · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning . PMLR, 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Cited alongside, same era.
VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation. In 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks
Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al · 2021
Cited alongside, same era.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. 2021b · 2021
Cited alongside, same era.
3D ResNet Video Classification in PyTorch
PyTorch. 2022 · 2022
Later among the works it cites.
Optimizing Video Analytics with Declarative Model Relationships
Francisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang, Matei Zaharia, and Christos Kozyrakis. 2022 · 2022
Later among the works it cites.
PLAID: an efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 1747–1756
Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022 · 2022
Later among the works it cites.
Yolov5 Object Detection in PyTorch
Ultralytics. 2022 · 2022
Later among the works it cites.
EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views. In SIGMOD
Zhuangdi Xu, Gaurav Kakker, Joy Arulraj, and Umakishore Ramachandran. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021 · 2021
Cited alongside, same era.
Laion-400-Million Open Dataset
Christoph Schuhmann. 2021 · 2021
Cited alongside, same era.
TorchMetrics v0.3 — Information Retrieval metrics and more
PyTorch Lightning Team. 2021 · 2021
Cited alongside, same era.
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021 · 2021
Cited alongside, same era.
FILIP: fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021 · 2021
Cited alongside, same era.
CLIP Prompts for Image Classification
2022 · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Zhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu, Chen Li, and X. Sean Wang. 2022 · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022c · 2022
Later among the works it cites.
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra. 2022a · 2022
Later among the works it cites.
Boggart: Towards General-Purpose Acceleration of Retrospective Video Analytics. In NSDI
Neil Agarwal and Ravi Netravali. 2023 · 2023
Closest in time.
VOCALExplore: Pay-as-You-Go Video Data Exploration and Model Building
Maureen Daum, Enhao Zhang, Dong He, Stephen Mussmann, Brandon Haynes, Ranjay Krishna, and Magdalena Balazinska. 2023 · 2023
Closest in time.
The Full Guide to Video Annotation for Computer Vision
Encord. 2023 · 2023
Closest in time.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. 2023 · 2023
Closest in time.
A Dive into Vision-Language Models
HuggingFace. 2023 · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Closest in time.
NVIDIA T4
NVIDIA. 2023 · 2023
Closest in time.
A Dive into Vision-Language Models
OpenAI. 2023 · 2023
Closest in time.
Free Movies
1091 Pictures. 2023 · 2023
Closest in time.
Pinecone: Vector Database for Vector Search
PineconeDB. 2023 · 2023
Closest in time.
Querying Large Language Models with SQL
Mohammed Saeed, Nicola De Cao, and Paolo Papotti. 2023 · 2023
Closest in time.
Laion-400-Million Open Dataset: KNN Front-end
Christoph Schuhmann. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.