Fetching the paper…
Reading the bibliography…
We present IMU2CLIP, a novel pre-training approach to align Inertial Measurement Unit (IMU) motion sensor recordings with video and text, by projecting them into the joint representation space of Contrastive Language-Image Pre-training (CLIP).
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2011
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“A gaussian process regression model for walking speed estimation using a head-worn imu,”
Shaghayegh Zihajehzadeh and Edward J Park, · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al., · 2018
Earlier work this paper cites.
“Sentence-bert: Sentence embeddings using siamese bert-networks,”
Nils Reimers and Iryna Gurevych, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Earlier work this paper cites.
“Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping,”
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith, · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Cited alongside, same era.
“Gpt-3: Its nature, scope, limits, and consequences,”
Luciano Floridi and Massimo Chiriatti, · 2020
Cited alongside, same era.
“Charm-deep: Continuous human activity recognition model based on deep neural network using imu sensors of smartwatch,”
Sara Ashry, Tetsuji Ogawa, and Walid Gomaa, · 2020
Cited alongside, same era.
“Epic-kitchens-100-2021 challenges report,” 2021
Dima Damen, Adriano Fragomeni, Jonathan Munro, Toby Perrett, Daniel Whettam, Michael Wray, Antonino Furnari, Giovanni Maria Farinella, and Davide Moltisanti, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Later among the works it cites.
“Wearable imu-based human activity recognition algorithm for clinical balance assessment using 1d-cnn and gru ensemble model,”
Yeon-Wook Kim, Kyung-Lim Joa, Han-Young Jeong, and Sangmin Lee, · 2021
Later among the works it cites.
“Ego4d: Around the world in 3,000 hours of egocentric video,”
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al., · 2022
Closest in time.
“Aria pilot dataset,” 2022
Zhaoyang Lv, Edward Miller, Jeff Meissner, Luis Pesqueira, Chris Sweeney, Jing Dong, Lingni Ma, Pratik Patel, Pierre Moulon, Kiran Somasundaram, Omkar Parkhi, Yuyang Zou, Nikhil Raina, Steve Saarinen, Yusuf M Mansour, Po-Kang Huang, Zijian Wang, Anton Troynikov, Raul Mur Artal, Daniel DeTone, Daniel Barnes, Elizabeth Argall, Andrey Lobanovskiy, David Jaeyun Kim, Philippe Bouttefroy, Julian Straub, Jakob Julian Engel, Prince Gupta, Mingfei Yan, Renzo De Nardi, and Richard Newcombe, · 2022
Closest in time.
“Multi-category gesture recognition modeling based on semg and imu signals,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A probability distribution model-based approach for foot placement prediction in the early swing phase with a wearable imu sensor,”
Xinxing Chen, Kuangen Zhang, Haiyuan Liu, Yuquan Leng, and Chenglong Fu, · 2021
Cited alongside, same era.
“CLIP4Clip: An empirical study of clip for end to end video clip retrieval,”
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
Yujian Jiang, Lin Song, Junming Zhang, Yang Song, and Ming Yan, · 2022
Closest in time.
“Egocentric video-language pretraining,”
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al., · 2022
Closest in time.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Closest in time.
“Socratic models: Composing zero-shot multimodal reasoning with language,”
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al., · 2022
Closest in time.