Fetching the paper…
Reading the bibliography…
Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a diverse variety of tasks and domains.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019 · 1910
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Exploring Contrastive Learning in Human Activity Recognition for Healthcare
Tang, C. I.; Perez-Pozuelo, I.; Spathis, D.; and Mascolo, C. 2020 · 2011
Earlier work this paper cites.
Introducing a new benchmarked dataset for activity monitoring
Reiss, A.; and Stricker, D. 2012 · 2012
Earlier work this paper cites.
mHealthDroid: a novel framework for agile development of mobile health applications
Banos, O.; Garcia, R.; Holgado-Terriza, J. A.; Damas, M.; Pomares, H.; Rojas, I.; Saez, A.; and Villalonga, C. 2014 · 2014
Earlier work this paper cites.
Smart devices are different: Assessing and mitigatingmobile sensing heterogeneities for activity recognition
Stisen, A.; Blunck, H.; Bhattacharya, S.; Prentow, T. S.; Kjærgaard, M. B.; Dey, A.; Sonne, T.; and Jensen, M. M. 2015 · 2015
Earlier work this paper cites.
Human daily activity and fall recognition using a smartphone’s acceleration sensor
Chatzaki, C.; Pediaditis, M.; Vavoulas, G.; and Tsiknakis, M. 2016 · 2016
Earlier work this paper cites.
Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition
Ordóñez, F. J.; and Roggen, D. 2016 · 2016
Earlier work this paper cites.
On-body localization of wearable devices: An investigation of position-aware activity recognition
Sztyler, T.; and Stuckenschmidt, H. 2016 · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Thomee, B.; Shamma, D. A.; Friedland, G.; Elizalde, B.; Ni, K.; Poland, D.; Borth, D.; and Li, L.-J. 2016 · 2016
Earlier work this paper cites.
Myogym: introducing an open gym data set for activity recognition collected using myo armband
Koskimäki, H.; Siirtola, P.; and Röning, J. 2017 · 2017
Earlier work this paper cites.
Data augmentation of wearable sensor data for parkinson’s disease monitoring using convolutional neural networks
Um, T. T.; Pfister, F. M.; Pichler, D.; Endo, S.; Lang, M.; Hirche, S.; Fietzek, U.; and Kulić, D. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Protecting sensory data against sensitive inferences
Malekzadeh, M.; Clegg, R. G.; Cavallaro, A.; and Haddadi, H. 2018 · 2018
Earlier work this paper cites.
Statistical machine learning of sleep and physical activity phenotypes from sensor data in 96,220 UK Biobank participants
Willetts, M.; Hollowell, S.; Aslett, L.; Holmes, C.; and Doherty, A. 2018 · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Cited alongside, same era.
Multi-task self-supervised learning for human activity detection
Saeed, A.; Ozcelebi, T.; and Lukkien, J. 2019 · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Cited alongside, same era.
Testing self-report time-use diaries against objective instruments in real time
Gershuny, J.; Harms, T.; Doherty, A.; Thomas, E.; Milton, K.; Kelly, P.; and Foster, C. 2020 · 2020
Cited alongside, same era.
Imutube: Automatic extraction of virtual on-body accelerometry from video for human activity recognition
Kwon, H.; Tong, C.; Haresamudram, H.; Gao, Y.; Abowd, G. D.; Lane, N. D.; and Ploetz, T. 2020 · 2020
Cited alongside, same era.
Imu2clip: Multimodal contrastive learning for imu motion sensors from egocentric videos and text
Moon, S.; Madotto, A.; Lin, Z.; Dirafzoon, A.; Saraf, A.; Bearman, A.; and Damavandi, B. 2022 · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N.; Kirillov, A.; Wagner, D.; and Xie, S. 2022 · 2022
Later among the works it cites.
K-lite: Learning transferable visual models with external knowledge
Shen, S.; Li, C.; Hu, X.; Xie, Y.; Yang, J.; Zhang, P.; Gan, Z.; Wang, L.; Yuan, L.; Liu, C.; et al. 2022 · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
Wu, H.-H.; Seetharaman, P.; Kumar, K.; and Bello, J. P. 2022 · 2022
Later among the works it cites.
Unified contrastive learning in image-text-label space
Yang, J.; Li, C.; Zhang, P.; Xiao, B.; Liu, C.; Yuan, L.; and Gao, J. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Capture-24: Activity tracker dataset for human activity recognition
Chan Chang, S.; and Doherty, A. 2021 · 2021
Cited alongside, same era.
Applying machine learning for sensor data analysis in interactive systems: Common pitfalls of pragmatic use and ways to avoid them
PlÖtz, T. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Zero-shot learning for imu-based activity recognition using video embeddings
Tong, C.; Ge, J.; and Lane, N. D. 2021 · 2021
Cited alongside, same era.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Xu, H.; Ghosh, G.; Huang, P.-Y.; Okhonko, D.; Aghajanyan, A.; Metze, F.; Zettlemoyer, L.; and Feichtenhofer, C. 2021 · 2021
Cited alongside, same era.
A CLIP-Hitchhiker’s Guide to Long Video Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2022 · 2022
Cited alongside, same era.
A personalized approach for developing a snacking detection system using earbuds in a semi-naturalistic setting
Bin Morshed, M.; Haresamudram, H. K.; Bandaru, D.; Abowd, G. D.; and Plötz, T. 2022 · 2022
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Elizalde, B.; Deshmukh, S.; Al Ismail, M.; and Wang, H. 2023 · 2023
Later among the works it cites.
Improving CLIP Training with Language Rewrites
Fan, L.; Krishnan, D.; Isola, P.; Katabi, D.; and Tian, Y. 2023 · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K. V.; Joulin, A.; and Misra, I. 2023 · 2023
Later among the works it cites.
Investigating enhancements to contrastive predictive coding for human activity recognition
Haresamudram, H.; Essa, I.; and Plötz, T. 2023 · 2023
Later among the works it cites.
Anymal: An efficient and scalable any-modality augmented language model
Moon, S.; Madotto, A.; Lin, Z.; Nagarajan, T.; Smith, M.; Jain, S.; Yeh, C.-F.; Murugesan, P.; Heidari, P.; Liu, Y.; et al. 2023 · 2023
Later among the works it cites.
If only we had more data!: Sensor-Based Human Activity Recognition in Challenging Scenarios
Plötz, T. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Later among the works it cites.
Learning Video Representations from Large Language Models
Zhao, Y.; Misra, I.; Krähenbühl, P.; and Girdhar, R. 2023 · 2023
Later among the works it cites.
Tent: Connect language models with iot sensors for zero-shot activity recognition
Zhou, Y.; Yang, J.; Zou, H.; and Xie, L. 2023 · 2023
Later among the works it cites.
Verma, G.; Choi, M.; Sharma, K.; Watson-Daniels, J.; Oh, S.; and Kumar, S. 2024 · 2024
Closest in time.
TS2ACT: Few-Shot Human Activity Sensing with Cross-Modal Co-Learning
Xia, K.; Li, W.; Gan, S.; and Lu, S. 2024 · 2024
Closest in time.