Fetching the paper…
Reading the bibliography…
We introduce an open-source system called SIGMA (short for "Situated Interactive Guidance, Monitoring, and Assistance") as a platform for conducting research on task-assistive agents in mixed-reality scenarios.
Augmented reality: an application of heads-up display technology to manual manufacturing processes
T. Caudell and D. Mizell · 1992
Earlier work this paper cites.
Principles of mixed-initiative user interfaces
E. Horvitz · 1999
Earlier work this paper cites.
Comparative effectiveness of augmented reality in object assembly
A. Tang, C. Owen, F. Biocca, and W. Mou · 2003
Earlier work this paper cites.
Larri: A language-based maintenance and repair assistant
D. Bohus and A. I. Rudnicky · 2005
Earlier work this paper cites.
Diamondhelp: A generic collaborative task guidance system
C. Rich and C. L. Sidner · 2007
Earlier work this paper cites.
Exploring the benefits of augmented reality documentation for maintenance and repair
S. Henderson and S. Feiner · 2010
Earlier work this paper cites.
An augmented reality training platform for assembly and maintenance skills
S. Webel, U. Bockholt, T. Engelke, N. Gavish, M. Olbrich, and C. Preusche · 2013
Earlier work this paper cites.
Using in-situ projection to support cognitively impaired workers at the workplace
M. Funk, S. Mayer, and A. Schmidt · 2015
Earlier work this paper cites.
Eye-wearable technology for machine maintenance: Effects of display position and hands-free operation
X. S. Zheng, C. Foucault, P. Matos da Silva, S. Dasari, T. Yang, and S. Goose · 2015
Earlier work this paper cites.
Rapid development of multimodal interactive systems: a demonstration of platform for situated intelligence
D. Bohus, S. Andrist, and M. Jalobeanu · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu, et al · 2017
Earlier work this paper cites.
Smart glasses based intelligent trainer for factory new recruits
C.-F. Liu and P.-Y. Chiang · 2018
Earlier work this paper cites.
Hololens-based vascular localization system: precision evaluation study with a three-dimensional printed model
T. Jiang, D. Yu, Y. Wang, T. Zan, S. Wang, and Q. Li · 2020
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Earlier work this paper cites.
Hololens 2 research mode as a tool for computer vision research
D. Ungureanu, F. Bogo, S. Galliani, P. Sama, X. Duan, C. Meekhof, J. Stühmer, T. J. Cashman, B. Tekin, J. L. Schönberger, P. Olszta, and M. Pollefeys · 2020
Earlier work this paper cites.
Platform for situated intelligence, 2021
D. Bohus, S. Andrist, A. Feniello, N. Saw, M. Jalobeanu, P. Sweeney, A. L. Thompson, and E. Horvitz · 2021
Earlier work this paper cites.
Augmented reality maintenance assistant using yolov5
A. Malta, M. Mendes, and T. Farinha · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Developing mixed reality applications with platform for situated intelligence
S. Andrist, D. Bohus, A. Feniello, and N. Saw · 2022
Cited alongside, same era.
My view is the best view: Procedure learning from egocentric videos
S. Bansal, C. Arora, and C. Jawahar · 2022
Cited alongside, same era.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, J. Ma, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video, 2022
https://github.com/microsoft/psi/tree/master/Sources/MixedReality/HoloLensCapture
Platform for situated intelligence hololens capture tools · 2023
Later among the works it cites.
https://learn.microsoft.com/en-us/windows/mixed-reality/design/spatial-anchors/
Spatial anchors · 2023
Later among the works it cites.
https://stereokit.net
Stereokit · 2023
Later among the works it cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, et al · 2023
Later among the works it cites.
Can foundation models watch, talk and guide you step by step to make a cake?
Y. Bao, K. Yu, Y. Zhang, S. Storks, I. Bar-Yossef, A. de la Iglesia, M. Su, X. Zheng, and J. Chai · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Radosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, D. Batra, V. Cartillier, S. Crane, T. Do, M. Doulaty, A. Erapalli, C. Feichtenhofer, A. Fragomeni, Q. Fu, A. Gebreselasie, C. Gonzalez, J. Hillis, X. Huang, Y. Huang, W. Jia, W. Khoo, J. Kolar, S. Kottur, A. Kumar, F. Landini, C. Li, Y. Li, Z. Li, K. Mangalam, R. Modhugu, J. Munro, T. Murrell, T. Nishiyasu, W. Price, P. R. Puentes, M. Ramazanova, L. Sari, K. Somasundaram, A. Southerland, Y. Sugano, R. Tao, M. Vo, Y. Wang, X. Wu, T. Yagi, Z. Zhao, Y. Zhu, P. Arbelaez, D. Crandall, D. Damen, G. M. Farinella, C. Fuegen, B. Ghanem, V. K. Ithapu, C. V. Jawahar, H. Joo, K. Kitani, H. Li, R. Newcombe, A. Oliva, H. S. Park, J. M. Rehg, Y. Sato, J. Shi, M. Z. Shou, A. Torralba, L. Torresani, M. Yan, and J. Malik · 2022
Cited alongside, same era.
Augmented reality-based surgery on the human cadaver using a new generation of optical head-mounted displays: Development and feasibility study
B. Puladi, M. Ooms, M. Bellgardt, M. Cesov, M. Lipprandt, S. Raith, F. Peters, S. C. Möhlhenrich, A. Prescher, F. Hölzle, et al · 2022
Cited alongside, same era.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, M. Attarian, B. Ichter, K. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, et al · 2022
Cited alongside, same era.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Cited alongside, same era.
https://azure.microsoft.com/en-us/products/ai-services/ai-speech
Azure ai speech · 2023
Cited alongside, same era.
https://azure.microsoft.com/en-us/products/ai-services/openai-service
Azure openai service · 2023
Cited alongside, same era.
S. Castelo, J. Rulff, E. McGowan, B. Steers, G. Wu, S. Chen, I. Roman, R. Lopez, E. Brewer, C. Zhao, et al · 2023
Later among the works it cites.
Alexa arena: A user-centric interactive platform for embodied ai
Q. Gao, G. Thattai, X. Gao, S. Shakiah, S. Pansare, V. Sharma, G. Sukhatme, H. Shi, B. Yang, D. Zheng, et al · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu, K. Aggarwal, Z. Chi, J. Bjorck, V. Chaudhary, S. Som, X. Song, and F. Wei · 2023
Later among the works it cites.
LAVIS: A one-stop library for language-vision intelligence
D. Li, J. Li, H. Le, G. Wang, S. Savarese, and S. C. Hoi · 2023
Later among the works it cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
X. Wang, T. Kwon, M. Rad, B. Pan, I. Chakraborty, S. Andrist, D. Bohus, A. Feniello, B. Tekin, F. V. Frujeri, N. Joshi, and M. Pollefeys · 2023
Later among the works it cites.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Gao, and Y. J. Lee · 2023
Later among the works it cites.