Fetching the paper…
Reading the bibliography…
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
Iterative reclassification procedure for constructing an asymptotically optimal rule of allocation in discriminant analysis
McLachlan, G. J · 1975
Earlier work this paper cites.
Active touch and robot perception
Goldberg, K. Y. and Bajcsy, R · 1984
Earlier work this paper cites.
Multiresolution gray-scale and rotation invariant texture classification with local binary patterns
Ojala, T., Pietikainen, M., and Maenpaa, T · 2002
Earlier work this paper cites.
The skin and its receptors 148 pathways to cortex and major cortical areas
Klatzky, R. L. and Lederman, S. J · 2003
Earlier work this paper cites.
The psychology of multimodal perception
Bertelson, P. and De Gelder, B · 2004
Earlier work this paper cites.
A comparison of learning with haptic and visual modalities
Jones, M. G., Bokinsky, A., Tretter, T., and Negishi, A · 2005
Earlier work this paper cites.
Semi-supervised self-training of object detection models
Rosenberg, C., Hebert, M., and Schneiderman, H · 2005
Earlier work this paper cites.
Vision and touch are automatically integrated for the perception of sequences of events
Bresciani, J.-P., Dammeier, F., and Ernst, M. O · 2006
Earlier work this paper cites.
Memory for curvature of objects: Haptic touch vs. vision
Ittyerah, M. and Marks, L. E · 2007
Earlier work this paper cites.
Development of a tactile sensor based on biologically inspired edge encoding
Chorley, C., Melhuish, C., Pipe, T., and Rossiter, J · 2009
Earlier work this paper cites.
Tactile sensing—from humans to humanoids
Dahiya, R. S., Metta, G., Valle, M., and Sandini, G · 2009
Earlier work this paper cites.
Coding and use of tactile signals from the fingertips in object manipulation tasks
Johansson, R. S. and Flanagan, J. R · 2009
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Lee, D.-H. et al · 2013
Earlier work this paper cites.
Sensing and recognizing surface textures using a gelsight sensor
Li, R. and Adelson, E. H · 2013
Earlier work this paper cites.
Multimodal interaction: A review
Turk, M · 2014
Earlier work this paper cites.
The contributions of vision and haptics to reaching and grasping
Stone, K. D. and Gonzalez, C. L · 2015
Earlier work this paper cites.
Multi-sensorial and explorative recognition of garments and their material properties in unconstrained environment
Kampouris, C., Mariolis, I., Peleka, G., Skartados, E., Kargakos, A., Triantafyllou, D., and Malassiotis, S · 2016
Earlier work this paper cites.
Combining finger vision and optical tactile sensing: Reducing and handling errors while cutting vegetables
Yamaguchi, A. and Atkeson, C. G · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Wearable haptic systems for the fingertip and the hand: Taxonomy, review, and perspectives
Pacchierotti, C., Sinclair, S., Solazzi, M., Frisoli, A., Hayward, V., and Prattichizzo, D · 2017
Earlier work this paper cites.
Gelsight: High-resolution robot tactile sensors for estimating geometry and force
Yuan, W., Dong, S., and Adelson, E. H · 2017
Earlier work this paper cites.
More than a feeling: Learning to grasp and regrasp using vision and touch
Calandra, R., Owens, A., Jayaraman, D., Lin, J., Yuan, W., Malik, J., Adelson, E. H., and Levine, S · 2018
Earlier work this paper cites.
Verbal labels facilitate tactile perception
Miller, T. M., Schmidt, T. T., Blankenburg, F., and Pulvermüller, F · 2018
Earlier work this paper cites.
Active clothing material perception using tactile sensing and deep learning
Yuan, W., Mo, Y., Wang, S., and Adelson, E. H · 2018
Earlier work this paper cites.
Connecting touch and vision via cross-modal prediction, 2019
Li, Y., Zhu, J.-Y., Tedrake, R., and Torralba, A · 2019
Earlier work this paper cites.
Neuronal correlates of label facilitated tactile perception
Schmidt, T. T., Miller, T. M., Blankenburg, F., and Pulvermüller, F · 2019
Earlier work this paper cites.
Design, motivation and evaluation of a full-resolution optical tactile sensor
Sferrazza, C. and D’Andrea, R · 2019
Earlier work this paper cites.
Tactile image sensors employing camera: A review
Shimonomura, K · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation
Lambeta, M., Chou, P.-W., Tian, S., Yang, B., Maloon, B., Most, V. R., Stroud, D., Santos, R., Byagowi, A., Kammerer, G., Jayaraman, D., and Calandra, R · 2020
Cited alongside, same era.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E. D., Kurakin, A., Zhang, H., and Raffel, C · 2020
Cited alongside, same era.
Long-tailed recognition by routing diverse distribution-aware experts
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P · 2023
Later among the works it cites.
Safe self-supervised learning in real of visuo-tactile feedback policies for industrial insertion, 2023
Fu, L., Huang, H., Berscheid, L., Li, H., Goldberg, K., and Chitta, S · 2023
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, X., Lian, L., Miao, Z., Liu, Z., and Yu, S. X · 2020
Cited alongside, same era.
Integration of haptics and vision in human multisensory grasping
Camponogara, I. and Volcic, R · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers, 2021
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations
Gao, R., Chang, Y.-Y., Mall, S., Fei-Fei, L., and Wu, J · 2021
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio, 2021
Guzhov, A., Raue, F., Hees, J., and Dengel, A · 2021
Cited alongside, same era.
Openclip, July 2021
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Crossmodal associations with olfactory, auditory, and tactile stimuli in children and adults
Speed, L. J., Croijmans, I., Dolscheid, S., and Majid, A · 2021
Cited alongside, same era.
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I · 2023
Later among the works it cites.
Point-bind and point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following, 2023
Guo, Z., Zhang, R., Zhu, X., Tang, Y., Ma, X., Han, J., Chen, K., Gao, P., Li, X., Li, H., and Heng, P.-A · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning, 2023
Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., Guo, Z., Lu, X., Ren, S., Wen, Y., Chen, X., Yue, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Instruct-nerf2nerf: Editing 3d scenes with instructions
Haque, A., Tancik, M., Efros, A., Holynski, A., and Kanazawa, A · 2023
Later among the works it cites.
Learning to read braille: Bridging the tactile reality gap with diffusion models
Higuera, C., Boots, B., and Mukadam, M · 2023
Later among the works it cites.
Self-supervised visuo-tactile pretraining to locate and follow garment features, 2023
Kerr, J., Huang, H., Wilcox, A., Hoque, R., Ichnowski, J., Calandra, R., and Goldberg, K · 2023
Later among the works it cites.
Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models, 2023
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., Han, J., Huang, S., Zhang, Y., He, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action, 2023
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A · 2023
Later among the works it cites.
Anymal: An efficient and scalable any-modality augmented language model, 2023
Moon, S., Madotto, A., Lin, Z., Nagarajan, T., Smith, M., Jain, S., Yeh, C.-F., Murugesan, P., Heidari, P., Liu, Y., Srinet, K., Damavandi, B., and Kumar, A · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B · 2023
Later among the works it cites.
General In-Hand Object Rotation with Vision and Touch
Qi, H., Yi, B., Ma, Y., Suresh, S., Lambeta, M., Calandra, R., and Malik, J · 2023
Later among the works it cites.
Robot learning with sensorimotor pre-training
Radosavovic, I., Shi, B., Fu, L., Goldberg, K., Darrell, T., and Malik, J · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all, 2023
Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D · 2023
Later among the works it cites.
Generative multimodal models are in-context learners, 2023
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Luo, Z., Wang, Y., Rao, Y., Liu, J., Huang, T., and Wang, X · 2023
Later among the works it cites.
Codi-2: In-context, interleaved, and interactive any-to-any generation, 2023
Tang, Z., Yang, Z., Khademi, M., Liu, Y., Zhu, C., and Bansal, M · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Cut and learn for unsupervised object detection and instance segmentation, 2023
Wang, X., Girdhar, R., Yu, S. X., and Misra, I · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S · 2023
Later among the works it cites.
Visual-tactile learning of garment unfolding for robot-assisted dressing
Zhang, F. and Demiris, Y · 2023
Later among the works it cites.
Llama-adapter: Efficient finetuning of language models with zero-init attention
Zhang, R., Han, J., Liu, C., Gao, P., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Multimodal visual-tactile representation learning through self-supervised contrastive pre-training
Dave, V., Lygerakis, F., and Rueckert, E · 2024
Closest in time.
Binding touch to everything: Learning unified multimodal tactile representations
Yang, F., Feng, C., Chen, Z., Park, H., Wang, D., Dou, Y., Zeng, Z., Chen, X., Gangopadhyay, R., Owens, A., et al · 2024
Closest in time.