Fetching the paper…
Reading the bibliography…
Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound.
Measuring room impulse responses: Impact of the decay range on derived room acoustic parameters
Hak, C. C., Wenmaekers, R. H., and Van Luxemburg, L · 2012
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, K., Van Merriënboer, B., Bahdanau, D., and Bengio, Y · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., et al · 2017
Earlier work this paper cites.
AudioSet: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., et al · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
Adavanne, S., Politis, A., Nikunen, J., and Virtanen, T · 2018
Earlier work this paper cites.
Semi-orthogonal low-rank matrix factorization for deep neural networks
Povey, D., Cheng, G., Wang, Y., Li, K., Xu, H., et al · 2018
Earlier work this paper cites.
Pyroomacoustics: A python package for audio room simulation and array processing algorithms
Scheibler, R., Bezzam, E., and Dokmanic, I · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Joint measurement of localization and detection of sound events
Mesaros, A., Adavanne, S., Politis, A., Heittola, T., and Virtanen, T · 2019
Earlier work this paper cites.
The LOCATA challenge: Acoustic source localization and tracking
Evers, C., Löllmann, H. W., Mellmann, H., Schmidt, A., Barfuss, H., et al · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., et al · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2020
Earlier work this paper cites.
Learning representations from audio-visual spatial alignment
Morgado, P., Li, Y., and Nvasconcelos, N · 2020
Cited alongside, same era.
The USTC-Iflytek system for sound event localization and detection of dcase2020 challenge
Wang, Q., Wu, H., Jing, Z., Ma, F., Fang, Y., Wang, Y., Chen, T., Pan, J., Du, J., and Lee, C.-H · 2020
Cited alongside, same era.
Telling left from right: Learning spatial correspondence of sight and sound
Yang, K., Russell, B., and Salamon, J · 2020
Cited alongside, same era.
AST: Audio spectrogram transformer
Gong, Y., Chung, Y.-A., and Glass, J · 2021
Cited alongside, same era.
ACCDOA: Activity-coupled cartesian direction of arrival representation for sound event localization and detection
Shimada, K., Koyama, Y., Takahashi, N., Takahashi, S., and Mitsufuji, Y · 2021
Cited alongside, same era.
Pano-AVQA: Grounded audio-visual question answering on 360deg videos
Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models
Chu, Y., Xu, J., Zhou, X., Yang, Q., Zhang, S., et al · 2023
Later among the works it cites.
Pengi: An audio language model for audio tasks
Deshmukh, S., Elizalde, B., Singh, R., and Wang, H · 2023
Later among the works it cites.
LLaMA-Adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., et al · 2023
Later among the works it cites.
AudioGPT: Understanding and generating speech, music, sound, and talking head
Huang, R., Li, M., Yang, D., Shi, J., Chang, X., et al · 2023
Later among the works it cites.
AD-YOLO: You look only once in training multiple sound event localization and detection
Kim, J. S., Park, H. J., Shin, W., and Han, S. W · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yun, H., Yu, Y., Yang, W., Lee, K., and Kim, G · 2021
Cited alongside, same era.
Zhang, Y., Wang, S., Li, Z., Guo, K., Chen, S., and Pang, Y · 2021
Cited alongside, same era.
Mae-ast: Masked autoencoding audio spectrogram transformer
Baade, A., Peng, P., and Harwath, D. F · 2022
Cited alongside, same era.
SoundSpaces 2.0: A simulation platform for visual-acoustic learning
Chen, C., Schissler, C., Garg, S., Kobernik, P., Clegg, A., et al · 2022
Cited alongside, same era.
L3DAS22 challenge: Learning 3d audio sources in a real office environment
Guizzo, E., Marinoni, C., Pennese, M., Ren, X., Zheng, X., et al · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., et al · 2022
Cited alongside, same era.
Masked autoencoders that listen
Huang, P.-Y., Xu, H., Li, J., Baevski, A., Auli, M., et al · 2022
Cited alongside, same era.
Later among the works it cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Later among the works it cites.
GPT-4V(ision) system card
OpenAI · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Later among the works it cites.
Shimada, K., Politis, A., Sudarsanam, P., Krause, D., Uchida, K., et al · 2023
Later among the works it cites.
The nerc-slip system for sound event localization and detection of dcase2023 challenge
Wang, Q., Jiang, Y., Cheng, S., Hu, M., Nian, Z., et al · 2023
Later among the works it cites.
Listen, think, and understand
Gong, Y., Luo, H., Liu, A. H., Karlinsky, L., and Glass, J · 2024
Closest in time.
Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., et al · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., et al · 2024
Closest in time.