Fetching the paper…
Reading the bibliography…
In this paper, we propose a Collaboration of Experts (CoE) framework to pool together the expertise of multiple networks towards a common aim.
The transportation problem and the vogel approximation method
Shore, H. H · 1970
Earlier work this paper cites.
Learning specialized activation functions with the piecewise linear unit
Zhou, Y., Zhu, Z., and Zhong, Z · 1970
Earlier work this paper cites.
Neural network ensembles
Hansen, L. K. and Salamon, P · 1990
Earlier work this paper cites.
Hd-cnn: hierarchical deep convolutional neural networks for large scale visual recognition
Yan, Z., Zhang, H., Piramuthu, R., Jagadeesh, V., DeCoste, D., Di, W., and Yu, Y · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzębski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., and Lacoste-Julien, S · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollar, P., Tu, Z., and He, K · 2017
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ma, N., Zhang, X., Zheng, H.-T., and Sun, J · 2018
Earlier work this paper cites.
Hydranets: Specialized dynamic architectures for efficient inference
Mullapudi, R. T., Mark, W. R., Shazeer, N., and Fatahalian, K · 2018
Earlier work this paper cites.
Practical block-wise neural network architecture generation
Zhong, Z., Yan, J., Wu, W., Shao, J., and Liu, C.-L · 2018
Earlier work this paper cites.
Embedding complementary deep networks for image classification
Chen, Q., Zhang, W., Yu, J., and Fan, J · 2019
Cited alongside, same era.
Addressing failure prediction by learning model confidence
Corbière, C., THOME, N., Bar-Hen, A., Cord, M., and Pérez, P · 2019
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
Searching for mobilenetv3
Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et al · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Ghostnet: More features from cheap operations
Han, K., Wang, Y., Tian, Q., Guo, J., Xu, C., and Xu, C · 2020
Later among the works it cites.
Weightnet: Revisiting the design space of weight networks
Ma, N., Zhang, X., Huang, J., and Sun, J · 2020
Later among the works it cites.
Glance and focus: a dynamic approach to reducing spatial redundancy in image classification
Wang, Y., Lv, K., Huang, R., Song, S., Yang, L., and Huang, G · 2020
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Wen, Y., Tran, D., and Ba, J · 2020
Later among the works it cites.
Hyperparameter ensembles for robustness and uncertainty quantification
Wenzel, F., Snoek, J., Tran, D., and Jenatton, R · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Condconv: Conditionally parameterized convolutions for efficient inference
Yang, B., Bender, G., Le, Q. V., and Ngiam, J · 2019
Cited alongside, same era.
S4l: Self-supervised semi-supervised learning
Zhai, X., Oliver, A., Kolesnikov, A., and Beyer, L · 2019
Cited alongside, same era.
Beyer, L., Hénaff, O. J., Kolesnikov, A., Zhai, X., and van den Oord, A · 2020
Cited alongside, same era.
Once-for-all: Train one network and specialize it for efficient deployment
Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S · 2020
Cited alongside, same era.
Dynamic convolution: Attention over convolution kernels
Chen, Y., Dai, X., Liu, M., Chen, D., Yuan, L., and Liu, Z · 2020
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Dynamic graph: Learning instance-aware connectivity for neural networks
Yuan, K., Li, Q., Chen, D., Zhou, A., and Yan, J · 2020
Later among the works it cites.
Dynet: Dynamic convolution for accelerating convolutional neural networks
Zhang, Y., Zhang, J., Wang, Q., and Zhong, Z · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Closest in time.
Tradeoffs in data augmentation: An empirical study
Gontijo-Lopes, R., Smullin, S., Cubuk, E. D., and Dyer, E · 2021
Closest in time.
Training independent subnetworks for robust prediction
Havasi, M., Jenatton, R., Fort, S., Liu, J. Z., Snoek, J., Lakshminarayanan, B., Dai, A. M., and Tran, D · 2021
Closest in time.
Logme: Practical assessment of pre-trained models for transfer learning
You, K., Liu, Y., Wang, J., and Long, M · 2021
Closest in time.
Basisnet: Two-stage model synthesis for efficient inference
Zhang, M., Chu, C.-T., Zhmoginov, A., Howard, A., Jou, B., Zhu, Y., Zhang, L., Hwa, R., and Kovashka, A · 2021
Closest in time.