Fetching the paper…
Reading the bibliography…
Exponential families are widely used in machine learning; they include many distributions in continuous and discrete domains (e.g., Gaussian, Dirichlet, Poisson, and categorical distributions via the softmax transformation).
Sur les lois de probabilitéa estimation exhaustive
Georges Darmois · 1935
Earlier work this paper cites.
Sufficient statistics and intrinsic accuracy
Edwin James George Pitman · 1936
Earlier work this paper cites.
On distributions admitting a sufficient statistic
Bernard Osgood Koopman · 1936
Earlier work this paper cites.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Quantification method of classification processes. concept of structural a a -entropy
Jan Havrda and František Charvát · 1967
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Lev M Bregman · 1967
Earlier work this paper cites.
Non-parametric estimation of a multivariate probability density
Vassiliy A Epanechnikov · 1969
Earlier work this paper cites.
Gini-Simpson index of diversity: a characterization, generalization, and applications
R.A. Rao · 1982
Earlier work this paper cites.
Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory
Lawrence D Brown · 1986
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
John S. Bridle · 1990
Earlier work this paper cites.
Approximation of dynamical systems by continuous time recurrent neural networks
Ken-ichi Funahashi and Yuichi Nakamura · 1993
Earlier work this paper cites.
Adaptive sparseness using Jeffreys prior
M. Figueiredo · 2001
Earlier work this paper cites.
Sparse Bayesian learning and the relevance vector machine
M. Tipping · 2001
Earlier work this paper cites.
Geometry of escort distributions
Sumiyoshi Abe · 2003
Earlier work this paper cites.
Entropy and diversity
L. Jost · 2006
Earlier work this paper cites.
Generalized Maximum Entropy, Convexity and Machine Learning
Timothy Sears · 2008
Earlier work this paper cites.
The q-exponential family in statistical physics
Jan Naudts · 2009
Earlier work this paper cites.
t-logistic regression
Nan Ding and S.V.N. Vishwanathan · 2010
Earlier work this paper cites.
Geometry of q-exponential family of probability distributions
Shun-ichi Amari and Atsumi Ohara · 2011
Cited alongside, same era.
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Cited alongside, same era.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Heinz Bauschke and Patrick Combettes · 2011
Cited alongside, same era.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
Geometry for q-exponential families
Hiroshi Matsuzoe and Atsumi Ohara · 2012
Cited alongside, same era.
Measure Theory , volume 18
Paul R Halmos · 2013
Cited alongside, same era.
Schnet: A continuous-filter convolutional neural network for modeling quantum interactions
Kristof Schütt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert Müller · 2017
Later among the works it cites.
Deep parametric continuous convolutional neural networks
Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, and Raquel Urtasun · 2018
Later among the works it cites.
Neural ordinary differential equations
Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Later among the works it cites.
Von mises-fisher loss for training sequence to sequence models with continuous outputs
Sachin Kumar and Yulia Tsvetkov · 2018
Later among the works it cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André F.T. Martins · 2019
Later among the works it cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ole Barndorff-Nielsen · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Rezende, and Daan Wierstra · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Later among the works it cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Later among the works it cites.
Latent ordinary differential equations for irregularly-sampled time series
Yulia Rubanova, Tian Qi Chen, and David K Duvenaud · 2019
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Later among the works it cites.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Tianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang, and Vicente Ordonez · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Joey nmt: A minimalist nmt toolkit for novices
Julia Kreutzer, Joost Bastings, and Stefan Riezler · 2019
Later among the works it cites.
Learning with fenchel-young losses
Mathieu Blondel, André FT Martins, and Vlad Niculae · 2020
Closest in time.
Hard-coded Gaussian attention for neural machine translation
Weiqiu You, Simeng Sun, and Mohit Iyyer · 2020
Closest in time.