Azzalini A. A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. 1985;12(2):171–178.
Azzalini A. Further Results on a Class of Distributions Which Includes the Normal Ones. Statistica. 1986;46:199–208.
Azzalini A, Capitanio A. Distributions Generated by Perturbation of Symmetry with Emphasis on a Multivariate Skew t Distribution. Journal of the Royal Statistical Society: Series B. 2003;65(2):367–389.
Azzalini A, Dalla Valle A. The Multivariate Skew-Normal Distribution. Biometrika.1996;83(4):715–726.
Bai Z, Rao CZ, Wu CZ. Model Selection: The Efficient Determination Criterion. Journal of Multivariate Analysis. 1989;30(2):180–198.
Basford K, Greenway D, McLachlan G, Peel D. Standard errors of fitted component means of normal mixtures. Computational Statistics. 1997;12(1):1–18.
Baudry JP, Celeux G. EM for mixtures: convergence, likelihood monotonicity, and initialization. Statistics and Computing. 2015;25(4):713–728.
Biernacki C, Celeux G, Govaert G. Assessing a Mixture Model for Clustering with the Integrated Completed Likelihood. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2000;22(7):719–725.
Capitanio A, Azzalini A, Stanghellini E. Graphical models for skew-normal variates. Scandinavian Journal of Statistics. 2003;30(1):129–144.
Chamroukhi F. Skew t mixture of experts. Neurocomputing. 2017;266:390–408.
Dempster AP, Laird NM, Rubin DB. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B. 1977;39(1):1–38.
Gormley IC, Murphy TB. A mixture of experts model for rank data with applications in election studies. 2008;.
Hubert L, Arabie P. Comparing Partitions. Journal of Classification. 1985;2(1):193–218.
Jacobs RA, Jordan MI, Nowlan SJ, Hinton GE. Adaptive mixtures of local experts.
Neural Computation. 1991;3(1):79–87.
Jamalizadeh A, Lin TI. A general class of scale-shape mixtures of skew-normal distributions: properties and estimation. Computational Statistics. 2017;32(2):451–474.
Liu C, Rubin DB. The ECME Algorithm: A Simple Extension of EM and ECM with Faster Monotone Convergence. Biometrika. 1994;81(4):633–648.
Louis TA. Finding the observed information matrix when using the EM algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology. 1982;44(2):226–233.
Mazza A, Punzo A. Model-based clustering and classification with mixtures of contaminated Gaussian regression models. Advances in Data Analysis and Classification.2017;11:753–777.
Mazza A, Punzo A. Mixtures of multivariate contaminated normal regression models. Statistical Papers. 2020;61(2):787–822.
McNeil AJ, Frey R, Embrechts P. Quantitative risk management: concepts, techniques and tools-revised edition. Princeton university press; 2015.
MeilĖa M. Comparing Clusterings—An Information Based Distance. Journal of Multivariate Analysis. 2007;98(5):873–895.
Meng XL, Rubin DB. Maximum likelihood estimation via the ECM algorithm: A general framework. Biometrika. 1993;80(2):267–278.
Naderi M, Mirfarah E, Wang WL, Lin TI. Robust mixture regression modeling based on the normal mean-variance mixture distributions. Computational Statistics & Data Analysis. 2023;180:107661.
Nguyen HD, McLachlan GJ. Laplace mixture of linear experts. Computational Statistics & Data Analysis. 2016;93:177–191.
Peng F, Jacquier E, Strahan P. Bayesian Inference for Mixture Models. Journal of the American Statistical Association. 1996;91(433):187–196.