MIHMMs: Mutual Information Hidden Markov Models

Approach

Hidden Markov Models are usually trained to maximize the likelihood of the observations and hidden states. That objective gives a good generative model, but it does not directly reward hidden states for being informative about the observations—a useful property in classification and clustering.

Mutual Information Hidden Markov Models (MIHMMs) keep the graphical structure of an HMM but change the training objective to a weighted combination of likelihood and mutual information between observations and hidden states. We derived parameter-estimation equations for discrete and continuous observations, in supervised and unsupervised settings.

Across the synthetic and real classification tasks reported in the paper, MIHMMs outperformed standard maximum-likelihood HMMs. These experiments established the behavior of the proposed objective on selected tasks rather than a universal advantage: performance still depended on the data, model structure, and weighting between likelihood and mutual information.

Publications

Conference paper
PDF External link
Cite
Formatted citation

Nuria Oliver, Ashutosh Garg (2002). MIHMM: Mutual Information Hidden Markov Models. International Conference on Machine Learning (ICML 2002). https://www.microsoft.com/en-us/research/publication/mmihmm-maximum-mutual-information-hidden-markov-models/

BibTeX
@inproceedings{oliver2002mihmm,
  author = {Nuria Oliver and Ashutosh Garg},
  title = {{MIHMM}: Mutual Information Hidden {Markov} Models},
  booktitle = {International Conference on Machine Learning (ICML 2002)},
  year = 2002,
  url = {https://www.microsoft.com/en-us/research/publication/mmihmm-maximum-mutual-information-hidden-markov-models/}
}

MIHMM: Mutual Information Hidden Markov Models

Nuria Oliver, Ashutosh Garg
International Conference on Machine Learning (ICML 2002) · 2002