MAMI: Multimodal Automatic Mobile Annotations

MAMI was built around a simple idea: a person should be able to organize and find photographs by speaking to the phone. Its full name is Multimodal Automatic Mobile Indexing.

When taking a photograph, the user can record a spoken annotation or add one later, while MAMI automatically stores the date, time, and location. Finding the photograph requires another spoken query, which the system compares with the stored recordings.

Nothing is sent to a server. Instead of converting speech into text, MAMI does all the processing on the phone and compares the acoustic pattern of the query with the existing recordings, allowing photographs to be indexed and retrieved without speech recognition or network access.

A companion field study compared spoken and typed annotations on camera phones. It showed why the choice could not be reduced to input speed alone: the setting in which a photograph was captured, the effort of typing, and the social acceptability of speaking to a phone all shaped which modality people used.

MAMI was a preliminary prototype for a personal collection. Matching acoustic patterns does not understand the words being spoken, and recordings can vary with the speaker and surrounding noise. The work therefore established a mobile annotation and retrieval approach, not a general speech-search system.

MAMI recording a spoken annotation for a photograph. MAMI displaying photographs returned by a spoken search.

MAMI in capture/annotation mode (left) and search mode (right).

Publications

Conference paper
DOI
Cite
Formatted citation

Xavier Anguera, JieJun Xu, Nuria Oliver (2008). Multimodal Photo Annotation and Retrieval on a Mobile Phone. Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval (MIR 2008), 188-194. https://doi.org/10.1145/1460096.1460127

BibTeX
@inproceedings{anguera2008photo,
  author = {Xavier Anguera and JieJun Xu and Nuria Oliver},
  title = {Multimodal Photo Annotation and Retrieval on a Mobile Phone},
  booktitle = {Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval (MIR 2008)},
  pages = {188--194},
  year = 2008,
  doi = {10.1145/1460096.1460127},
  cites = 45,
  citesdate = {2026-03-23}
}

Multimodal Photo Annotation and Retrieval on a Mobile Phone

Xavier Anguera, JieJun Xu, Nuria Oliver
Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval (MIR 2008) · 2008
Conference paper
DOI
Cite
Formatted citation

Xavier Anguera, Nuria Oliver (2008). MAMI: Multimodal Annotations on a Camera Phone. Proceedings of the 10th International Conference on Human Computer Interaction with Mobile Devices and Services (MobileHCI 2008), 379-382. https://doi.org/10.1145/1409240.1409293

BibTeX
@inproceedings{anguera2008mami,
  author = {Xavier Anguera and Nuria Oliver},
  title = {{MAMI}: Multimodal Annotations on a Camera Phone},
  booktitle = {Proceedings of the 10th International Conference on Human Computer Interaction with Mobile Devices and Services (MobileHCI 2008)},
  pages = {379--382},
  year = 2008,
  doi = {10.1145/1409240.1409293},
  cites = 8,
  citesdate = {2026-03-23}
}

MAMI: Multimodal Annotations on a Camera Phone

Xavier Anguera, Nuria Oliver
Proceedings of the 10th International Conference on Human Computer Interaction with Mobile Devices and Services (MobileHCI 2008) · 2008
Conference paper
PDF DOI
Cite
Formatted citation

Mauro Cherubini, Xavier Anguera, Nuria Oliver, Rodrigo De Oliveira (2009). Text Versus Speech: A Comparison of Tagging Input Modalities for Camera Phones. Proceedings of the 11th International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI 2009), 1-10. https://doi.org/10.1145/1613858.1613860

BibTeX
@inproceedings{cherubini2009text,
  author = {Mauro Cherubini and Xavier Anguera and Nuria Oliver and Rodrigo De Oliveira},
  title = {Text Versus Speech: A Comparison of Tagging Input Modalities for Camera Phones},
  booktitle = {Proceedings of the 11th International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI 2009)},
  pages = {1--10},
  year = 2009,
  doi = {10.1145/1613858.1613860},
  award = {Best Paper Award Nominee},
  cites = 14,
  citesdate = {2026-03-23}
}
★ Best Paper Award Nominee

Text Versus Speech: A Comparison of Tagging Input Modalities for Camera Phones

Mauro Cherubini, Xavier Anguera, Nuria Oliver, Rodrigo De Oliveira
Proceedings of the 11th International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI 2009) · 2009