In plain English
A multimodal AI system is not limited to a single kind of input or output. It may be able to read text, inspect an image, listen to audio, or combine several of those information types in the same task.
Example
You could upload a photo of a damaged appliance and ask a multimodal assistant to identify visible problems and explain possible next steps in text.