Multimodal AI
A system able to receive, relate, or produce more than one modality, such as text, image, audio, and video. Multiple modalities do not solve the fusion problem: each carries its own noise, ambiguity, and cost. The architecture must decide which signal prevails when they conflict.