The analogy
Imagine an assistant who can only understand you by voice, blindfolded. It's useful, but blind: everything you want to show it you have to describe in words, and some things are a nightmare to explain in words. "There's a stripped screw at the top right, no, further right..."
Give it eyes and hearing too — let it see the photo, hear your voice — and suddenly it can help you with things that were impossible before. The multimodal is the AI's extra senses: besides reading, it looks and listens, and connects what it sees with what you ask it.
How it really works
A multimodal AI is trained to handle several types of input within the same space, so it can reason across them: it sees an image and answers your question about that image, connecting the two planes. The most recent models are born multimodal from the start, built to interweave text, images, audio and video instead of having them added on later. For you, at a practical level, this translates into concrete features: uploading a photo, talking by voice, having it read a screenshot or a scanned letter, generating images.
What you can do in practice
- Photograph instead of describing: a broken object, an error code, a sick plant, a handwritten note. The AI reads the photo and works on that.
- Talk instead of writing, hands-free while cooking or walking.
- Have it read a screenshot, a photographed document, a chart, and ask for an explanation or a summary.
- Have it describe an image, or generate one, for projects, ideas, accessibility.
A common misconception
People think that if the AI "sees" an image, then it understands it perfectly. It sees and interprets many things well, but it can get the details wrong, misread small numbers, confuse an unclear text, or "see" something that isn't there. Visual analysis is a powerful help, not an infallible eye: the results that matter must be checked, especially figures and texts inside images.
Frequently asked questions
With the free versions can I upload photos and use voice?
Generally yes: uploading an image and talking by voice are by now common features even in the free plans, with the version's usage limits. With many images analyzed in a row a daily cap may kick in, but for normal use you won't hit it.
Does it read handwriting?
Often yes, especially if the handwriting is legible and the photo is in focus. On difficult or faded writing it can get it wrong: always reread the transcription of important things, like a medical prescription or a number.
Can it watch a video?
Some models do, analyzing the images and audio of the clip and pinpointing precise moments. It's a more recent and less widespread capability than reading photos: it depends on the tool you use.
Does multimodal mean it sees the way we see?
No, and that's the underlying confusion. It has no eyes nor experience of the world: it turns the image into data and recognizes its patterns, a bit like it does with text. That's why it sometimes "sees" details that aren't there or misses obvious ones. It understands images in a statistical way, not by looking at them like a person: extremely useful, but with a margin of error that human sight doesn't have on the same details.