Clicked Gallery

What is a Multimodal Model?

Highlighted from a real engineering doc. Explained by Clicked.

Used in a sentence

Engineering Notes · AI Systems

The lab's newest multimodal model reads a photograph, a spreadsheet and a spoken question in the same request.

The reader highlighted one word in the docs. Clicked made the technical term “multimodal model” easy to understand:

Explained in three depths

Same facts, different vibe — Slang mode 😎

The Clicked way

●○○

Overview

A multimodal model is an AI that handles more than one kind of input, such as text, images, audio and video, inside a single system. It is why you can photograph a page and ask a question about it in the same breath. Earlier models could only read text, so anything else had to be described to them first.
●○○

Overview

A multimodal model can take pictures, sound and video, not only typing. That is why you can point a phone at something and ask about the thing itself. The older sort could read and nothing else, so you had to put the world into words first. 😎

A quick take — often all you need.

●●○

Detail

Older models could only read text, so anything else had to be described in words first, and most things do not survive that translation. A multimodal model removes that step by turning every kind of input into the same internal format before it processes anything. A sentence, a photograph and a few seconds of speech all become long lists of numbers. Once they are numbers the model treats them alike, which is what lets a question about an image sit in the same request as the image itself. So a photograph of a receipt or an error screen goes straight in, and it can answer things needing both at once, such as what is wrong with this chart. The cost is size and speed, since images and audio become far more numbers than a sentence does. That eats into the context window and makes each answer slower and more expensive. Being able to take an input is not the same as handling it well, so a model that reads a typed page reliably may still misread a hand-drawn diagram.
●●○

Detail

Older models had one door in and you were the doorman, so whatever you could not put into words never got inside, and your vocabulary was the ceiling on what the thing could know. The multimodal sort knocked the wall down. Pictures, sound and video all get melted down into the same numbers before it starts thinking, so a photo and a question about that photo travel together. Point your phone at a menu in Lisbon and ask what a vegetarian can safely order, and it just answers. Two bills arrive for this. A photo devours far more of its attention span than a sentence, so you burn through the window faster and wait longer. And seeing something is not the same as seeing it correctly. The model that reads a page of dense text without blinking will study your whiteboard scrawl and confidently describe a diagram that was never on it. 😎

Want more? One click digs deeper.

●●●

Analogy

Talking someone through a computer problem over the phone. Everything you know arrives through their description, so if they never mention the flashing red light, it does not exist as far as you are concerned. Stand next to the machine and you read the error yourself, spot the unplugged cable and hear the clicking. You did not get smarter, you got access to information that was there the whole time.
●●●

Analogy

Describing a song by typing the three lyrics you half remember, versus just humming it into your phone. Same song, same you, wildly different odds of being told what it is. One route has to squeeze the whole thing through your vocabulary before anything else can happen. The other lets the machine hear the actual noise and skip you entirely.

Unfamiliar concept? A real-world example makes it click — fresh analogies on tap.

AI explanations may contain errors · Not professional advice

Formal definition — The same term, explained the usual way

A multimodal model is a machine learning system trained to accept inputs of multiple types, typically text, images, audio or video, by projecting each into a shared representation space that permits joint processing within a single forward pass. This enables cross-modal reasoning without intermediate textual description, at the cost of increased token consumption, latency and modality-dependent variation in reliability.

Want Clicked to explain terms like “multimodal model” directly in your browser — including on PDFs?

Add to Chrome — Free

50 free Explanations · No credit card required