Skip to main content

AI coding · Models

Multimodal

Also called: Multimodal model, Vision model, Image input

The model can handle more than text. It can understand images, screenshots, design mockups, and other kinds of input, and some can also generate images or audio.

In detail

A text-only model can only "hear your description." A multimodal model can also "look at the picture." That's really handy for frontend beginners: one screenshot often explains a problem better than a long paragraph, like "the spacing here is off," "the button is covered," or "I want this kind of layout."

Common uses: send the AI a screenshot of a design or reference page and have it build the same thing (Screenshot to Code), or send a screenshot of a broken layout and have it track down the styling issue.

Keep in mind that models also misread details in images, like exact pixel values, color values, and small text. It's best to add a text note next to the screenshot that points out the area you care about. For error messages, copy the actual text instead of just sending a screenshot.

Developer info
Term ID
ai-multimodal
DOM selectors
No DOM cues. This concept isn't detected directly on a page.
Priority
1 · when several match at the same level, the higher priority wins
Version
v1 · updated Sep 29, 2026