AI coding · Models
Multimodal
Also called: Multimodal model, Vision model, Image input
The model can handle more than text. It can understand images, screenshots, design mockups, and other kinds of input, and some can also generate images or audio.
In detail
A text-only model can only "hear your description." A multimodal model can also "look at the picture." That's really handy for frontend beginners: one screenshot often explains a problem better than a long paragraph, like "the spacing here is off," "the button is covered," or "I want this kind of layout."
Common uses: send the AI a screenshot of a design or reference page and have it build the same thing (Screenshot to Code), or send a screenshot of a broken layout and have it track down the styling issue.
Keep in mind that models also misread details in images, like exact pixel values, color values, and small text. It's best to add a text note next to the screenshot that points out the area you care about. For error messages, copy the actual text instead of just sending a screenshot.
Developer infoTerm ID, DOM cues, match priority
- Term ID
ai-multimodal- DOM selectors
- No DOM cues. This concept isn't detected directly on a page.
- Priority
- 1 · when several match at the same level, the higher priority wins
- Version
- v1 · updated Sep 29, 2026