Image SEO for AI Search: How Multimodal Models Read Your Visuals
Alt text used to be the whole conversation about images and search: describe the picture, include a keyword, move on. Multimodal models changed the underlying situation. They process the image itself. They can tell you what is in a photograph, read the text inside a screenshot, and interpret a chart, without ever seeing your alt attribute.
That does not make alt text obsolete. It changes what it is for.
What multimodal models can and cannot get from an image
They can determine content. Objects, scenes, layout, colours, and text rendered inside the image, including words in a screenshot, a diagram label, or a slide.
They cannot determine identity or provenance. A model can see a modern office building. It cannot know it is your office, in which city, photographed by whom, or that the chart in it is your original research rather than someone else’s.
Scroll the diagram sideways to see all of it.
That gap is the entire job of the text around the image. You are no longer describing the picture for a machine that cannot see. You are supplying the facts a machine cannot infer from looking.
What this changes in practice
| Element | Old job | New job |
|---|---|---|
| Alt text | Describe the visual for accessibility and crawlers | Name the entities and supply context the pixels cannot carry |
| Filename | Keyword signal | Weak but free context signal |
| Caption | Optional design element | High-value, visible, extractable context |
| Surrounding copy | Loosely related | The primary source of the image’s meaning |
ImageObject markup | Rarely bothered with | Attaches the asset to your entity graph |
| Text inside images | Invisible to search | Readable, and now a liability if wrong |
That last row deserves attention. Text baked into images used to be simply unreadable. Now it is read. Which means an outdated price on a graphic, an old company name on a slide, or a stale statistic in an infographic is no longer harmlessly invisible. It is a contradictory fact about your brand, sitting on your own site, undermining the consistency that entity SEO depends on.
The practical checklist
1. Write alt text that names things
Not “team photo” and not “web design agency team photo Austin.” Write “Olga Kunger, founder of Ambeltek, at the studio’s Austin workspace.” Named entities, real relationships, plain language. This serves screen reader users properly and gives the model the identity layer it cannot see.
2. Use visible captions for anything that carries a claim
Captions are text on the page, so they chunk and get retrieved like any other passage. For charts, screenshots and data visuals, a caption stating what the image shows, the source and the date makes the visual quotable. An uncaptioned chart is a picture; a captioned one is evidence.
3. Audit text inside your images
Go through your graphics for outdated prices, old branding, superseded statistics and wrong dates. Anything contradicting your current site text is now actively read. Fix or replace.
4. Add ImageObject where it matters
For original photography, diagrams and research charts, ImageObject with caption, creator pointing to your Organization or author Person, contentUrl and licence information attributes the asset to your entity. This slots into the connected graph described in Organization schema for AI search and is part of our schema markup service.
5. Put the image near the text it supports
Retrieval works on passages. An image separated from the paragraph explaining it loses the context that makes it meaningful. Keep visual and explanation in the same section.
6. Replace stock photography with something specific
Generic stock contributes nothing a model can associate with you, and it looks the same as every competitor using the same library. Original photography, real screenshots of your work, and custom diagrams are all entity signals. Stock is not.
7. Do not let images break the technical basics
Descriptive filenames, correct dimensions, modern formats, lazy loading below the fold, and images that do not tank your Core Web Vitals. Speed still gates everything.
Where original visuals actually earn citations
The highest-value images for AI visibility are the ones that carry information nobody else has:
- Original data charts. A model summarizing a topic wants a source for the number. If your chart is the source, you get named.
- Process diagrams. A clear diagram of how something works is frequently the most quotable asset on a page.
- Real screenshots. Concrete, verifiable, and impossible for a competitor to duplicate.
- Comparison visuals. Well-structured, extractable, and matched to the comparison queries that fan-out generates.
These are the same properties that make text citable: specific, attributable, self-contained. Images are not a separate optimization discipline. They are the same discipline applied to a different asset type, which is the point made throughout our complete guide to generative engine optimization.
Where this leaves you
Stop thinking of alt text as the accessibility box you tick and start thinking of the whole image context as fact supply. The model sees the picture; you provide the names, dates, sources and relationships it cannot see. Then audit what your existing images are silently asserting, because whatever text is baked into them is now part of what AI systems believe about your brand.
If you want your visual assets audited alongside the rest of your entity signals, that is included in our AI SEO service.
Frequently asked questions
Do AI models actually look at images or just read alt text?
Multimodal models process the image itself, so they can describe what is in a picture without any alt text at all. Alt text, captions and surrounding copy still matter because they supply the context an image cannot carry on its own, such as who or what it depicts by name.
Is alt text still worth writing for AI search?
Yes, for two reasons. It remains a legal and accessibility requirement, and it supplies the naming and context that visual analysis cannot infer. A model can see a building; only your text tells it which building and whose.
What is ImageObject schema for?
ImageObject lets you attach structured facts to an image: its caption, creator, licence, and the URL it represents. It connects a visual asset to your entity graph so an image can be attributed to your organization rather than floating unattributed.
Do images help my brand get cited in AI answers?
Indirectly. Distinctive original imagery strengthens entity recognition and gives visual surfaces something to retrieve, while generic stock photography contributes nothing a model can associate specifically with your brand.