Review in context
The draft can identify visible content but cannot know the page purpose, surrounding copy, or why the image matters.
Advertisement
Generate a local first draft, then edit it before copying. The model can describe visible content; it cannot know why the image appears on your page, whether it is decorative, or which name and detail matter to your reader.
or drop one here — JPEG, PNG or WebP, up to 32 MB. It never leaves this device.
Checking what this device can do with Image captioning…
A captioning model reads the image and writes a description of it. It is a ~120 MB one-time download, cached afterwards, and it runs entirely on this device — the image is never sent anywhere.
If you decline: this tool cannot run. Unlike the other AI tools there is no classical shortcut: describing a photograph in words requires a model that has seen photographs.
Result proof
Input
Source image
Local pass
Local visual caption
Result
Editable text
How to verify it: Read the draft without looking at the image, add the page-specific purpose, remove repetition, and verify any names or text shown inside the pixels.
Text result, not a pixel edit
The caption model returns editable text for review. It does not touch the bitmap or create a misleading canvas-history entry, and it does not run until its dedicated model download is approved.
Editable text
Result type
The draft can identify visible content but cannot know the page purpose, surrounding copy, or why the image matters.
There is no classical caption shortcut, so declining consent leaves the image unchanged and explains why no result was generated.
Check names, identities, emotion, sensitive traits, and uncertain objects before publishing.
Know before you process: Useful alt text communicates the image purpose in its destination context; treat the generated description as a starting point, never an automatic accessibility guarantee.
Alt text replaces the image’s job in its surrounding page. A model sees visual objects, but it cannot know whether the same bridge photograph illustrates a location, a lighting technique or a construction detail. The generator solves the blank-page problem; the author still makes the accessibility decision.
Draft consistent descriptions in batches, then add product names, variants and the detail relevant to the listing.
Start from visible facts and add the people, place or event verified by the editor.
Fill missing alternatives efficiently while keeping every result editable before it reaches a publishing system.
Use the draft to spot whether the real requirement is concise alt text, empty alt, or a longer adjacent description.
Treat it as a first draft. Add the purpose, names and context that only the page author knows, then remove details the surrounding text already provides.
No. Decorative images should usually use empty alt text so a screen reader can skip them. Functional images need the action or destination, not merely their appearance.
No. Dense screenshots, charts and embedded text need human review and often a longer nearby explanation. Verify every word that matters.
No. The caption model runs locally after a one-time explicit download. The image and edited draft remain in this tab.
A useful visual description needs a trained vision-language model. The tool fails clearly when the model is unavailable instead of returning generic or misleading text.
As short as possible while replacing the image’s function in that page. Complex charts often need concise alt text plus a nearby detailed explanation.
The image stays local and the text stays editable until you approve it.
Draft alt text →