A couple years ago, I wrote about a simple recipe for good image alt text: good alt text lives at the intersection of three ingredients. Ask three simple questions about an image, content, context, purpose, and you’ll land on a description that’s meaningful instead of generic.
- What is in the image? (content)
- What is around the image? (context)
- Why is the image present? (purpose)
That post was about how a human should think through alt text. This one is about how to get an AI to think through it the same way, on the first try, instead of forty follow-up prompts in.
The problem with “generate alt text for this image”
It’s an easy trap to fall into once you start leaning on AI for accessibility tasks: upload an image, type “generate alt text for this image,” and get back something like “a photo of a building and some people outside.” Technically accurate. Completely useless. It’s the same complaint I raised in the original recipe post about AI-generated descriptions in general: the tool has no sense of context, so it can’t tell you why an image is there or what it’s doing for the page around it. The AI has none of the three ingredients. It’s guessing at content and skipping context and purpose entirely, because nobody handed those to it.
The fix isn’t a better model. It’s a better prompt. And the prompt writes itself once you notice that the three questions from the recipe map directly onto three prepositions.
Guided alt text generation
Here’s the template:
Give me alt text for this image of [content] for [purpose] about [context].
Fill in the three blanks and you’ve handed the AI everything it needs, in a form it can actually use.
“Of” carries the content. This answers what is in the image? What’s literally in the frame? A crowd, a chart, a red brick building. This is the part AI is already decent at guessing on its own from the pixels. It’s the other two it can’t infer.
“For” carries the purpose. This answers why is the image present? Why does it exist on your page at all? Is it proving a point, setting a mood, illustrating a step? A photo of a chair reads differently depending on whether it’s decorating a lifestyle post or anchoring a furniture listing.
“About” carries the context. This answers what is around the image? What’s it sitting next to? A picture of a building means one thing in an article about spring weather and something else entirely in an article about the building’s history. Same pixels, different job, same lesson from the recipe post: content and context together are what tell you what actually matters.
Three blanks, three prepositions, the same three questions from the original post. Once you see the pattern, you stop typing “generate alt text” and start typing a sentence that already contains the answer.
Putting it to the test
Harvard’s Digital Accessibility Services has a great teaching example on their describing images page: a single photo of Harvard Yard, students lounging in colorful Luxembourg chairs with Hollis Hall’s brick facade behind them. They use it to show how the same image needs different alt text depending on the article it’s placed in, which is exactly the point the recipe post made with its own conveyor belt and STEM toy examples. It’s a good testbed for Guided Alt Text Generation, so I ran it through three scenarios.

Photo courtesy of Harvard’s Digital Accessibility Services, used here as the shared image for all three scenarios below.
Scenario 1: the wrong way
Prompt: “Generate alt text for this image.”
Response: “Image of a building and people sitting outside.”
No content specifics, no purpose, no context. The AI filled the gap with the blandest true statement it could make, because that’s all it had to work with. This is what happens to most AI-generated alt text, and it’s a big part of why I tell clients that AI descriptions need a human review pass before they ship. The tool wasn’t given a chance to succeed.
Scenario 2: “Spring is here”
Imagine this photo running in an article celebrating the return of warm weather and the reappearance of Harvard’s Luxembourg chairs across the Yard.
Prompt: “Give me alt text for this image of students relaxing in colorful lawn chairs for an article entitled “Spring is Here” about the return of spring and the Luxembourg chairs to Harvard Yard.”
Response: “Students stretch out in bright yellow and green Luxembourg chairs scattered across the lawn of Harvard Yard, with Hollis Hall’s red brick facade in the background.”
Notice what made the cut and what didn’t. The chairs and their colors matter here because they’re the whole point of the article. Hollis Hall gets a passing mention for orientation, not a starring role.
Scenario 3: “Famous residents of Hollis Hall”
Now imagine the same photo illustrating a history piece about Hollis Hall itself and its place on campus.
Prompt: “Give me alt text for this image of Hollis Hall’s brick exterior for an article on famous residents of the dorm about the building’s notable occupants and prime location next to Harvard Yard’s open lawn.”
Response: “Hollis Hall’s red brick facade rises beside the open green lawn of Harvard Yard, with students relaxing in colorful chairs on the grass in front.”
Same photo, same three prepositions, completely different emphasis. Hollis Hall now leads. The chairs are still there, because they’re still part of the frame, but they’ve been demoted to scene-setting instead of the headline.
That’s the whole trick. The image never changes. The alt text does, because the purpose and context did, and the AI knew that because we told it.
Try it yourself
Next time you’re tempted to type “generate alt text for this image” and hope for the best, stop and fill in the blanks first: what’s in it, why it’s there, what it’s sitting next to. You’ll spend fifteen extra seconds writing the prompt and save yourself several rounds of rewriting the output. And the same rule from the original recipe post still applies here: take whatever the AI hands back and finesse it until it’s clear, concise, and uncluttered. Guided prompting gets you a much better first draft. It doesn’t replace the human touch.


