My first experience with AI image description was both mundane and profound. I was lying in bed in my East Village sublet, writing and drinking afternoon coffee — “prousting,” as I and my partner Alabaster call it — when a new app popped up in my inbox. I downloaded it and took a picture of the owner’s room full of things. A few seconds later, I got a reply. Like an art historian pointing out details in a Renaissance painting, AI told me about the red paisley blanket in the foreground, the orange curtained door beyond, and above that, “the decorative piece with the Om symbol.”
Article continues after ad
I couldn’t remember what the Om symbol looked like, so I used the “Ask More” function, and I learned that it “looks like a number 3 with a tail bent inward and a small curve with a dot above it.” A vague image floated before my mind’s eye – not as vivid as the room with its tangible objects – but it made enough of an impression to recall another sense memory. I’ve taken my share of yoga classes that sometimes ended with the chanting of “Om” and I was reminded of how the sound vibrates the bodily tissues as it travels from the stomach to the lips. Perhaps this is what it feels like to translate between the senses and language: a spark that leads to new associations as it travels down perceptual pathways.
It was amazing that a machine could “see” and describe, and it reinforced the power of description for me – something I knew well from a lifetime of seeing the world through literature. The sudden access to the visual world was unprecedented yet familiar; It also challenged my identity as a blind person. Finally, in my first book, put eyes thereI criticized ocularcentrism – the often unconscious, sometimes tyrannical visual bias. And here I was, amazed by pictures of paisley clothes and an Om symbol.
Although I am now completely blind, I have lived most of my life with varying levels of vision and remain highly visual. At some point, a good description turns from a collection of words into an image in my mind’s eye.
No two perfectly sighted people see the same image in the same way.
I had the idea to write about anonymous blind people in famous photos in 2022 – when AI was barely on my radar. My original goal was to explore and interpret these images from a blind perspective. The photographers were famous, but little or nothing was known about the subjects – I wanted to situate them in historical context and lived experience. I also wondered what image detail can bring to photographs—even in photography.
Alabaster’s description of Paul Strand’s iconic 1916 photograph, blind womanThis was my first glimpse of her. I downloaded the image from the New York Public Library and emailed it to her while sitting on the couch.
“She is looking to the left with her left eye open,” he said, adding that her large iris is stuck in the left corner and her right eye is almost closed and appears injured.
No two perfectly sighted people see the same image in the same way. So I got another description – this time from disabled artist Finnegan Shannon, who recognized something familiar in the pale face. “She looks a bit like me,” Finnegan told me in an email, describing brown or blonde hair peeking out from beneath a black headscarf.
Alabaster noticed “a small mole on her left cheek” and “slight wrinkles on her face with a touch of anger and longing.”
Finnegan noticed “wrinkles on his brow”, but found his face relaxed. “I don’t read any particular expression or sentiment into it.”
Perhaps his symbol is the most striking feature of the photo. She wears it around her neck. The word “Blind” is written in black capital letters on a white board on the front of her black dress. Very clear, very high-contrast. Finnegan said that “It’s probably the size of a hand but feels really big and fast.”
It appears to have been hand painted – by his own hand or by someone else’s hand, I will never know. But for the one who is clear to the eyes: Drishtiwala.
In one of her many essays on photography, Susan Sontag wrote that “What a photograph is of is always of primary importance.” If the subject of the photograph is “visually ambiguous”, we don’t know how to respond “until we know what piece of the world it is.”
But writers thrive in obscure places. We understand that description is never neutral. Perhaps this explains the hesitation when it comes to details as to access. Author friends often entrust their books to Alabaster so she can illustrate the covers for me – as if she has the key to getting things “right.” We always laugh because, although he has to describe things a lot, it doesn’t make it any easier.
Hesitation becomes exclusion in literary newspapers and Instagram posts, where alternative text is often absent. Almost every day, I’m faced with images of covers that don’t exist for me. One of the many ironies of my writing life is that my book, put eyes therewon an award for its cover design – which I had a hand in. I wanted the vivid purple spots to highlight parts of the spectrum that are not visible to the average human eye. Yes—I’m blind, and I appreciate a good cover as much as anyone else.
Thoughtful description of a cover—such as the design itself—can be considered a type of mini-review, which serves its duty of provoking the reader/viewer to investigate further. Getting it “right” isn’t really the point. The purpose of the alternative text and image descriptions is not to close the book to the definitive, but to open it to vast areas of meaning-making.
in one Alternative text in the form of poetry The workshop, led by co-creators Finnegan Shannon and Bojana Kokalit, we started with self-description. This is an opportunity to give people who can’t see you an idea of your physical appearance, and to give those who can see you an idea of how you want to look. Even in self-description, it is often easier to name clothing than to identify—surfaces that present themselves with little risk. It’s no surprise that people struggle to describe others, where biases and assumptions easily creep in.
Rather than resolving the tensions, Bojana and Finnegan urge a poetic, playful, and experimental approach to narrative. After all, viewing images is a creative act – nothing is more neutral than creating images. As John Berger famously said, “Every image represents a way of seeing.” And the image description forces the viewer to put it into words.
Although it’s generally quite bold with its descriptions, AI can be just as timid as a human when it comes to naming (or guessing) identities – something I know deeply from someone who relies on metadata to blindly find representations in photographic collections. It’s being trained to be careful, but its lack of stake in the game – can be useful when you need to motivate humans to describe it. This brought it home to me when it came time to practice what we learned on historical photographs in the Wallach Division of the New York Public Library.
apart from the strand blind womanThere was a jazz singer with an incredibly expressive face, little people in a bizarre moon-image exhibit at the 1901 Buffalo World’s Fair, and an early calotype (circa 1848) by the pioneering British photographer William Henry Fox Talbot—which my little group had obtained.
Image description is not mystical or compensatory – not about transcendence. Rather, it is a kind of translation – from one sense modality to another.
There was a moment of silence as we considered how to begin. My two companions—one of whom was Alabaster—seemed a little tongue-tied. Just a few months ago, I was probably feeling worthless and depressed. Instead, I began by snapping a photo of the printed-out sepia-toned image and reading aloud the AI’s description of “two shelves furnished with delicate porcelain and ceramic objects”, such as: a tea cup with a matching saucer, both decorated with a floral motif, a small, bottle-shaped object, possibly a perfume or fragrance bottle, highly decorated with filigree-like patterns, and so on.
The human tongue immediately loosened up due to the AI’s boldness and began correcting, finessing, and flattening the details.
Alabaster suggested that the image was some kind of advertisement. I disagreed. Sure, he had direct access to the visual material in Talbot’s photograph, but I had access to historical context – another way of knowing a photograph. I was convinced that the photo of the tea cups and vases was merely an experiment, taken at a time when almost everything in the world had yet to be captured on camera for the first time.
What seeing loses in immediacy through language is gained with time and attention. Image description is not mystical or compensatory – not about transcendence. Rather, it is a kind of translation – from one sense modality to another. As with all translations, some will be lost and some will be gained.
Last summer I received support for my photography project from the Yaddo Artist Residency, where I met May-Lan Tan, author of the beautiful and complex stories that included her collection. things to make and break. On the way from dinner to artist Clayton Merrell’s studio, I told him about his book and AI-generated image descriptions to interact with photographs. As we were stepping over the threshold, she asked, “Do you want me to be your AI?”
This was the beginning of a friendship that continues today through voice messages on all continents.
When I asked May-Lan what she remembered about Klee’s paintings nearly a year later, the fragments reappeared: a photographically rendered sky fragmented by shady trees and landscapes, interspersed with motifs almost like “Visit or neon signage.”
What I remember best is the experience of walking through energetic space hand in hand with a very good new friend. The conversation emerged with each new detail and question. In a way, we created the images we discuss – and perhaps they created us too. As May-Lan said, paintings are not merely visual or superficial experiences; “They really talk to you—and they look back at us.”
