Delphi Research · Crowdsourcing

Art Beyond Sight

• Ran a 3-phase research study for The Hepworth Wakefield Museum to test whether non-experts in art could write artwork descriptions good enough to make museums accessible for blind and partially sighted visitors. • Used the Delphi method with 18 art professionals and 10 blind/partially sighted people to refine a set of guidelines, then had 23 members of the public apply them to real artworks, then had the original blind/partially sighted participants rate the results. Applied 2 Statistical tests like Friedman's ANOVA and Wilcoxon across 3 phases of the study. • Non-expert describers, working from tested guidelines, produced descriptions rated as genuinely useful (about size, colour, description length). • Delivered as a recommendation for a museum accessibility program.

**Alt text:**  A dimly lit contemporary gallery with cool slate-grey concrete walls and deep shadows. At the center, a large abstract stone sculpture with a circular opening is illuminated by a single warm spotlight, creating a dramatic contrast against the dark interior. A blurred silhouette of a visitor stands nearby with one hand raised toward the artwork. A tall window on the right reveals a dusky outdoor scene with bare trees and a faint warm light in the distance. Large white sans-serif text in the lower-left portion of the image reads, “CAN YOU SEE IT WITH WORDS.” The overall mood is cinematic, quiet, and contemplative, emphasizing accessibility and the challenge of describing visual art through language.

Role

UX Researcher

Timeline

20 weeks

Team

Solo - Supervised by Prof. Helen Petrie and Prof. Michael White

Platform

Qualtrics · Prolific

Role

UX Researcher

Timeline

20 weeks

Team

Solo - Supervised by Prof. Helen Petrie and Prof. Michael White

Platform

Qualtrics · Prolific

Starting Point

Museums call it accessible. It rarely is.

This project was carried out for The Hepworth Wakefield, a modern and contemporary art museum in West Yorkshire. Like most museums, it has made real progress on physical access such as ramps, lifts, accessible routes. But most visual art still isn't described in any meaningful way for blind and partially sighted visitors. Where descriptions exist at all, they're usually minimal alt text, written by whoever had five minutes, with no shared standard for what "good" even means.

The instinct is to solve this with money and specialists: hire professional describers, write bespoke audio guides. But museums don't have that budget, and most artworks in most collections will never get that treatment. So the real question isn't "what does the perfect description look like." It's:

Can non-art experts write art descriptions good enough to actually serve blind and partially sighted visitors? And what would make that possible at scale?

**Alt text:**  Infographic titled **“Three-Study Research Framework”** showing a three-stage accessibility research process arranged left to right. Three connected panels are linked by arrows:  **Study 1 – Expert Evaluation:** 18 art professionals (n = 18) evaluate preliminary guidelines for clarity, relevance, and completeness, prioritize recommendations, and provide qualitative feedback. Output: refined guidelines (v1).  **Study 2 – Public Implementation:** 23 members of the public (n = 23) apply the guidelines to describe selected artworks, submit descriptions through an online platform, and contribute data for analysis. Output: art descriptions dataset.  **Study 3 – User Evaluation:** 10 blind or partially sighted users (n = 10) review and experience the artwork descriptions, rate usefulness, clarity, and engagement, and provide feedback for improvement. Output: validated final guidelines.  A feedback loop labeled **“Iterative Refinement”** runs beneath the three studies, indicating that insights from each stage inform the next to improve accessibility and user experience. The design uses dark blue and gold section headers, simple participant and activity icons, and a clean academic presentation style on a light background.

Study 1

Why Delphi? I needed consensus, not a single opinion.

Before any of the studies, I needed something concrete for people to work with:

A working set of eight short guidelines, plain-language instructions covering things like how to describe size, how to handle colour, how much of your own interpretation to include that anyone, expert or not, could follow to write a description of a painting for a blind or partially sighted person.

Think of it as a style guide for describing what you're looking at, in words that don't assume the reader has ever seen it.

An earlier MSc project (Bai, 2022) had already built a first draft of these guidelines, but it had one major gap: it was tested only on sighted members of the public. Nobody who was actually blind or partially sighted had ever weighed in on whether the guidelines described the right things.

A single round of feedback via a survey, a focus group would have given me opinions. It wouldn't have given me convergence. Art professionals and blind/partially sighted people come at "what makes a good description" from genuinely different places: one group thinks about accuracy and technique, the other thinks about what's actually useful to hear. I needed a method that let both groups react to each other's reasoning, not just register their own.

That's why I used the Delphi method: participants rate and comment on each guideline, I anonymously summarise the reasoning behind the ratings, and the group reconsiders in light of it. The guidelines that have actually been stress-tested against disagreement, not just collected as a list of preferences.

I recruited 18 people working in the art world and 10 blind and partially sighted people, and ran the questionnaire through a first round of ratings and open comments.

Study 2

What I found: Can non-experts actually use this?

With the revised guidelines in hand, I needed to know something Study 1 couldn't tell me: could general public actually apply this in practice? I recruited 23 members of the public through Prolific and asked each to write descriptions for three of six artworks, then reflect on the experience. An example of an artwork image:

**Alt text:**  A screenshot of an online survey or data collection form displaying a portrait painting and instructions for participants. At the top is a seventeenth-century portrait of a woman against a dark background. She faces forward with a calm expression and wears a large, pleated white ruff collar framing her face, a white cap covering her hair, and a dark garment with a reddish-brown bodice.  Below the image, identifying information reads: *“Portrait of a lady with a white ruff. Nicolaes Eliasz Pickenoy. 1640. 49 cm x 42 cm. Oil on panel.”* Further down, participants are instructed: *“Please create a description of this work of art in the box below.”* The screenshot captures the task setup used to collect participant-generated descriptions of the artwork.

The results were genuinely encouraging on output quality. Descriptions averaged 111 words, and adherence to the guidelines was strong: 95.9% of descriptions included objective/subjective balance, and 93.2% used clear, appropriate language. Confidence was reasonably high as 60% of participants felt confident using the guidelines.

But there was a real gap between what people produced and how they felt doing it. About 35% found the task difficult, and colour was the single most cited pain point:

"It was difficult not to use visual terms and to know how to describe colours." P23

"Trying to see this from a visually impaired point of view and not assuming people would know all about colours." P20

The deeper issue wasn't colour specifically, it was empathetic distance. Several participants struggled less with technique and more with imagining an audience they'd never had to write for before:

"It's quite hard describing things so another person can imagine what you are seeing." P9

Guidelines that work in theory can still be too much to hold in your head

A recurring, very practical complaint: the guidelines themselves were too long to hold in working memory while writing.

"I enjoyed this task but had to remember all the things to include. Perhaps in addition to the guidelines there could be a short summary in bullet point form that's easy to glance at?" P21

**Alt text:**  A clean infographic presents a bar chart titled “% Adherence per Guideline Aspect Across Six Artworks.” The chart compares how consistently different artwork description guidelines were followed across six analysed artworks.  Seven vertical bars are arranged from highest to lowest adherence. The highest-scoring category is **Objective/Subjective language** at **95.9%**, followed by **Language** at **93.2%**. **Describing People** scores **84.76%**, and **Colour** scores **84%**. Mid-range adherence is shown for **Medium, Style, and Technique** at **75.13%**. The lowest adherence rates are **Perspective and Composition** at **65.5%** and **Size** at **57.7%**.  Each category is represented by a coloured bar and a corresponding icon beneath it. A red annotation on the right side highlights that **Size** and **Perspective/Composition** are the least-adhered-to aspects. The visual emphasises that while most guideline categories were followed consistently, information about artwork size and composition was omitted much more frequently.

This was the moment the project's core tension became concrete: detailed guidelines produce better descriptions, but they're also harder to hold onto while actually writing. That tradeoff became the thing Study 3 had to help resolve.

Study 3

What I found: What blind and partially sighted readers actually valued

This was the study that mattered most, because it closed the loop: the same blind and partially sighted participants from Study 1 came back to rate real descriptions written by real non-experts from Study 2. The guidelines were now of three lengths (short, medium, long) per painting, three paintings.

I used a Related-Samples Friedman's Two-Way Analysis of Variance to test whether length preference differed significantly within each painting, then Wilcoxon signed-rank tests to pin down exactly which pair of descriptions differed.

The statistical pattern was more interesting than "longer is better":

**Alt text:**  A presentation-style infographic titled **“How Description Length Affected Ratings Across Three Artworks”** compares participant ratings for short, medium, and long artwork descriptions across three different paintings. The layout is divided into three vertical columns, each dedicated to one artwork: a portrait by Pickenoy, a landscape by Goldberg, and a still life by Scott.  The **left column** features a seventeenth-century portrait of a woman wearing a large white ruff collar against a dark background. Beneath the image, a bar chart shows average ratings of **4.8** for the short description (61 words), **7.2** for the medium description (170 words), and **6.9** for the long description (228 words). A summary notes that medium and long descriptions were rated similarly, while both performed better than the short version.  The **middle column** shows a landscape painting of a village viewed from a country lane lined with stone walls and trees. Its bar chart shows ratings of **5.6** for the short description (61 words), **7.6** for the medium description (143 words), and **5.1** for the long description (214 words). A summary highlights that the medium-length description received the highest ratings and significantly outperformed the longest version.  The **right column** displays a still life painting of green leaves and small purple flowers arranged in a blue vase against a bright green and orange background. The accompanying chart shows ratings of **4.2** for the short description (61 words), **6.0** for the medium description (132 words), and **7.4** for the long description (196 words). A summary notes that the longest description was rated significantly better than the shortest.  The overall visual demonstrates that preferred description length varied by artwork: medium-length descriptions worked best for the portrait and landscape, while the longest description was preferred for the still life. The figure emphasizes that there is no single optimal description length across all artworks.

Longer wasn't always better but detailed always was

In other words: the right length depends on the artwork, not on a fixed rule. The effect was statistically real, but it pointed in a different direction for each painting. What stayed constant across every preferred description was specific, well-organised detail, not word count for its own sake.

Participants were remarkably precise about what made a description work. Perspective came up unprompted, again and again:

"The exciting part of the description is again perspective. It's impossible for me to understand how this can be conveyed in a picture, but here, it was described perfectly." VI1

"I got a feeling for where we were viewing from and an appreciation of where we were in the room." VI2

Shade and colour detail, when specific, was valued rather than skipped past which directly contradicted the assumption (visible in Study 2's participant comments) that colour description might not matter to this audience:

"Having info on the perspective is very helpful. Also giving measurements makes it clearer and easier to understand. Having shade information is also good e.g. light blue table." VI3

And the objective/subjective balance turned out to be exactly what blind and partially sighted readers were listening for:

"I liked the topographical positioning of objects. This, for me, is quite important. The person describing this picture was objective, which again, I liked." VI1

Even word choice within a single sentence registered:

"The description of the clothing. Also, terminology. In the first description, she is described as smiling. In this one, it's a grin." VI1

Reflection

So what this means for museums?

  • Optimise guidelines for scannability. Study 2 showed people didn't fail because the content was wrong, they failed to hold onto it mid-task. A one-page reference card alongside the full guidelines would likely close much of that 35%-difficulty gap without cutting a single guideline.

  • Let a non-sighted intuition decide what to keep. The clearest miss in this project was size comparisons. A guideline that read as helpful to sighted reviewers and split blind participants right down the middle. Any guideline aimed at a specific audience needs that audience in the room before it ships, not after.

  • Treat "objective vs. subjective" as a feature. It was the most confusing guideline to sighted describers and the most valued trait to blind and partially sighted readers. That gap is worth designing around directly, to give describers permission and structure for including their own reaction, rather than hedging it as optional.

  • Flexible description length. What "worked" varied by painting, not by a general preference for long or short. A flexible length, guided by the artwork's own complexity, will likely outperform a house style that mandates one word count for every piece.

Non-experts weren't the risk. Untested assumptions were.

Going in, the implicit worry behind this whole project was that the accessible description was something you needed art expertise to do well. That wasn't what the data showed. Non-experts, working from well-tested guidelines, produced descriptions that blind and partially sighted participants rated as genuinely useful, and prior experience with visually impaired people had no measurable effect on description quality.

What actually put descriptions at risk wasn't the describer's expertise. It was guidelines shaped by sighted intuition about what blind and partially sighted people would want. The intuitions on aspects like size and colour, turned out to be wrong or at best incomplete until tested directly against the people who'd be listening.

That reframes the design question for museums thinking about scale. It's not "how do we find and train enough expert describers." It's "how do we build a system good enough that anyone who cares can contribute meaningfully". The only way to know if a system is good enough is to close the loop back to the people it's actually for, the way Study 3 did here.

The guidelines that came out of this project were delivered as a recommendation for a museum description program. There's a lot more behind the numbers than a scroll can hold. If you want the deeper cut, I'm one email away.

ux.megha@gmail.com

Megha Upadhyaya

UX Researcher trained in rigor and drawn to work where research leads to what ships.

Contact

ux.megha@gmail.com

Megha Upadhyaya

UX Researcher trained in rigor and drawn to work where research leads to what ships.

Contact

ux.megha@gmail.com

Megha Upadhyaya

UX Researcher trained in rigor and drawn to work where research leads to what ships.

Contact

ux.megha@gmail.com

Create a free website with Framer, the website builder loved by startups, designers and agencies.