House

Captioning and audio description as production decisions

Two different jobs for two different audiences, sharing equipment and almost nothing else, and each needing its own channel.

Production teams applying this principle can also compare practical guidance on hours tracker, keeping time and activity records separate from the artistic and technical judgement they are meant to inform.

Captioning and audio description are frequently discussed together, budgeted together and delivered on the same hardware, and they are different jobs for different audiences with almost nothing in common except the equipment. Treating them as one line item is how productions end up doing neither properly.

Captions are not subtitles

The article on surtitles deals with translation, which assumes the audience can hear. Captioning assumes they cannot, and that changes the content rather than only the language.

A caption identifies who is speaking, because a listener who cannot hear cannot tell two offstage voices apart. It carries sound that carries meaning: a door, a gunshot, an approaching train, the fact that music has become ominous. And it does not condense as aggressively, because the text is the only channel available rather than a support for one.

Open, closed, and the trade nobody escapes

Open captions appear on a display everybody can see. They are reliable, need no device, cannot be left uncharged, and cannot be refused by somebody who does not want to identify themselves at the box office. They are also visible to the whole house, which productions sometimes resist and audiences generally stop noticing.

Closed captions go to individual devices, seatback units or glasses. They are unobtrusive and they reintroduce every operational failure described in the article on assistive listening: charging, distribution, staff briefing, and a patron who has to ask.

Placement follows the same arithmetic as titles. The higher and further from the stage the text sits, the more of the performance is missed reading it, and a caption user is reading continuously rather than occasionally.

Prepared and live are different products

A scripted play can be captioned in advance and cued like titles, with accurate text and controlled timing. Unscripted material needs a live captioner working in real time, and that introduces both delay and error.

Automatic speech recognition performs worse in a theatre than almost anywhere else, and the reasons are structural rather than a matter of maturity: reverberation, music under dialogue, overlapping speech, invented proper nouns, heightened or archaic language, and accents used deliberately. Any assessment made in a quiet office does not predict behaviour in a hall.

Assistive listeningthe programme soundno added wordsfor a hearing aid or receiverCaptioningevery word, as textwho is speakingsound that carries meaningAudio descriptiona voice in the gapsentrances, action, costumewritten, and therefore authoredone distribution system, and each needs its own channel
Figure 1Three services with different audiences and different content, usually sharing one distribution system that has to be specified for all three from the start.

Audio description is a written work

A describer speaks into a channel heard only by users of the service, fitting description into the gaps between lines. What goes in those gaps is a decision: an entrance, a costume, a look between two characters, a change of light. It cannot all fit, so somebody chooses.

That makes description authored rather than transcribed, and it means a describer needs the production, not just the script. Watching a run, writing against it, and rehearsing the delivery is the work, and a describer arriving on the day produces a commentary rather than a description.

The associated practice of a touch tour, in which patrons come early to handle props, walk the set and meet the cast in costume, does work no amount of spoken description can, because it establishes scale and material directly. It costs stage time and it is frequently the part audiences mention.

The describer needs a position, and it is not the last one left

Description is delivered live, which means somebody is speaking throughout the performance from somewhere in the building. That position needs an unobstructed view of the whole stage, because a describer who cannot see an entrance cannot describe it, and it needs acoustic isolation, because a voice audible to the surrounding seats is a distraction rather than a service.

In houses built before any of this existed, the position is usually improvised, and the improvisation is usually a corner with a partial view. Where a booth cannot be provided, a remote position fed by a camera is workable and imports the latency described elsewhere in this publication, which has to be measured rather than assumed.

Three services, three channels

Assistive listening carries the programme sound. Description carries a different voice mixed with the programme sound. Translation, where it exists, carries something else again. These cannot share one channel, and a distribution system specified for one service will not deliver three.

This is a specification decision made when the equipment is bought, and discovering it when a described performance is scheduled is discovering it too late.

The scheduling question is the equivalence question

Most companies offer captioned and described performances on selected dates. That is understandable and it means a patron who needs one has a choice of two evenings while everyone else has a choice of thirty, and cannot easily attend with a group that picked a different night.

Increasing the number of such performances costs money. Saying plainly how many there are, early enough to plan around, costs nothing and is frequently what is missing.

What we cannot verify

Accuracy claims for speech recognition and captioning products come from their vendors, measured on material that does not resemble a performance, and we reproduce none. Practices for description differ between countries and between companies, and no standards body governs the editorial content of a description. Whether a particular provision works is answered by asking the people who use it, which is a step frequently skipped.

The short version

  1. Captions name the speaker and carry meaningful sound; translation assumes hearing.
  2. Open captions are reliable and visible; closed ones are discreet and fail operationally.
  3. A caption user reads continuously, so placement costs more than it does for titles.
  4. Speech recognition performs worse in a theatre than almost anywhere, for structural reasons.
  5. Description is authored, needs the production rather than the script, and is rehearsed.
  6. Three services need three channels, and that is decided when equipment is bought.

Further context

For a primary, standards or institutional reference, see the W3C Web Content Accessibility Guidelines.