Opens in a new tab
Loudness Project

Dialogue Intelligibility

Why are consumers increasingly using subtitles?

Consumer surveys show a rising use of captions or subtitles, with some estimates approaching 50% of viewing. As a press article stated:

“Whether you have lousy TV speakers, are hard of hearing, are distracted by the kids, or are watching a film with actors who mumble, chances are you are using the captions option while watching TV…”

Consumers use subtitles for various reasons, from avoiding disturbing others in public to a tool in learning another language.

Bar chart showing subtitle use most of the time by generation in 2022: Baby Boomers about 35%, Gen X 38%, Millennials 53%, Gen Z 70%, and 50% overall.

However, in other cases consumers are using captions because the dialogue is not intelligible. When listeners struggle to understand what is being said, they may become annoyed and disengage from the content. While subtitles may help with viewers following the story of the program, they detract from the quality of the visual images and engagement with the story. The best viewing experience is enjoying a great story without significant listening effort.

Intelligibility is affected by many human, technical, and creative factors and has been the subject of an AES Broadcast and Online Delivery working group studying the issue. Technical Document AES TD 1009: “Improving Dialogue Intelligibility in Media” is the result of this work. A summary of its points follows below:

What is Dialogue Intelligibility?​

Speech intelligibility is a continuum between effortless understanding and total unintelligibility regardless of listening effort. Intelligibility is affected by:

  • The speaker’s delivery and listener’s familiarity with what is said.
  • Degradation or masking of the speech signal during acoustic and electronic transmission.
  • The listener’s hearing ability and fluency with the language.

Masking

Elements of speech such as phonemes or speech cues can be masked by interfering sounds, both in the content and in the listening environment. Competing sounds in the same frequency range as a given speech element are most able to mask the speech signal. Masking decreases as the frequency separation between the competing sound and the dialogue increases. This phenomenon is asymmetrical with respect to frequency. With typical program material, a commonly troublesome scenario is heavy bass masking higher frequency dialogue elements.

Graph showing that higher masker levels produce stronger and broader auditory masking across frequencies.
Increased hearing thresholds of a pure-tone signal masked by 410 Hz narrow-band noise increased in level. Tones under the curve are masked. Adapted from Egan and Hake, as reprinted in Moore, Psychology of Hearing, 5th ed., p. 88.

Loudness

Loudness refers to the perceived level of a program, whereas dialogue intelligibility refers to how much of the dialogue is understood by the listener. Two programs that have the same loudness may differ in intelligibility. 

Loudness management may affect intelligibility when there are quiet passages in the dialogue or the dynamic range is too large, leading to undesirable consumer “volume-riding”. Loudness normalization to a uniform level at playback helps maintain intelligibility by avoiding a listener having to adjust the volume.

Listening Effort and Cognitive Load

Understanding dialogue involves both bottom-up processing and top-down processing. 

Top-Down Processing: Listeners using context along with their experience and expectations to help understand speech that cannot be understood by bottom-up processing alone.

Bottom-Up Processing: Speech understanding that occurs subconsciously and effortlessly. Typically, listeners will subconsciously augment it with top-down processing.

Bottom-up processing of the phonemes of speech occurs subconsciously and effortlessly. When bottom-up processing is incomplete, top-down processing is used by the listener to “fill in the blanks” by considering context, grammar, and expectations of what is said. Top-down processing requires mental effort, or cognitive load, that compromises the listener’s ability to comprehend or follow the story. Lack of clear speech, due to masking in the mix or deficiencies of the playback device or a noisy environment, will increase listening effort and may lead to listening fatigue.

Hearing Ability

Comprehension depends on a listener’s cognitive and hearing ability. Young children and the elderly generally do not understand as well as adults with normal hearing acuity. Listeners may have hearing loss due to age or noise exposure or may be listening to a program not in their native language. Because of this, creators with normal hearing can fool themselves into thinking that their productions are intelligible to all audience members.

The Dialogue Intelligibility Ecosystem

The Dialogue Intelligibility Ecosystem diagram below shows the many points in the content delivery chain where intelligibility issues can arise as content moves from production to distribution to the consumer. 

(Click to pan, double-click to zoom, touch two fingers to zoom/pan, or use control buttons in lower right corner)

There are multiple points in the chain where dialogue intelligibility may be impaired, including:

  • Speech may not be articulated clearly
  • Accents can be hard to understand
  • Dense soundtracks mask dialogue
  • Content dynamic range is excessive 
  • Noise in the listening environment masks dialogue
  • Non-optimal reproduction
  • Hearing impairment 

Thus, loss of intelligibility often does not have a single cause, but rather a combination of factors.

Issues in Professional Practice – Content Creation

On-stage performance and recording

Content for home playback has evolved over the last two decades in ways that can have a negative impact on dialogue intelligibility:

  • Acting styles have become more subdued and dialogue performances have become more “natural”, typically quieter. 
  • Television content has become much more cinematic, increasing soundtrack complexity and dynamic range. 
  • In some cases, dialogue may be mixed for lower intelligibility due to artistic intent, or actors may have strong accents or mumble. 

Mixing

Content mixed in larger rooms at high levels (SMPTE RP200, A/85), as with contemporary cinema soundtracks, typically results in content with a wider dynamic range. These large rooms are frequently used for content intended for home viewing as well. This wide dynamic range content typically does not translate well when played back in home environments that have higher background noise and less controlled acoustics. 

Studies have shown that the median home listening level is 58 dBA sound pressure level (SPL) for TV speakers and 65 dBA SPL for home theater/soundbars, as shown in the figure below. Variations in dialogue level and intelligibility will be more pronounced at a lower playback SPL because quieter passages will be masked by noise in the home environment.

Bar chart showing preferred television listening levels, with most participants selecting 58 dBA.

Most consumers listen in stereo, and unless the production budget allows for a separate manual/artistic stereo mix, the consumer will typically hear an automatic stereo downmix. Unless this downmix is checked during the mixing process to ensure there is no masking of the dialogue, it could introduce intelligibility issues.

Dynamics Processing in Mixing

The ratio of dialogue to background noise will impact listening effort, so processing tools that reduce noise or enhance dialogue can be beneficial. However, misuse of these tools (e.g., noise reduction, de-reverberation, compression/limiting, expansion) impairs intelligibility and causes listener fatigue. 

Tools that apply downward expansion (such as noise reduction and de-reverberation) can render quieter components of words inaudible, degrading intelligibility. While compression can be a useful tool to improve intelligibility of a performance (particularly at low listening levels preferred by many consumers), over-compression or limiting of dialogue can sound unnatural and distracting. Furthermore, if carelessly applied to the full mix, dynamics processing can have a negative effect on dialogue clarity when other high-energy mix elements compete and cause unintended ducking. Misuse of compression and limiting can be detrimental from the initial stages of production through to the final sound mix of post-production and during distribution. [ref Issues in Professional Practice – Distribution]

Additionally, unnecessarily low peak limiting thresholds have proved to be problematic because they may sacrifice the impact of the dialogue and compromise audio quality.

Mixing to Loudness Specifications

Dialogue-loudness-based normalization of the content has proved to be the most effective means of maintaining dialogue intelligibility and consistency between different programs and channels. 

Distributors and broadcast regulators have set requirements for loudness, program-to-dialogue loudness ratio, and loudness range (LRA). Often these requirements are misapplied or misunderstood.

Mixing to specific loudness range (LRA) targets is discouraged due to unintuitive results when modifying the mix.

Checking Content for Intelligibility

During quality checking (QC), it is beneficial for dialogue intelligibility to be assessed by a native speaker unfamiliar with the script. This should be performed with subtitles / captioning turned off.

The mix should be checked at consumer listening levels on loudspeakers having approximately flat frequency response. Automatic or manual stereo downmixes should also be checked as there could be additional masking due to downmixing of the other channels. Background noise at 35 to 40 dBA SPL, such as from a fan or air cleaner, can help to simulate consumer listening.

Emerging metering tools [ref “Monitoring Tools to Estimate Intelligibility”] may be able to identify segments of the content needing human review for potential intelligibility issues.

Issues in Professional Practice – Distribution

Distribution channels to the consumer may affect intelligibility. Primary issues are differences in loudness of various content items and sources, the application of real-time loudness processing, and the lack of dynamic range control metadata. The provision of stereo downmixes is also a factor. To ensure consumers receive content with consistent loudness, distributors may loudness normalize content received from creators or transmit its loudness metadata using a metadata-based codec.

Real-time processing

For live linear content requiring a loudness adjustment to match a target, an LKFS-based automatic loudness control (ALC) processor can be used. ALC can be used in real time. However, because it changes the dynamic range of the content to achieve a result, creative intent is altered. Even so, these processors can be quite effective when set up by an expert using typical source content.  Problems with real-time processing can stem from at least two causes. First is using processors that are not purpose-built for unattended, real-time processing of soundtracks, such as studio and production-targeted compressors without gain-freezing silence gating.  Second is incorrect setup of purpose-built processors. TD1009 offers detailed recommendations. 

Dynamic Range Control

Audio codecs [pop-up: This document primarily addresses video content. Audio-only content may or may not be transmitted with audio codecs or systems that support loudness normalization or DRC.] may include a Dynamic Range Control (DRC) feature that can boost portions of the content so that quiet dialogue is heard above the ambient noise in the user’s environment. It also reduces uncomfortably loud portions that would otherwise require the user to lower the playback volume, possibly compromising dialogue intelligibility. Codec-based DRC does not alter the encoded audio. Gains and control parameters are carried as metadata that can be applied by the user during decoding. When active, codec-based DRC may improve dialogue intelligibility for users in challenging listening environments, while remaining inactive for users who wish to hear the content unmodified.

Downmixing

5.1 surround or immersive content must be downmixed to stereo for playback on consumer devices that have stereo speakers, such as TV sets. According to major streaming distributors, 90% of consumers listen over their TVs in stereo. Thus, they will always be listening to a downmix.  This can happen in several ways:

  • Automatic downmixing done in the audio decoder of the playback device. In most broadcasting, this is the only method possible.
  • Creation of a stereo mix by the content provider, which is sent as a separate content stream to the consumer. This is the preferred method if available. 
  • Delivery of a supplemental stereo downmix with boosted dialogue, which is sent as a separate stream to consumers with impaired hearing or difficult listening environments. 

Downmixing may affect intelligibility due to auditory masking. [ref Masking heading]  One form of masking occurs because all the content comes from the stereo speakers. This prevents binaural separation of the surround or overhead content by the listener. Another form is the addition of surround or overhead components to the stereo channels, which may increase the level of other program elements such as music or sound effects relative to the dialogue. Of course, legacy or secondary content produced only in stereo will not incur any additional masking. Automatic downmixing may be controlled to produce either a Lt/Rt [pop-up: >Lt/Rt is a stereo downmix specified in ATSC A/52, Section 7.8.2, and is intended to be used by a matrix surround decoding process (e.g., K. Gundry, “A New Active Matrix Decoder for Surround Sound,” in Proc. AES 19th Int. Conf., June 2001).] or a Lo/Ro [pop-up: Lo/Ro is a stereo downmix specified in Rec. ITU-R BS.775.] downmix during audio decoding.  The Lo/Ro downmix without a 90° phase shift is preferred for intelligibility.

Issues in Consumer Playback

Limitations of Consumer Devices

TV set sound quality has become secondary to cost and aesthetics. Sets have evolved from cathode ray tube (CRT) displays with front-firing speakers to thin flat-panel displays with narrow bezels and hidden rear or down-firing low-excursion speakers that rely on surface reflections. Often these speakers have limited frequency response and flatness.

Photo comparing the speaker assemblies from two televisions, showing the larger speakers from the better-performing TV.
Downward-firing speakers similar to the best and worst (left) examples in the figures below

Speakers are often controlled by driver circuits that monitor voice coil temperature and/or coil displacement indirectly to maximize output. Speaker signal processing will likely also be used to increase the loudness. This induces dynamic compression or limiting of the playback signal at high playback levels, compromising intelligibility.

Bar chart showing the distribution of preferred television listening levels among study participants. The most common preferred level is 58 dBA (11 subjects), with fewer participants preferring levels between 50 and 68 dBA. The chart is titled "Distribution of preferred listening levels for television viewing" and cites Benjamin, AES Paper 6233.
Frequency response of the best and worst (blue dashed line) 2023 TVs measured by rtings.com

Manufacturers have responded to consumer complaints with post-processing intended to improve dialogue intelligibility. This may improve or degrade intelligibility and sound quality which may affect creative intent.

Some consumer devices have an automatic room equalization feature. This can help dialogue intelligibility by improving the frequency response at the listening position.

Scatter plot showing reverberation times in consumer listening rooms are generally consistent across frequency.

The consumer environment often impacts dialogue intelligibility.  

Reverberation times of consumer listening environments are typically in the range of 400-600 ms, likely much higher than those of a near-field mixing room. This may lead to time masking of dialogue.

Consumers routinely have noise sources—such as dishwashers, clothes dryers, HVAC systems, and other occupants—that can mask audio content. They also often face limits on maximum playback level because of equipment constraints or the objections of family or neighbors. These factors can restrict usable dynamic range, causing dialogue to be masked by environmental noise, or leading viewers to keep the volume too low while “volume-riding.”

Understanding signal flow

Intelligibility can be affected by the complex setup options presented by multiple consumer devices in the playback chain. As the DRC metadata in the audio bitstream is optionally applied during decoding, the DRC and processing options available to the consumer may depend on where the audio is actually decoded.

The consumer audio signal path can be quite complex, and TD1009 (Section 6.3.4) describes some common examples. Often the best option for intelligibility is to use Bitstream Passthrough to a soundbar or audio-video receiver (AVR), if possible.

Television audio settings menu showing HDMI output configured for Primary Pass-Thru.
Consider selecting “bitstream passthrough” on your TV and/or source device

Consumer Quality Hierarchy

The best TV dialogue intelligibility is obtained with external speakers, connected to an audio-video receiver (AVR). Some premium TVs and soundbars can be networked with satellite speakers, making this very convenient.  

AVR systems can be costly. A close level of performance can be obtained with a premium soundbar with more than three channels. 

 A lower-cost soundbar with at least a Left–Center–Right speaker arrangement will likely improve intelligibility. Even a stereo soundbar may provide better intelligibility than a typical TV set. Note that inexpensive soundbars will be stereo, even if they support surround or immersive decoding.

Consumer Recommendations

  • Consider buying a soundbar, preferably with a center speaker.
  • Experiment with TV settings such as DRC (“night mode”). DRC will help avoid “volume-riding”.
  • Try turning TV post-processing on/off. It may help or hurt.
  • Use room measurement/correction if available.
  • Investigate if the content offers other mixes, such as a manual stereo mix or one with boosted or clean dialogue  [ref Clear or Boosted Speech Streams] .
  • Consumers should be aware that playback systems may offer multiple output modes and settings. They may need to experiment to identify which options are active in a given situation and which provide the best intelligibility. Sending bitstream audio directly to the final playback device may be a good choice.

Emerging Solutions to Consider

Chart comparing methods of increasing dialogue audibility using dialogue boost, higher playback levels, and dynamic range compression.
Methods of boosting dialogue level

Clear or Boosted Speech Streams

There are two common methods available to allow consumers to adjust dialogue level with respect to other elements. The first is using the dialogue object/channel capabilities of a Next-Generation Audio (NGA) codec to allow consumers to fine-tune the dialogue level. The second is to provide one or more isolated or boosted dialogue streams that the consumer can choose.

  • The first method requires use of an NGA codec where dialogue can be conveyed as a separate stream that allows consumers to adjust the dialogue level either directly or by presets. This requires dialogue to be available as an independent stem (sub-mix) from production or to be separated automatically during the encoding process. 
  • The second method requires distributors to prepare an isolated dialogue or boosted dialogue version of the content for use in cases of poor intelligibility or for users with hearing impairment. As above, this may be done with dialogue stems from production or with automatic dialogue separation. The advantage of this method is that distributors have more control over the mix, including processing of the dialogue stem.
  • An alternative method is for consumers to adjust the playback level of the center speaker in their audio-video receiver (AVR). This is the least desirable method since the center channel may contain other program elements, and dialogue may be present in other channels.

Monitoring Tools to Estimate Intelligibility

Several tools and metrics have been developed and proposed to identify content sections needing human review for potential intelligibility issues. One example is an automatic speech recognizer that uses a deep neural network to estimate the listening effort needed to correctly perceive individual phonemes. [pop-up: R. Huber, H. Baumgartner, S. Goetze, J. Rennies, “ASR-Based, Single-Ended Modeling of Listening Effort—A Tool for TV Sound Engineers,” Forum Acusticum, pp. 2441-2445 (2020 December).] Another separates dialogue and then measures the short-term loudness ratio between dialogue and background sound. Other methods of estimating intelligibility exist or are under development.

Listening on Mobile Devices

Tablets and mobile phones have issues that are an order of magnitude worse than those of TV sets. Speakers may be only 10-12 mm in diameter and have a 0.5 mm maximum displacement. The frequency response of a typical tablet or phone rolls off below 700 Hz due to acoustic effects. Active feedback drivers are necessary for tolerable output SPL.  Compounding the problem, these devices are often used in environments that have more ambient noise than the home TV room. Speaker playback in this case can require both dynamic range control (DRC) and dialogue boost for successful intelligibility. This may already be performed on the device. Earphones and headphones (particularly in-canal or noise-cancelling ones) can be superior to mobile device speakers. They reduce the ambient noise, improve maximum SPL, avoid the need for compression or limiting, isolate listening from the environment, and may improve frequency response and reduce distortion. These advantages will benefit dialogue intelligibility. Binaural rendering technology will improve immersion and envelopment, and can affect intelligibility positively or negatively.

To Learn More

AES TD 1009: “Improving Dialogue Intelligibility in Media”  available on this website, offers much more detail and recommendations on each of these topics.