Skip to main content
OpenTrials
Completed

NCT Number: NCT07020728

Clarity in Motion Phase-1 Perceptual Study of Speech Intelligibility

Participants received a bilateral pure-tone hearing screen administered by the research team. All potential participants who failed the hearing screen were provided with information about its meaning and referral for further audiological testing.

Participants who passed the hearing screen and other inclusion criteria were divided into 6 groups, each of which were presented with 144 stimuli equally distributed among processing conditions. Listeners choose a comfortable listening level using supplied headphones and were able to control the rate of presentation. Following a short practice session, listeners were be asked to transcribe each target sentence. The intelligibility of each stimulus was estimated by determining the mean percentage of content words correctly transcribed. After transcription, listeners were asked for two qualitative judgments: (1) the "clarity" of the stimulus, and (2) the "listening effort" involved. The quality of each stimulus was estimated by the median quality judgment, and the effort likewise. Listening sessions were located in a quiet room and presentation was controlled by the Superlab presentation software program.

The Stimuli consisted of audio recordings of target spondaic words embedded in a carrier sentence produced by a male and a female native speaker of American English recorded under quiet conditions. Each stimulus presented to the listeners for identification was either unmasked pristine speech or speech that had been processed in one of five ways with different mixtures of noise and sensor movement. The latter are identified as QoS Levels 1-5.

Collectively, the estimates of word intelligibility, clarity, and listening effort under the different conditions shed light on the effectiveness with which the tested algorithm preserves listener intelligibility with acceptable effort and quality.

Completed

Looking for future studies?

Notify Me

Key information

Conditions

Age range

18 year and older

Sex eligibility

All sexes

Study type

Interventional

Phase

Not applicable

Primary location

University of Cincinnati

Cincinnati, Ohio, 45221, United States

About this study

Participants received a bilateral pure-tone hearing screen administered by a clinically trained member of the research team. The threshold criteria was 20 dB SPL at 250 and 500 Hz, and 25 dB SPL at 1000, 2000, 4000, and 8000 Hz. All potential participants who failed the hearing screen were provided with information about its meaning and referred for further audiological testing.

Participants who passed the hearing screen and other inclusion criteria were divided into 6 groups, each of which was presented with 144 stimuli equally distributed among processing conditions (Pristine Non-Moving Speech plus QoS Levels 1-5). Listeners self-selected a comfortable listening level using supplied headphones and were able to control the rate of presentation. Following a short practice session, listeners were asked to transcribe the target sentences, and the intelligibility of each stimulus was estimated by determining the mean percentage of content words correctly transcribed. After transcription, listeners were also asked for two qualitative judgments using a visual analog scale: (1) the "clarity" of the stimulus, and (2) the "listening effort" involved. These measures are sensitive to situations where listeners manage to extract the uttered words from the signal, but with increasing difficulty. The quality of each stimulus was estimated by the median quality judgment, and the effort likewise. Listening sessions took place in a quiet room and presentation was controlled by the Superlab presentation software program.

The Stimuli consisted of audio recordings of target spondaic words embedded in a carrier sentence produced by a male and a female native speaker of American English recorded under quiet conditions. Each stimulus presented to the listeners for identification was either unmasked pristine speech or speech that had have been processed in one of five ways with different mixtures of noise and sensor movement. The latter are identified as QoS Levels 1-5. Each type of processing was expected to have a different effect on the underlying probability that the listener would be able to correctly identify the spoken word, the effort required to do so, and the quality of the presented recording. Data on familiarity ratings for the target spondaic words and relative intelligibility of the two speakers under different conditions of noise masking has been previously reported.

Collectively, the estimates of word intelligibility, clarity, and listening effort under the different conditions is expected to shed light on the effectiveness with which the tested algorithm preserves listener intelligibility with acceptable effort and quality.

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Inclusion criteria will include adult native speakers of American English with auditory thresholds within age-normal limits, defined as passing a pure tone screening test of 20 dB SPL at 250 and 500 Hz, and 25 dB SPL at 1000, 2000, 4000, and 8000 Hz. Exclusion criteria will include self-report of hearing difficulties and failure to pass hearing screen.

No special populations requiring special protections will be utilized in this study. Potential participants who fail the hearing screen will be provided with results and appropriate clinical referral as per the IRB protocol.

Exclusion criteria

  • Exclusion criteria will include self-report of hearing difficulties and failure to pass hearing screen.

Treatment and study plan

Solo: Unmasked Speech Stimuli

Behavioral

Speech stimuli recorded using non-moving speakers and mics. No masking sources present. No BSS applied to multi-channel recordings. Very high output QoS values.

Other names: Treat1

Raw: Fully masked speech--no motion stimuli

Behavioral

Speech stimuli recorded using non-moving speakers and mics. All masking sources present. No speech separation or extraction methods applied to multi-channel recordings. Very low output QoS values.

Other names: Treat2

StatScrub: Extracted Speech--no motion stimuli

Behavioral

Speech stimuli recorded using non-moving speakers and mics. All masking sources present. Joint ACES scrubbing of both noise sources applied to multi-channel recordings. Very high output QoS values.

Other names: Treat3

SlideSpch: Scrubbed Speech emitted from linearly moving speaker stimuli

Behavioral

Speech stimuli recorded using linearly moving speech source and stationary masking sources and mics. All masking sources present. Joint ACES scrubbing of both noise sources applied to multi-channel recordings. Moderately high output QoS values.

Other names: Treat4

SlideNoise: Speech Scrubbed from linearly moving and stationary noise stimuli

Behavioral

Mixed speech and noise sources recorded using a stationary speech source, a stationary noise source, and a linearly moving noise source. A valid source hypothesis of the speech source is used to extract the speech source. High output QoS values.

Other names: Treat5

SlideMic: Stationary sources scrubbed from a linearly moving mic stimuli

Behavioral

Mixed speech and noise sources recorded using all stationary sources, and a linearly moving microphone (mic 1). Joint ACES scrubbing of both noise sources is used to reduce the response of Mic 1 to a residue of speech. Low output QoS values.

Other names: Treat6

Primary outcomes

  1. Target Word Transcription Accuracy

    Time frame: The subject transcribed the target word immediately after hearing the carrier phrase.

    For each audio stimulus of a target word in a carrier phrase, the subject transcribed the target word they heard (by typing that word in a field in the display. Their transcription was later graded as correct or incorrect by a researcher.

  2. Listening Effort

    Time frame: Immediately after hearing the carrier phrase.

    Subject-reported effort required to identify word in carrier phrase, reported by moving a slider on a 0% to 100% scale.

  3. Naturalness of speech

    Time frame: Immediately after hearing the carrier phrase.

    Subject-reported perceived naturalness of target word in carrier phrase, reported by moving a slider on a 0% to 100% scale.

Sponsors and collaborators

Lead sponsor

Speech Technology and Applied Research Corp.

Industry

Collaborators

  • National Institute on Deafness and Other Communication Disorders (NIDCD)
  • University of Cincinnati

Registry information

Official study title

Speech Intelligibility and Listening Effort

Acronym: CIM PhI PercX

Important dates

Study start
2024
Primary completion
2024
Study completion
2024
First posted
Jun 13, 2025
Registry last updated
Jun 25, 2025

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.