Skip to main content
OpenTrials
Completed

NCT Number: NCT07835347

Large Language Models for Stuttering Assessment and Therapy

This observational study aims to compare the quality of responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20) to questions about stuttering. The study focuses on three main areas: general information about stuttering, clinical assessment, and therapy. A total of nine questions were developed based on evidence-based clinical practice guidance, including the American Speech-Language-Hearing Association (ASHA) Practice Portal. Each question is presented to each language model in separate sessions, and the generated responses are recorded for evaluation.

Five speech-language therapists with clinical experience in stuttering independently evaluate the model-generated responses. Each response is rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale. The responses are also compared with guideline-based reference information. The study does not involve any clinical intervention or treatment of patients. The aim is to determine how closely large language model responses align with current clinical practice guidance and to identify their potential strengths and limitations when used as informational or clinical support tools in the field of stuttering.

Completed

Looking for future studies?

Notify Me

Key information

Age range

18 year and older

Sex eligibility

All sexes

Study type

Observational

Primary location

Istanbul Gelisim University

Istanbul, 34310, Turkey (Türkiye)

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Having an undergraduate or postgraduate degree in speech and language therapy.
  • Actively practicing clinically in the field of stuttering/fluency disorders.
  • Voluntarily agreeing to participate in the study.
  • Completing the evaluation form in full.

Exclusion criteria

  • Not having active clinical experience in the field of stuttering.
  • Incomplete completion of the evaluation form.
  • Withdrawal of voluntary participation during the evaluation process.

Treatment and study plan

ChatGPT-Generated Responses

Other

ChatGPT responses to nine standardized questions on stuttering information, assessment, and therapy were generated in separate sessions and recorded for independent expert evaluation. Questions were repeated under comparable conditions to allow assessment of response consistency.

Claude-Generated Responses

Other

Claude responses to nine standardized questions on stuttering information, assessment, and therapy were generated in separate sessions and recorded for independent expert evaluation. Questions were repeated under comparable conditions to allow assessment of response consistency.

Google Gemini-Generated Responses

Other

Google Gemini responses to nine standardized questions on stuttering information, assessment, and therapy were generated in separate sessions and recorded for independent expert evaluation. Questions were repeated under comparable conditions to allow assessment of response consistency.

Grok-4.20-Generated Responses

Other

Grok-4.20 responses to nine standardized questions on stuttering information, assessment, and therapy were generated in separate sessions and recorded for independent expert evaluation. Questions were repeated under comparable conditions to allow assessment of response consistency.

Primary outcomes

  1. Expert-Rated Quality of Large Language Model Responses

    Time frame: During the single cross-sectional expert evaluation conducted over approximately 1 month

    Responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20) to nine standardized questions about stuttering were independently evaluated by five speech-language therapists with clinical experience in stuttering. Each response was rated using a 5-point Likert scale (1 = strongly disagree to 5 = strongly agree) across five criteria: relevance, accuracy, clarity, completeness, and consistency. Responses were evaluated against guideline-based reference information derived from evidence-based clinical practice guidance. Higher scores indicate better response quality and closer alignment with clinical practice guidance. Mean scores were calculated for each large language model and each evaluation criterion.

Interested in participating?

Completed

Looking for future studies?

Notify Me

Sponsors and collaborators

Lead sponsor

Istanbul Gelisim University

Other

Registry information

Official study title

Alignment of Large Language Model Responses on Stuttering Definition, Assessment, and Therapy With Clinical Practice Guidelines: A Multi-Model Comparison

Important dates

Study start
2026
Primary completion
2026
Study completion
2026
First posted
Sep 22, 2026
Registry last updated
Sep 22, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.