Visual Attention Lab / Brigham and Women's Hospital
Boston, Massachusetts, 02215, United States
NCT Number: NCT05272189
The study is one part of a "bundle" of experiments that constitute Project Three of a National Eye Institute grant. Project Three includes a series of experiments that investigate how changing the input from a simulated AI can affect the decisions made by human observers in a two-alternative forced choice task (like the decision to recall a woman for further examination in mammography). HAICT 7, the experiment described here, investigates how changing prevalence affects human performance when AI is used as a Second Reader.
Looking for future studies?
Notify Me18 year and older
All sexes
Interventional
Not applicable
Boston, Massachusetts, 02215, United States
This text is the text of the pre-registration for the HAICT 7 experiment as described on the Open Science Framework. https://osf.io/hngu4/
NOTE: This study is representative of studies conducted in Project 3 of this grant. There are multiple experiments in the bundle of experiments represented by Project 3 but it is not possible to register a bundle of studies on CT.gov.
NOTE: Since the pronoun comment is advisory, we will leave it for now.
Human-AI Collaboration Tester (HAICT) Exp. 7 (lightly edited from OSF)
yes
Background: In a variety of search experiments, both basic and clinical, the data have been consistent with a situation where the variability of the signal (or target) is greater than the variability of the noise (distractors). The classic sign of this is a zROC function with a slope < 1 - typically around 0.6. A slope of 1.0 is indicative of an equal variance 2AFC task. For the HAICT task that we have been testing, we would expect equal variance, but we think it would be worth checking so we will systematically vary prevalence which will shift criterion. That will sweep out an ROC curve that we can examine.
We will also test the Second Reader faux-AI in order to determine if low prevalence makes Second Reader worse.
The main dependent variables of interest are accuracy (and the signal detection derivatives of accuracy, d' and c), reaction time, and subjective ratings on the survey following each block.
This series of experiments investigates how changing the input from a simulated AI can affect the decisions made by human observers in a two-alternative forced choice task (like the decision to recall a woman for further examination in mammography). We have developed a paradigm called the Human-AI Collaboration Tester (HAICT) that allows for efficient testing of interactions between a human and a simulated AI.
The observers' task in all conditions is to give a 2AFC decision about whether a stimulus is "bad" or "not bad." To use language roughly mimicking a medical diagnosis, each stimulus is referred to as a "case." Observers are asked to make a 2AFC decision about arrays of colored shapes. The decision is made based on the predominant color of the case. The number of elements of each color are drawn from one of two normal distributions, one for positive (bad) stimuli and the other for negative (not bad) stimuli.
The results from previous HAICT experiments (3 and 4) showed that human performance in the Second Reader condition drops off significantly at low prevalence. Performance in the Second Reader condition was better than Baseline when the prevalence of bad cases was 50% but was significantly worse than Baseline when prevalence was only 10%. In this experiment, we manipulate the prevalence of "bad" cases in the Second Reader and Baseline conditions. Four different prevalence rates will be tested - 10%, 33%, 67%, and 90%. Observers will complete 8 blocks (2 AI rules x 4 prevalence rates), and block order is random.
AI rules to be tested:
As in Experiments 1-5, the AI d-prime is fixed at 2.2. Feedback is known to increase the prevalence effect, so feedback will be given in both the practice and the test trials. Observers will complete 20 practice trials and 200 test trials in each block. Immediately after each block is completed, observers will be shown a summary of their performance. After the Second Reader blocks, they will also be asked to answer three subjective questions about the usefulness of the AI (see "Files" for more details).
First, we summarize the number of hits, true negatives, misses, and false alarms in each block. From this, we can calculate the accuracy, the positive predictive value, sensitivity (d-prime), and the criterion for each observer under each of the different conditions. Given measures of performance at 4 levels of prevalence, we can estimate the ROC curve (pHit x pFA) and the zROC function (zHit x zFA). We will test the hypothesis that the slope of the zROC is equal to 1 (the consequence of an equal variance 2AFC task).
We will look to see if the observers' subjective opinions about the AI are correlated with variables such as the empirical d-prime, or the positive predictive value.
We will test 12 observers. This is consistent with the sample sizes of previous experiments.
N/A
Healthy volunteers accepted: Yes
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
In this experiment, in some conditions, the participant makes their decision in the presence of information about a simulated artificial intelligence decision.
The frequency with which targets are presented varies from 10% to 90%
Other names: Base Rate
Time frame: Data are collected within a session of about an hour.
D' (d-prime) is the signal detection theory measure of the level of performance on a task. It is computed by calculating the proportion of true positive responses =(true positive trials)/(true positive + false negative trials) = p(TP) and by calculating the proportion of false positive responses =(false positive trials)/(false positive + true negative trials) = p(FP). These values are transformed into 'z-scores' (for example, using NORMSINV in Excel to calculate the inverse of the standard normal distribution). D' is defined as Z(TP)-Z(FP). Its range is from 0 for cases where no signal can be discriminated from the noise, to ~4.0. The upper limit is not defined, but 4 would mean that and observer is essentially perfect at discriminating signal from noise.
Time frame: Data are collected within a session of about an hour.
Criterion, like D' (see above) is calculated from z(TP) and z(FP). Criterion ( c ) = (z(TP)+z(FP))/-2. A value of zero means that the observer is equally likely to make a positive (e.g. 'target present') response as a negative (absent) response. Positive values mean that the observer is more likely to say "absent" (a "conservative" criterion). Negative values mean the observer is more likely to say "present" (a "liberal" criterion). Liberal and conservative have no political connotations in this case. Criterion values almost always fall between -2 and 2.
Time frame: Data are collected within a session of about an hour.
This is the measure of how long it takes to make a response.
Brigham and Women's Hospital
Other
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT07692698
Decision Making, Executive Functions (EF)
Elâzığ, Merkez, Turkey (Türkiye)
View Trial DetailsNCT07479368
Decision Making, Football Players
Medellín, Colombia
View Trial DetailsNCT05173922
Alzheimer Disease, Behavior
Aurora, Colorado, United States
View Trial DetailsNCT04466865
Behavior, Chronic Disease
Denver, Colorado, United States
View Trial Details