Study design and setting
This is a multicenter, prospective, longitudinal observational cohort study conducted at tertiary-care obstetric centers. The coordinating center is British Columbia Women's Hospital (Vancouver, Canada), which performs approximately 15 to 20 elective cesarean deliveries under neuraxial anesthesia per week. Additional participating centers are recruited from the academic obstetric anesthesia community through the PIONEER multicenter collaborative and are added as each completes local regulatory approval and site initiation. Every center follows a common protocol with identical intraoperative assessment instruments, identical surgical timepoint definitions, identical eligibility criteria, and an identical longitudinal follow-up schedule, while preserving local institutional standards for anesthetic and surgical care. The study is conducted in compliance with the Declaration of Helsinki and the International Council for Harmonisation Good Clinical Practice guidelines.
The study is strictly observational. The research team does not participate in, direct, or influence any aspect of clinical care. All decisions about anesthesia, surgery, and the management of patient comfort rest entirely with the attending clinical team. The research team documents only what the participant reports and what care the clinical team provides.
Study procedures
Preoperative phase. Eligible patients are identified from the elective surgical schedule and approached on the day of surgery after standard preoperative nursing assessment. Following written informed consent, baseline questionnaires are completed by the participant directly in REDCap on a study device. Demographic and clinical data are collected from the participant and the medical record, including age, height, weight, body mass index, ethnicity, race, primary language, gravidity, gestational age, number of previous cesarean deliveries, surgical indication, comorbidities, psychiatric history and medication history, current medications, and baseline vital signs. Participants receive a structured orientation to the two intraoperative tools (a visual analog scale for sensation intensity and a standardized body diagram for sensation location) and practice using them in the configuration that will be used in the operating room, where the tools are positioned within the participant's field of vision and within reach of the arm bearing the pulse oximeter. Participants are told once, in a standardized manner, that any sensation they report is of interest, that there is no expected answer, and that they may decline to respond at any time.
Intraoperative phase. Anesthetic details are recorded, including neuraxial technique (single-shot spinal or combined spinal-epidural), intrathecal drugs and doses, local anesthetic agent, dose and baricity, intrathecal opioid type and dose, epidural test dose where applicable, time of neuraxial placement, dermatomal level of block to cold, pinprick, and light touch, and motor block (Bromage scale) at the start and end of surgery. For combined spinal-epidural recipients, any intraoperative epidural catheter activation is documented as a distinct event with drug, volume, and timing.
A brief standardized self-assessment is administered at six predefined surgical timepoints: (1) the surgical sharp-stimulus test performed by the surgeon immediately before skin incision, (2) skin incision, (3) uterine incision, (4) end of uterine closure (defined as completion of the final hysterotomy suture, which standardizes this timepoint across centers that do and do not perform uterine exteriorization), (5) end of fascia closure, and (6) skin closure. The timing of each assessment is recorded. At each timepoint the participant is first asked whether they feel any sensation. If they do, they are asked to describe it in their own words (recorded verbatim), to indicate whether they would like the team to do something about it, and, if so, what type of action they are requesting. Participants who report a sensation also rate its intensity on the visual analog scale and mark its location on the body diagram. Each assessment is designed to take one to two minutes.
Fetal delivery is not a standalone assessment timepoint because extraction occurs soon after uterine incision and the associated abdominal pressure precludes meaningful engagement with the assessment tools. Instead, a research team member observes and documents during the delivery interval whether the participant spontaneously reports pain or discomfort and whether the participant exhibits distress (vocalization or facial grimacing). These observer-assessed data are interpreted with appropriate caution as less reliable than self-report. Participants are also told during orientation that they may volunteer any sensation, comment, or concern at any point during surgery and need not wait for a scheduled assessment; any such spontaneous report is documented with timing relative to the nearest surgical timepoint and analyzed separately from the protocol-driven assessments.
Operative times, supplemental intravenous analgesia (by drug type and dose), anxiolytic administration before delivery, uterine exteriorization status (binary), and intraoperative events such as hemodynamic instability, nausea, vomiting, shivering, additional procedures, and estimated blood loss are recorded. At the conclusion of surgery, all participants are asked a single standardized question about whether they would have preferred general anesthesia for any part of the procedure.
Postoperative and follow-up phase. Within 24 to 48 hours after surgery and while still hospitalized, participants complete a further set of questionnaires in REDCap. In-hospital opioid consumption is recorded from end of surgery to discharge and converted to oral morphine milligram equivalents, length of stay is calculated in hours, and any postpartum complications are recorded. Longitudinal follow-up is conducted at 6 weeks, 3 months, and 6 months postpartum using secure individualized REDCap survey links, with automated invitations and telephone reminders from the local site team when questionnaires are not completed within the expected window.
Study instruments
Validated instruments are used throughout. Psychological symptoms are assessed with the PTSD Checklist for DSM-5 (PCL-5), the Edinburgh Postnatal Depression Scale (EPDS), and the Peritraumatic Distress Inventory (PDI). Preoperative anxiety is measured once with the 6-item short form of the State-Trait Anxiety Inventory (STAI-6). Sensation intensity is measured with a visual analog scale, and sensation location is recorded on a standardized body diagram coded to predefined anatomical regions. Established clinical thresholds on the psychological instruments are used to flag participants for safety follow-up.
Sample size
The sample size is determined by powering each participating center independently for the primary hypothesis test, consistent with the site-stratified meta-analytic framework. For each center, a one-sided one-sample test of a proportion was used with a null benchmark of 11.5 percent, an alternative of 17 percent, alpha of 0.05 (one-sided), and 80 percent power, yielding a minimum of 234 analyzable participants per center. To allow for an anticipated 30 percent attrition across six months of follow-up, each center targets enrollment of 335 participants. This per-site target also ensures an expected 27 to 40 events (patient-initiated requests for pharmacological intervention) per center, sufficient to support within-site multivariable adjustment for the three prespecified confounders at an events-per-variable ratio of approximately 6.75 to 10, and to provide meaningful weight in the random-effects pooling. The pooled analytical cohort scales with the number of participating centers: a minimum of two centers (at least 468 analyzable participants) is required for the multicenter analysis, with three or four centers providing progressively greater precision and enabling subgroup analyses that would be underpowered at a single site.
Statistical analysis plan
All quantitative analyses follow a site-stratified meta-analytic framework in which each center is treated as an independent analytical unit and site-specific estimates are pooled using the DerSimonian-Laird random-effects model. This approach was chosen over a single pooled regression model with site as a covariate because it preserves between-site variation as an explicit output, is better behaved with a small number of centers, and supports transparent forest-plot presentation and leave-one-center-out sensitivity analyses. Between-site heterogeneity is quantified with Cochran's Q (with p below 0.10 indicating substantive heterogeneity), the I-squared statistic (25, 50, and 75 percent interpreted as low, moderate, and substantial), and tau-squared.
The primary analysis estimates the proportion of participants who request a pharmacological intervention in response to an intraoperative sensory stimulus at one or more of the six surgical timepoints. Within each center this proportion is estimated with a 95 percent Clopper-Pearson exact confidence interval; site-specific proportions are transformed to the logit scale, pooled by the random-effects model, and back-transformed. The primary hypothesis is tested by rejecting the null (pooled incidence of 11.5 percent or less) at alpha of 0.05 when the lower bound of the one-sided 95 percent confidence interval of the pooled proportion exceeds 0.115. Site-specific one-sided exact binomial tests are reported descriptively to assess consistency of signal direction.
The primary hypothesis test is the only confirmatory analysis. All secondary analyses are exploratory and hypothesis-generating. Descriptive secondary analyses (sensory characterization, concordance between requests and supplemental analgesia, and clinical outcomes) and the qualitative thematic analysis are reported as proportions, medians, or frequencies with confidence intervals, pooled by random-effects meta-analysis where appropriate, without formal hypothesis testing. Inferential secondary analyses (associations between patient-initiated requests and exceeding clinical thresholds on the PCL-5, EPDS, and PDI, and associations between preoperative anxiety and intraoperative sensory outcomes) use within-site multivariable logistic or linear regression, with site-specific adjusted estimates pooled by random-effects meta-analysis and reported as pooled adjusted odds ratios or beta coefficients with 95 percent confidence intervals. Because the secondary analyses are exploratory, no multiplicity adjustment is applied, and results are interpreted in the context of effect size, confidence interval precision, and between-site consistency rather than reliance on p-value thresholds.
Confounder adjustment is restricted to three prespecified primary confounders to respect the events-per-variable constraint at the per-site event count: dermatomal level of sensory block (ordinal), number of previous cesarean deliveries (integer), and preoperative anxiety operationalized as the STAI-6 score (continuous). Documented psychiatric history is retained as a covariate for sensitivity analyses. Two additional prespecified variables (surgical duration and body mass index) are evaluated only in unadjusted within-site analyses and reported as hypothesis-generating. The confounder set was derived from a directed acyclic graph constructed in DAGitty to distinguish confounders from mediators (for example, supplemental analgesia administration) and colliders; the directed acyclic graph will be provided as a supplementary figure.
Longitudinal PCL-5 and EPDS trajectories are modeled within each center using generalized linear mixed-effects models with random intercepts and slopes across the five assessment timepoints; site-specific trajectory parameters are extracted and pooled by random-effects meta-analysis. The two PDI administrations (24 to 48 hours and 6 weeks) are analyzed independently as related but distinct constructs (acute versus recalled peritraumatic distress), with an exploratory within-participant change score. Qualitative responses to the open-ended sensation prompt are analyzed by inductive content analysis. Two independent investigators develop and refine a codebook through iterative open coding; inter-rater reliability is assessed on a random 20 percent subset using Cohen's kappa with a prespecified threshold of 0.70 before full coding proceeds, with a further 10 percent double-coded to confirm sustained agreement and disagreements resolved by a third investigator.
Prespecified sensitivity analyses include participant-level analyses (restriction to single-shot spinal recipients; exclusion of conversions to general anesthesia or major surgical complications; exclusion of participants with elevated baseline PCL-5 or EPDS; restriction of qualifying events to sensations rated at least 30 mm on the visual analog scale; exclusion of combined spinal-epidural participants who received epidural supplementation; and a Hawthorne-effect comparison between early and late enrollment tertiles) and site-level analyses (leave-one-center-out, a coordinating-center-only restriction, and exploratory meta-regression of site-specific incidence on site-level moderators if substantive heterogeneity is detected).
Missing data
The extent and pattern of missing data are assessed at each follow-up timepoint, overall and by center. Baseline characteristics of completers and non-completers are compared within each center and in the pooled cohort to evaluate attrition bias. The primary intraoperative analysis uses complete-case analysis because the primary outcome is captured in real time and missingness is expected to be minimal; per-site missingness rates are reported. Longitudinal models accommodate unbalanced data under a missing-at-random assumption. If missingness at any timepoint exceeds 20 percent at any center, multiple imputation by chained equations is performed at the site level as a sensitivity analysis, with re-estimated parameters pooled by random-effects meta-analysis. Because participants with adverse intraoperative experiences may be more likely to disengage from follow-up (missing not at random), a prespecified tipping-point analysis shifts imputed values toward worse outcomes in increments of 0.5 standard deviation up to 2.0 standard deviations and reports the departure from missing-at-random required to change the study's conclusions, both within each site and after pooling.