Background and Rationale
Systematic classification of intraoperative adverse events (iAEs) is a prerequisite for measuring surgical safety. In adults, ClassIntra (Dell-Kuster et al., 2020) grades intraoperative events on five severity levels (Grade I-V) and has been internationally validated. In children, no population-specific instrument exists and pediatric experience with ClassIntra is limited to a small number of reports: a neurosurgical pilot congress abstract (Drexler et al., 2022; n=21), a full-text prospective neurosurgical cohort from the same group (Middelkamp, Drexler et al., World Neurosurgery 2025; n=47), and a prospective multidisciplinary pediatric robotic surgery programme in which ClassIntra was recorded alongside Clavien-Dindo grading (Vinit et al., Annals of Surgery 2023; n=300, of which 83 digestive and 105 urological or gynaecological procedures). In none of these was ClassIntra applied to video recordings, in none was its inter-rater reliability assessed, and none focused on appendectomy. The iAE profile of pediatric laparoscopic appendectomy - the most common pediatric surgical procedure - has not been systematically documented.
For postoperative complications, a pediatric-validated instrument does exist: the Clavien-Madadi classification (Madadi-Sanjani et al., 2023; ERNICA validation 2024). The relationship between intraoperative event severity and postoperative complication severity has not been studied in children.
Design
Prospective, single-center, single-surgeon, video-based pilot study, positioned at Stage 2a (Development) of the IDEAL Framework for surgical innovation. Single-arm observational cohort; no control group and no comparison arm. The principal investigator performs all operations and does not serve as a rater; the investigator acts only as arbiter in case of unresolved disagreement between raters.
Video Assessment Protocol
Operations are recorded in 1920x1080 H.264 MP4 format. Recordings are de-identified, assigned a case code (PA-XXX), and distributed to raters in randomized order. Two independent pediatric surgery specialists from outside the operating institution serve as raters; a third specialist is pre-designated as reserve. Raters have no access to postoperative clinical data, which prevents halo effects. Raters undergo a three-stage training program before assessment begins: theoretical training (2 weeks), calibration sessions (2 weeks), and certification; an inter-rater reliability threshold of 0.60 is targeted at first calibration, with an additional calibration round if not reached.
The assessment window is deliberately restricted to the laparoscopic phase, from insertion of the laparoscope into the abdomen until its withdrawal. Skin incision, port placement, fascial closure and anesthesia-related events are outside the assessment window. This partial application of ClassIntra is a pre-specified scope decision made for reasons of pilot scope management and rater workload, and is reported as a limitation. The study is not a retrospective analysis of an existing video archive: all videos are obtained prospectively after ethics approval, in accordance with the study protocol.
Three-Level Event Characterization
Each identified intraoperative event is characterized on three levels:
- Level 1, event type: E1 bleeding, E2 thermal injury, E3 avulsion, E4 serosal tear, E5 injury to adjacent organ, E6 detached clip, E7 spillage (pediatric adaptation after Sanmoto 2025).
- Level 2, mechanism: GERT categories M1 excessive force/distance, M2 insufficient force/distance, M3 wrong orientation, M4 insufficient visualization (Bonrath 2013), a procedure-independent internationally validated framework.
- Level 3, severity: ClassIntra Grade I-V (Dell-Kuster 2020), applied to the laparoscopic phase.
In parallel and independently of the video assessment, postoperative complications occurring within 30 days are graded by the principal investigator using the Clavien-Madadi classification from routine clinical follow-up data. No automatic conversion between complication classification systems is performed.
Relationship to Institutional ERAS Care
The department applies a 20-item Enhanced Recovery After Surgery (ERAS) protocol as routine care for all children undergoing appendectomy, independently of this study. Protocol adherence and recovery data (including time to medical readiness for discharge) are therefore documented for all participants as part of routine care, and are used in this study only as exploratory secondary variables. Enrollment in any other study is neither an inclusion criterion nor a requirement for participation in this study.
Statistical Approach
Inter-rater reliability is assessed with Gwet's AC1 as the primary measure, with weighted Kappa and the intraclass correlation coefficient reported as supporting measures. As a pre-specified sensitivity analysis, AC1 is additionally computed in the subset of videos with at least one identified intraoperative adverse event; if this subset comprises fewer than 15 videos, the subset analysis is reported descriptively only. Event incidence is reported with Wilson score confidence intervals. Exploratory hypotheses are analyzed with logistic regression (ERAS protocol adherence) and Spearman rank correlation (time to medical readiness for discharge; Clavien-Madadi severity). No adjustment for multiple testing is applied, as these analyses are explicitly exploratory; effect sizes with 95% confidence intervals are emphasized over p-values. Missing data are managed with a pre-specified, proportion-dependent three-scenario strategy. Statistical analysis is conducted under the responsibility of the principal investigator.
Sample Size
A target of 80 patients is based on published precedent for reliability pilot studies rather than on a confirmatory power calculation. This sample supports estimation of AC1 with acceptable precision (approximately +/-0.10 around an expected AC1 of 0.65) but does not provide statistical power for confirmatory hypothesis testing; secondary results are therefore reported as effect-size estimates.
Known Limitations
The single-surgeon, single-center design limits generalizability and is a deliberate pilot choice. Findings are restricted to non-complicated appendicitis and cannot be extrapolated to other pediatric procedures. Video-based assessment may under-detect anesthesia-related and organizational events. Post-hoc exclusion of conversion cases lowers the estimated incidence of high-grade events. The hybrid ClassIntra/Clavien-Madadi structure has not previously been tested in a pediatric population, so no accuracy benchmark exists. Accordingly, this study provides feasibility evidence, not validity evidence; multicenter studies are required for the latter.