BACKGROUND
The GMFCS-E&R separates its five levels mainly by two things: what a child cannot do without help from another person, and what a child cannot do without a hand-held mobility device. The wording is explicit in the age bands relevant here. At Level I a child moves in and out of floor sitting and standing "without adult assistance." At Level III a child "may require adult assistance to assume sitting" and needs "adult assistance for steering and turning" when walking with a walker. Adult contact and external support are therefore not proxies chosen for convenience. They are part of the definition.
Scoring them is another matter. It takes an experienced clinician, it is slow, and between 2 and 6 years of age the assignment is known to be difficult.
STUDY DESIGN
This is a two-center feasibility study of a determination made from video. The index test is a binary determination produced by an open-weight vision-language model. The reference standard is a human annotator's record of the same clips. No clinical grade is requested from the model at any point.
Four elements of the design are fixed in advance.
First, analyses are stratified by participant and by movement class. Summary statistics are computed within a stratum and pooled afterward, so movement class cannot vary with the outcome.
Second, each movement attempt is reduced to exactly 16 frames, 8 from each of two camera views. The count is fixed and is never scaled to clip duration.
Third, the model receives the frames with no identifier, no movement name and no GMFCS level, and answers two questions: was the child touched by an adult, and did the child bear load through an external object.
Fourth, a deterministic rule combines the two answers. An attempt counts as performed unaided when both answers are negative.
BLINDING
No masking of intervention assignment applies, because the study has a single group and no assignment step. What is masked is the reading of the data, and it is masked on both sides.
The rater model sees only the extracted image frames. It is given no participant identifier, no movement name, no GMFCS level and no other clinical information, and it is never shown what the reference standard recorded for that attempt. Each request carries one attempt and nothing else, so nothing learned from one attempt can be carried into the next.
The human annotator who produced the reference standard worked from the clips alone, blind to clinical information, and recorded the two features before any model output existed. The stimulus set was fingerprinted and sealed before rating began, and the rating files are recorded as absent in that seal, which is positive evidence that no model output existed at the time the stimuli were fixed.
The index test and the reference standard were therefore read independently of each other. Neither reader saw the other's output.
PARTICIPANTS AND RECORDINGS
Twenty-six children with cerebral palsy were enrolled, 15 at Samsung Medical Center and 11 at Asan Medical Center. Ages ranged from 25 to 72 months, median 48.5 months. Fifteen were male and 11 female. All five GMFCS levels are represented: 6 children at Level I, 5 at Level II, 4 at Level III, 3 at Level IV and 8 at Level V.
Five movement classes were filmed: walking, crawling, side-rolling, floor sit-to-stand and stand-to-floor-sit. The unit of annotation is the movement attempt, filmed by up to three cameras at once. The dataset holds 536 annotated attempts across 1,554 clip files, recorded at 1920 x 1080 and 30 frames per second.
REFERENCE STANDARD
One annotator, blind to clinical information, recorded two features for every attempt: whether a caregiver physically assisted the child, and whether the child used an acrylic support stand or a walker.
ANALYSIS AND IMAGE-INDEPENDENT CONTROLS
A clinical video dataset carries the answer in places that have nothing to do with the images, so every reported figure is set against baselines that read no image content at all: majority class, movement repertoire alone, a clip-duration threshold and scene composition alone.
Two of these baselines are strong in this dataset, and they are the reason the design takes the form it does. The set of movement classes a clinician chose to film for a child recovers the dichotomized GMFCS group in 25 of 26 children, which is why analyses are stratified by movement class. Mean clip duration alone recovers it in 24 of 26, which is why the frame count per attempt is fixed. For the same reason the primary reporting level is the two video-derived determinations rather than a participant-level grade. Participant-level aggregates are reported as descriptive only and always beside their image-independent control.
RATER MODELS
Rating uses open-weight vision-language models run on hardware controlled by the investigators. Each model is identified by name and version, and no accuracy figure is carried beyond the version that produced it. Whether one rater model can be substituted for another is treated as something to measure rather than assume, and agreement between raters on identical image inputs is reported.
SCOPE
The study is designed to show whether adult contact and external support can be read from video by an open-weight vision-language model closely enough to a human annotator to be useful. It does not assign a GMFCS level.
Those two features are part of how the GMFCS defines its levels, so reading them reliably from video contributes to GMFCS assessment directly rather than alongside it. If the feasibility holds, the same two determinations can be put in front of a clinician as video-derived evidence for criteria that are at present judged by eye, and can serve as the input layer for later work on supporting GMFCS assessment between 2 and 6 years of age, the range in which that assessment is most difficult.