Database Credentialed Access
LENS: Lifespan and Sleep-Stage-Resolved Normative EEG Background Slowing - Data and Code
Jin Jing , ChenXi Sun , Wolfgang Ganglberger , Alice Lam , Haoqi Sun , Tianyu Zhang , Daniel Goldenholz , Fábio A. Nascimento , Doyle Yuan , Sándor Beniczky , Jennifer Kim , Aaron F Struck , Sahar F. Zafar , Robert Thomas , Mouhsin Shafi , M. Brandon Westover
Published: July 20, 2026. Version: 1.0.0
When using this resource, please cite:
(show more options)
Jing, J., Sun, C., Ganglberger, W., Lam, A., Sun, H., Zhang, T., Goldenholz, D., Nascimento, F. A., Yuan, D., Beniczky, S., Kim, J., Struck, A. F., Zafar, S. F., Thomas, R., Shafi, M., & Westover, M. B. (2026). LENS: Lifespan and Sleep-Stage-Resolved Normative EEG Background Slowing - Data and Code (version 1.0.0). Brain Data Science Platform. https://doi.org/10.60508/7060-qq30.
Abstract
Objective. Existing references for what constitutes pathological EEG background slowing are based on small samples and primarily awake recordings. To address this gap, we built LENS --- lifespan-and-sleep-stage-resolved EEG growth charts.
Methods. 25,536 clinical EEGs (21,757 patients; infancy to >90 y) were analyzed to extract measurements of spectral power and their ratios, and used to construct sleep-stage × age growth curves (GAMLSS) scored every 15-s segment as a deviation z from its matched normal. LENS was trained on labels extracted from clinical EEG reports to identify and localize pathological focal or generalized slowing (or both), and to provide validated verbal descriptions. LENS was externally validated on two held-out 100-EEG test sets: ON-100 (18 experts) and SAI-100 (14 experts).
Results. EEG growth curves captured development and sleep physiology. Against the panel majority LENS reached AUROC 0.946 (generalized) and 0.921 (focal), with most experts under each curve (78%, 71%), beating a foundation-model on both axes and the strongest published slowing index by 0.10--0.13 AUROC; on the external site it matched experts and beat SCORE-AI, a published reader (focal 0.93). Slowing is the least reliable expert judgement (κ 0.37--0.45); generated reports tracked statements in EEG reports. Clinical readers under-reported pathological slowing in sleep.
Conclusions. One normative field detects slowing at or beyond expert and foundation-model level, yielding validated, automated stage-and-age aware EEG reports.
Significance. LENS is the first lifespan- and sleep-stage-resolved deviation-from-normal instrument for EEG, enabling reproducible automated reporting.
Keywords: EEG; quantitative EEG; slowing; normative modelling; sleep; automated reporting
Background
Slowing of the EEG background is the most common and one of the most clinically consequential abnormalities a neurophysiologist reports. Such slowing takes two forms: focal slowing points to a localized structural or functional lesion; generalized slowing signals diffuse encephalopathy. Yet whether a given rhythm counts as "slow" is relative: the posterior dominant rhythm that is normal at 4 Hz in an infant would be markedly abnormal in an adult, and delta activity that is pathological in the waking adult is physiological in deep sleep. Interpretation therefore depends on age and state, and in practice remains qualitative and expert-dependent, causing inter-reader variability and blocking scalable, reproducible second reads.
The direction of lifespan spectral change is textbook-settled. Low-frequency (delta, theta) power dominates in infancy and declines with maturation; faster rhythms increase; the posterior dominant rhythm accelerates from ~3--4 Hz in infancy to the adult 8--12 Hz alpha by adolescence; aperiodic (1/f) activity flattens with development. In aging the picture is subtler and health-dependent, and much apparent "age-related slowing" reflects comorbidity rather than healthy aging. Modern quantification has taken two forms. First, EEG "brain age" regresses chronological age on resting spectral (and aperiodic) features, formalized as a reusable M/EEG benchmark by Engemann et al. [1]. Second, and conceptually closest to us, normative centile modeling of brain measures: Bethlehem et al. [2] built MRI "brain charts" across the lifespan (n = 101,457) using GAMLSS to yield individual deviation scores, the structural-imaging analog of what we do here functionally with EEG. Crucially, almost all lifespan quantitative EEG (qEEG) reports normal values, not the deviation of an individual clinical EEG from its matched norm, and essentially none is sleep-stage-specific.
Day-to-day clinical EEG reading rests on age norms established decades ago on modest samples. Petersén & Eeg-Olofsson [3] characterized the developing EEG in children aged 1--15 years, qualitatively-to-semiquantitatively and awake-focused. The most direct precedent for our deviation framing is John et al. [4], whose "developmental equations" gave 32 linear age regressions of band power; the companion paper explicitly used deviation from these equations to flag dysfunction, an early "deviation-from-normal" idea, but over a narrow age band, eyes-closed rest only, and a few hundred subjects. John et al. [5], "Neurometrics," and the commercial normative databases it seeded (NeuroGuide/NxLink lineage) generalized age-regressed z-scoring, but these remain wake-resting, age-banded, of modest and somewhat opaque N, and not reproducible.
A separate literature quantifies pathology directly without lifespan normalization. van Putten & Tavy [6] introduced the Brain Symmetry Index (BSI), a 0--1 measure of interhemispheric spectral asymmetry correlating strongly with stroke severity (n = 21); the revised pairwise BSI followed [7]. This is the intellectual ancestor of our homologous-channel asymmetry feature. Slowing ratios such as DAR (delta/alpha ratio) and (delta+theta)/(alpha+beta) track acute-stroke severity (Finnigan & van Putten [8]), and relative delta/theta and alpha-delta ratios discriminate ICU delirium and coma. However, these instruments are each tied to one disease and setting, use fixed thresholds, and ignore age, sex, and sleep stages.
The Temple University Hospital Abnormal corpus (TUAB; López et al. [13]; Obeid & Picone [12]) established binary normal/abnormal classification as a standard benchmark, with ConvNets reaching ~85% (Schirrmeister et al. [14]; Gemein et al. [15]). EEG foundation models (BENDR [17], LaBraM [18], and clinically grounded variants) now dominate representation learning; our foundation-model (which we call "Morgoth") sits in this family and serves here as a reference detector our interpretable model is measured against. Report NLP has extracted findings from free text (Biswal et al. [16]), and recent systems generate narrative from signal, but none ties generated findings to a lifespan-normative deviation model, nor validates stage-specific slowing sentences against the actual report corpus.
No prior work combines (a) lifespan-continuous, [(b) sex-specific,]{.mark} and (c) sleep-stage-specific normative modeling of clinical slowing features with (d) per-recording deviation-from-normal scoring, (e) on >20,000 patients from clinical practice, and (f) closes the loop to clinician-style narrative validated against real reports.
Here we introduce LENS (Lifespan EEG Normative Scoring): lifespan, sleep-stage-resolved EEG growth charts --- the functional-EEG analog of pediatric growth charts --- and the interpretable deviation-from-normal field they yield when a recording is scored against them, which serves both detection and description. Specifically, LENS (1) builds reproducible age × sleep-stage normative growth curves for clinical slowing features across the lifespan, in whole-head, regional, and scalp-topographic form; (2) derives from them a per-segment deviation field (stage- and age-matched z per region × feature), the shared substrate for downstream analysis; (3) detects slowing with an interpretable model based on that deviation field, externally validated on two multi-expert datasets, outperforming the current state-of-the-art foundation-model; (4) generates a structured, validated description (type, laterality, region, anterior--posterior gradient, persistence, sleep stage, electrode) that tracks the report by dose-response contrast; and (5) releases an open Python package with a single-command-reproducible pipeline and a published per-recording label set.
Methods
2.1 Cohort and data
The analysis cohort comprises 25,536 recordings from 21,757 unique patients curated from a single academic health system (MGB sites), spanning infancy to >90 years: 19,617 routine clinical EEGs ("cohort") and 5,919 overnight/long-term studies ("expansion"). Each recording carries report-derived structured finding flags (normal, abnormal, focal slowing, generalized slowing), which are non-exclusive (a report may note more than one). The normal reference is clean-normal: flagged normal in clinical EEG reports (n = 10,189). Pathological slowing groups are evaluated one-vs-clean-normal: focal slowing n = 8,016 and pathologic generalized slowing n = 6,841 (with substantial overlap). Critically, the generalized-slowing flag is split into pathologic vs physiologic slowing, to avoid labeling slowing that occurs in normal drowsy/sleep/hyperventilation as abnormal; only the pathologic set defines the generalized slowing class (physiologic generalized slowing, n = 3,382, is left in the clean-normal reference).
Age and sex are taken from the clinical records; sex is balanced at 49.2% female overall. Age is a critical confound that motivates the entire normative design: abnormal recordings are markedly older than clean-normals (median 53.9 y [IQR 24.3--69.7] vs 36.8 y [18.6--59.2]), so any unadjusted comparison would conflate slowing with age. Full cohort characteristics --- age bands, sex, recording length, usable segments, stage composition, and the abnormal-detail strata (focal side, generalized topography, band) --- are given in Table 1 (scripts/table1_sap.py, SAP §10). Segment-level feature tables are keyed on the recording; patient is the clustering unit for all confidence intervals (patient-clustered bootstrap), and report-derived quantities are computed only on the clean_pair set (§2.6). This work was conducted under IRB protocol number 2022P000417, with the BIDMC IRB granting a waiver of consent.
2.2 Reproducible feature extraction
To make the pipeline reproducible and extensible to new recordings, extraction is implemented in Python. The pipeline re-montaged the referential 10-20 EEG to 18 bipolar (double-banana) channels, applied a 0.5 Hz high-pass and 50/60 Hz notch, segmented into 15-second windows (3000 samples at 200 Hz, step 2800), and computes a multitaper power spectral density (time-bandwidth NW = 4, 7 tapers). Per segment we derive band powers (delta/theta/alpha/beta/gamma/total), relative powers, and inter-band ratios (DAR = delta/alpha, TAR = theta/alpha), for each of the 18 bipolar channels, plus 8 homologous left--right pairs for focal localization. Channels are aggregated into regions (whole-head, left/right temporal, left/right parasagittal, anterior, posterior) via config/channels_regions.yaml. Linear powers are log-transformed before z-scoring.
Calibration. The delta band is defined as 1--4 Hz, bringing normal whole-head relative delta to ~0.34 (physiological). Because all deviation scoring is z relative to each feature\'s own age/stage normal curve, the scoring, discrimination, and generated descriptions are scale-invariant.
Artifact rejection. A per-segment filter marks a 15-second bipolar segment unusable if it is flat/disconnected, high-amplitude (electrode pop/movement), or EMG-dominated; the usable-segment fraction is carried into scoring so low-yield recordings are flagged.
2.3 Sleep staging
Routine EEG feature sets are unstaged. We staged each recording from the raw EEG with the morgoth2 deep-learning automated sleep stager (5-class window-level model), assigning each 15-second feature segment its majority stage (W/N1/N2/N3/REM) [20]. The overnight recordings fill the deep-sleep coverage routine clips cannot. Staging runs in a dedicated virtual environment (the stager\'s pyhealth dependency pins pandas <2, incompatible with the analysis stack).
2.4 Normative growth curves (GAMLSS)
For each feature × region × sleep stage we estimate normal-population percentile "growth curves" as a continuous function of age using GAMLSS, the method behind clinical growth charts (Cole & Green [11]; Rigby & Stasinopoulos [10]). Age is entered as log10(age + 1/12), expanding infancy where maturation is fastest. Two families are used according to the feature\'s support: positive features (relative powers, ratios) are fit with a Box--Cox-t (BCT) distribution with penalized-spline median, dispersion, and age-varying skewness (mu, sigma, nu each smooth in log-age), which removes an infant-age median bias a constant-skewness fit produces; real-line log features (log_delta, log_theta, log_TAR, which take negative values in 15--37% of segments) are fit with a support-aware robust median-in-log-age model on the real line. The Python BCT z-scores are validated exact against R gamlss centiles.pred. Each curve is checked against a model-free rolling median in a sliding, age-widening window, which the fitted median tracks closely across the lifespan.
Sex is pooled. Conditioning norms on sex changes abnormality discrimination by ΔAUROC ≤ 0.002 in every feature × region × contrast we tested; we therefore pool sexes, doubling the effective sample per age.
Support-aware refit. An earlier fit exposed a stiff μ-spline; the support-aware fix (BCT for positive features, robust real-line for log features) lowered report/deviation discordance from 41.1% to 38.4% and is the fit used throughout.
2.5 The per-segment deviation field
Every 15-second segment is scored as a deviation z for each feature × region against its own (sleep-stage, age)-matched normal curve (data/derived/segment_deviation/, joinable 1:1 to the segment tables), using a precomputed normal grid (grid_norm.json regional, grid_anorm.json asymmetry). Because each segment is compared to its own stage\'s normal, delta that is abnormal in wake but normal in N2/N3 yields near-zero deviation in normals. This per-segment, per-region field is the single measurement layer both the detector and the description run on; it does not depend on any label (it is fit to the normal population only).
2.6 Report--recording pairing and label provenance
We joined reports to EEGs at the patient level upstream: a single report is stamped onto a mean of ~2.9 EEGs of that patient (maximum 170). We assign each report to the EEG nearest it in time, and define a recording as cleanly paired (clean_pair) when its report is claimed by no other EEG, or when it is that report\'s nearest-in-time owner; all report-text-derived quantities are computed on clean_pair recordings only (1,664 EEGs dropped by the filter). Corrected label rules (scripts/label_rederive_sap.py) make focal slowing always pathologic, generalized slowing pathologic only if the report names it among the abnormalities, and abnormal-without-slowing its own stratum.
2.7 Detection
Detection and description share the deviation field. We report three detectors.
2.7a LENS --- the deviation detector (primary). A single segment-level model with two independent heads, focal and generalized, trained on the single-scored clinical data, with a patient-stratified train/test split balanced over lifespan × {control, focal-only, gen-only, both, other} (scripts/53). Each 15-second segment yields stage-matched deviation features: a whole-head amount z for generalized, and localization features (peak-region z, focality = peak − median region, asymmetry z, spatial stability) for focal (scripts/54, 55). The model works on a lone 15-second clip (segment output) and aggregates for full recordings: generalized pools segment amount scores (top-5), while focal aggregates the localization features at the recording level, a split forced by the mechanics (generalized is diffuse; focal is spatial + intermittent). It is trained on clinical report labels and tested (external validation) unchanged on two independent, multi-site evaluation sets --- each 100 EEGs annotated by many experts and drawn entirely from hospitals outside the training health system: ON-100 (100 recordings from five US centers --- Barnes-Jewish Hospital, Washington University in St. Louis, the Dallas VA Medical Center, UT Southwestern, and Louisiana State University --- each read by 18 experts) and SAI-100 (the 100-recording holdout set of the SCORE-AI validation study [9], routine EEGs from Haukeland University Hospital (Norway), the Danish Epilepsy Centre (Denmark), and Mayo Clinic (USA), each read by 14 experts, and with SCORE-AI's own automated predictions available). Both benchmarks are held out entirely from fitting (no patient overlap with the report-trained cohort or with each other). ON-100 and SAI-100 are therefore both external validations, at institutions distinct from the training data and from one another.
2.7b The Morgoth gate (reference). A three-tier hierarchical foundation-model gate (abnormal → slowing → focal/generalized EEG-level heads), calibrated against expert reports, serves as the learned-representation reference detector. It is not part of the interpretable pipeline; it is the bar LENS is measured against.
2.7c The van Putten benchmark. We recomputed the van Putten family of indices: the Brain Symmetry Index and its revised pairwise form (r-sBSI), the diffuse-slowing index Q_SLOWING, the anterior--posterior gradient Q_APG, homologous-pair asymmetry Q_ASYM, and the slowing ratios DAR (delta/alpha) and DTABR = (δ+θ)/(α+β), plus the 95% spectral-edge frequency SEF95 (scripts/recompute_vanputten_fullcov.py). Each is evaluated as published and age-conditioned against our normative curves (scripts/recompute_vanputten_fullcov.py); the clean-panel head-to-head (Figure S7, scripts/vanputten_panel_s7.py) then pits the best index per axis against LENS and the Morgoth gate.
2.8 Description: reading the deviation field into words
Once slowing has been detected in a recording, LENS generates a structured verbal description of what kind of slowing it is, where it is, and how persistent it is (scripts/56 extracts descriptors; 57 renders the panels; 58 generates sentences). Per recording, from the deviation field we read off: type/amount (whole-head delta-excess and theta-excess z: p90, mean, prevalence), laterality (signed left-minus-right region z, + = left), region (per-lobe magnitude and relative prominence/focality), anterior--posterior gradient (anterior − posterior z), persistence (prevalence = fraction of abnormal segments, longest continuous run, number of episodes), and per-sleep-stage versions of each. Descriptors are assembled into a compact finding line and a full report-style paragraph, governed clause-by-clause by docs/claims_table.md: magnitude as SD and centile (never a severity adjective), prevalence as a percentage following ACNS [19] terminology (occasional/frequent/abundant/continuous) as an internal gloss only, band as a low-confidence δ/θ/mixed call read from absolute delta/theta power dominance (whole-head mean log(δ/θ), marginal-matched to the report distribution), side asserted with the maximum-deviation lobe flagged provisional, anterior--posterior predominance asserted only when it clears the normal centile, stage accentuation and "present only during sleep", and a required abstain path ("no lateralizing or regional spectral excess above the normal centile") so the system never invents a lobe. Validation is by contrast (dose-response), rather than binary classification: each continuous descriptor is compared between recordings whose report does and does not name the corresponding finding, and must be higher where the report names it.
2.9 Expert panels and the human ceiling
Agreement with a single clinical report is bounded by report reliability. We therefore evaluate against two independent, multiply-read external datasets --- ON-100 (100 EEGs from five US centers, each read by 18 electroencephalographers, a subset re-read) and SAI-100 (100 EEGs from the SCORE-AI validation study, each read by 14 experts) --- each recording judged for focal and generalized epileptiform and non-epileptiform abnormality. Ground truth is the panel majority. Our primary metric is the percentage of experts under the curve: we place each expert on the model\'s ROC curve (and its precision--recall, PR, curve) as one sensitivity--specificity point --- graded against the leave-one-out consensus of the other readers --- and count how many fall below the curve, i.e. how many the model matches or beats at that expert\'s own operating point. A model performing at the panel\'s own level would put roughly half of the experts under its curve. Inter-rater reliability (Fleiss κ, chance-corrected agreement) and each reader\'s self-consistency quantify the human ceiling. Band agreement is evaluated against experts\' per-band calls (no text extractor in the loop).
2.10 Statistics
All confidence intervals are patient-clustered bootstraps. Detection is reported as the area under the ROC curve (AUROC) vs the clean-normal (report) or panel-majority (panel) reference. Description contrasts report group medians with Mann--Whitney tests and Cohen\'s d. LENS is evaluated leave-one-out on the panel with no refitting or threshold tuning. The full pipeline is reproducible from the derived tables via a single ordered runner (§ Data and code availability).
Data Description
De-identified derived data reproducing every figure, table, and number in the paper. Source EEGs are referenced from the published BDSP EEG dataset (s3://bdsp-opendata-repository/EEG/bids/), not re-hosted; data/raw/manifest.csv in the code repo lists all 25,008 source EDFs.
Under s3://bdsp-opendata-credentialed/morgoth-slowing/:
- derived/ (66.7 GB):
segment_master/(per-segment x per-channel band powers, hive-partitioned byeeg_id),segment_deviation/(age/stage-matched deviation z),description_*,single_model_segfeats,occasion_*,grid_norm.json, gate tables. - panels/ (2.4 GB): ON-100 expert-panel inputs and per-reader votes.
- manifest_build/ (146 MB): manifest lineage + raw report-findings CSVs (regenerable; provenance).
De-identified (BDSP surrogate IDs, shifted dates). A directory-level bucket_manifest.csv is at the prefix root.
Usage Notes
Code on GitHub: https://github.com/bdsp-core/morgoth-slowing-growth-curves
Code: https://github.com/bdsp-core/morgoth-slowing-growth-curves
export AWS_PROFILE=opendata aws s3 sync s3://bdsp-opendata-credentialed/morgoth-slowing/derived/ data/derived/ pip install -e . bash scripts/reproduce_story.sh results
Load tables with pandas.read_parquet; the per-segment deviation field is the shared substrate for detection and description. See REPRODUCE.md for the figure/table -> script -> input map.
Release Notes
Version 1.0.0 - initial release accompanying the LENS paper (Jing, Sun, ... Westover).
Ethics
This work was conducted under IRB protocol number 2022P000417; the BIDMC IRB granted a waiver of consent. Data are de-identified.
Acknowledgements
Jin Jing and Chenxi Sun contributed equally to this work (co-first authors).
Dr. Westover\'s laboratory is supported by grants from the NIH (R01AG073410, R01HL161253, R01NS126282, R01AG073598, R01NS131347, R01NS130119) and by AWS.
Conflicts of Interest
Dr. Westover is a co-founder of, serves as a scientific advisor and consultant to, and has a personal equity interest in Beacon Biosignals. The remaining authors declare no competing interests.
Access
Access Policy:
Only credentialed users who sign the DUA can access the files.
License (for files):
BDSP Credentialed Health Data License 1.5.0
Data Use Agreement:
BDSP Credentialed Health Data Use Agreement
Required training:
Discovery
DOI:
https://doi.org/10.60508/7060-qq30
Project Website:
https://github.com/bdsp-core/morgoth-slowing-growth-curves
Corresponding Author
Files
- be a credentialed user
- sign the data use agreement for the project