
43 clinicians, 107 patients, 204 recordings, and the trouble with one night of sleep
This week's papers attach an error bar to the nightly sleep number, from single-night apnea thresholds to four scoring algorithms, and close with a seven-night test of your own measurement.
A single night of sleep is one sample, and this week's research measures how much a sample can settle. More than 80% of the clinicians surveyed across 28 sleep centres in 15 countries called the single-night apnea-hypopnea index thresholds their specialty diagnoses by insufficiently robust, and 83% agreed that measuring several nights would reduce how often patients are misclassified. 1 Seven of those 28 centres run multi-night protocols. In a separate study published the same week, wrist actigraphy and laboratory polysomnography, recorded on the same eight nights in 107 patients, agreed on group averages and diverged by up to several hours for a single person, with the 95% limits of agreement for one night's total sleep time spanning roughly four hours. 2 Three more papers this week take up the analysis side of the same question: which algorithm produced the number. Read together, they give the nightly score its proper status. It is a reading with a known error bar, and the width of that bar decides what the reading can support.
Single-night diagnosis is already conceded to be insufficient
Karkala and colleagues surveyed clinicians inside the European Sleep Apnea Database network about multi-night testing for obstructive sleep apnea: how much they believe in it, and whether their centres can deliver it. 1
The belief is not in question. Of 43 respondents drawn from 28 centres in 15 countries, more than 80% considered single-night AHI thresholds insufficiently robust, and 83% agreed that multi-night testing reduces the risk of misclassification. Practice is another matter. Only 7 of the 28 centres currently run multi-night protocols, and their diagnostic pathways differ from one another. Nearly half of the centres reported they would not be ready to adopt multi-night testing even with guaranteed reimbursement, citing inadequate staffing and the absence of standardized European or national guidelines ahead of cost. Clinicians supported the selective use of multi-night testing for diagnostic uncertainty or suspected night-to-night variability rather than its universal adoption, and the regional split was sharp: Western European centres were furthest along, while Southern, Central and Eastern European centres reported the tightest organizational and financial constraints. 1
The quality signal here is the design. This is an online survey of clinicians inside one research network, so it records opinion and readiness. What the people who diagnose apnea think of a one-night threshold is one question; whether multi-night testing improves anyone's outcome is a separate one the survey does not reach.
The cheap route to more nights carries an error that changes with the patient
If a lab night is expensive, a wrist device that records eight of them is the obvious substitute. Drakatos and colleagues tested how far that substitution holds, recording MotionWatch 8 actigraphy alongside inpatient polysomnography for eight nights in 107 patients at a tertiary sleep centre (mean age 37.6 years, SD 14.9; 56% female). 2
The device systematically flattered sleep:
- Total sleep time was overestimated by 27.31 minutes on average, sleep efficiency by 6.09 percentage points, and wake after sleep onset was underestimated by 14.86 minutes.
- The 95% limits of agreement for total sleep time ran from −96.78 to +151.40 minutes: on one person, the two methods can sit roughly four hours apart.
- Agreement was strong for total sleep time (r = 0.728), moderate for sleep efficiency (0.540) and wake after sleep onset (0.581), and weak for a movement-based fragmentation index (0.328).
- The error was not random across patients. Older age widened the overestimation of total sleep time and sleep efficiency; the 17% of patients with periodic limb movements showed a different pattern again, with less sleep overestimation but exaggerated movement-derived fragmentation.
The authors read this as clinically meaningful at the level of a group and as unreliable at the level of one person, and they note the setting matters: a single inpatient comparison should not be carried over to home recording. 2
Scoring the night is the second place error enters. Crosbie and colleagues compared four automated sleep-staging algorithms against manual scoring on 204 polysomnograms from 75 athletes, gathered retrospectively in a single laboratory. Macro F1 scores landed at 0.76 for GSSC, 0.71 for U-Sleep, 0.70 for YASA and 0.69 for Luna, which is close to the agreement two trained human scorers typically reach with each other. Robustness separated the field more than accuracy did: signal quality accounted for far less of the error in GSSC and Luna (R² = 0.08) than in YASA (0.16) or U-Sleep (0.29), which matters in the unquiet conditions of a real bedroom rather than a sleep laboratory. 3
Adding sensors is the third place, and it does not work the way device marketing suggests. Tran and colleagues trained random-forest classifiers on EEG, ECG, blood-oxygen saturation and abdominal effort from 148 patients, fused them at the probability level, and tested on 37 held-out patients, deliberately withholding oxygen data in the quality-rejected epochs that real recordings contain. Abdominal effort alone reached a macro F1 of 0.996. Adding oxygen saturation to abdominal effort moved that to 0.997, while adding oxygen saturation to the ECG-plus-EEG pair made performance worse (0.872, down from 0.899). Performance held steady with 30.7% of the oxygen epochs dropped. What carried the result was how much the included signals told the model about each other, and the count of sensors carried much less. 4 The near-perfect figure for abdominal effort is a single held-out set of 37 patients in one centre, and it is the kind of number that needs replication elsewhere before it means anything general.
A commentary in Sleep makes the same point about the most ambitious version of this work. Helmy and colleagues reviewed published sleep foundation models and found their training cohorts consistently skewed toward older, predominantly mono-ethnic populations with established comorbidities, with inconsistent evaluation frameworks and few supervised comparisons. Applying one such model without fine-tuning to a separate cohort of 51 narcolepsy type 1 patients and 28 healthy controls, they found zero-shot sleep staging modest and below supervised methods on the same cohort, and PSG-derived embeddings added little to disorder classification beyond demographics. Their conclusion is that sleep foundation models are a promising research direction and are not yet suitable for clinical deployment. 5
These are not only technical problems. In a Nature interview, the digital-health researcher Wuyou Sui argues that when researchers build studies on consumer wearables, the decisions that carry the most risk are who the data reaches and what the vendor's algorithms do with it, since the researcher shares the dataset with a company whose interests are outside the study. 6 A measurement instrument that comes with a data broker attached is a different proposition from a sensor, and the field is only beginning to price that in.
Which parameter you compute from a week of data decides what shows up
Collecting many nights does not finish the job, because the quantity you extract from them decides the answer. Chang and colleagues pooled seven cohort and cross-sectional studies covering 109,365 participants and asked which rest-activity rhythm parameters travel with cardiovascular disease and death. Rest-activity rhythm here is the 24-hour pattern of movement and rest, and the analysis used three standard nonparametric measures: interdaily stability, which describes how consistent the day-to-day pattern is; intradaily variability, which measures how fragmented it is; and relative amplitude, which measures how robust the day-to-night contrast is. 7
- Higher intradaily variability went with higher all-cause mortality (hazard ratio 1.05, 95% CI 1.02–1.09) and with cardiovascular disease prevalence (odds ratio 1.73, 95% CI 1.21–2.47).
- Lower relative amplitude went with higher cardiovascular mortality (hazard ratio 1.50, 95% CI 1.13–1.98) and higher all-cause mortality (hazard ratio 1.46, 95% CI 1.27–1.68).
- Interdaily stability showed no significant pooled association with any outcome. 7
The pattern is worth reading closely. The two parameters that carried associations describe the shape and the fragmentation of a person's rhythm over a week; the one that came back null describes how steady it is from day to day, which is the property a behavioral regularity programme can most directly reach. The authors present these measures as early, non-invasive risk indicators and call for work that tests whether interventions aimed at the rhythm reduce risk. 7 That test has not been run. And seven observational studies can produce a pooled estimate that misses a real association, so the null for interdaily stability is one estimate from a thin evidence base rather than a demonstration that day-to-day steadiness does nothing.
Behavioral programs move timing faster than they move biology
Two intervention studies this week give a partial answer, and both separate what changed in behavior from what changed in the body.
Duraccio and colleagues ran a six-week transdiagnostic intervention for sleep and circadian problems in 31 adolescents with self-reported circadian misalignment (mean age 15.9, SD 1.3; 45.2% female). The design was pre-post within the same participants, with actigraphy, dim-light melatonin onset, questionnaires and a depression-anxiety-stress scale measured on both sides. 8
- Sleep midpoint shifted about 30 minutes earlier (p = .023, g = 0.591) and weekend sleep onset shifted significantly earlier (p = .003, g = 0.856).
- Subjective sleep quality improved (p = .021, g = 0.452).
- Dim-light melatonin onset, the circadian misalignment index, total sleep duration, sleep onset latency, chronotype and every mental health outcome measured showed no significant change. 8
The adolescents moved their sleep earlier, felt better about their sleep, and the biological phase marker did not follow. The study is a pilot with no control group and no follow-up, so maturation and the academic season are uncontrolled across the six weeks.
Pérez-Manchón and colleagues randomized 40 first-year nursing students in Madrid to a brief group cognitive behavioral therapy for insomnia or to their usual routine, monitoring circadian rhythm and sleep for seven days before and after. The largest effect was on self-reported sleep latency, which improved by 19.78 minutes (95% CI −33.83 to −5.74; partial η² = 0.194); self-reported total sleep duration (β = 36.32 minutes, 95% CI −5.01 to 77.64) and sleep efficiency (β = 8.75%, 95% CI −0.97 to 18.47) also favored the intervention, though both intervals cross zero. Objective circadian and sleep outcomes from the ambulatory monitor showed generally small effects, and 65% of the intervention group rated the program very good. 9
Placed side by side on one dimension, the dimension being what actually moved, the studies gathered here answer the same way: the measured change depends on how long the measurement ran and which quantity was extracted from it. The actigraphy study is the first half of that, because it shows a single night's error can be larger than the effect anyone is chasing. The rest-activity meta-analysis is the second half, because it shows that even a week of continuous data answers a different question depending on whether you compute amplitude, fragmentation or day-to-day steadiness from it. The two intervention pilots put numbers on the consequence: a program that reliably moves bedtime by half an hour can leave the circadian phase marker and the objective monitor where they were. A nightly tracker number is a sample with a known error; a behavior change is a change in the schedule, and the body reads the schedule on its own time.
Wearable-maker watch
No new research release from the three named device makers landed in this window.
WHOOP published a research post on September 18 on sleep and menstrual cycle regularity, drawn from the Stanford Wu Tsai Human Performance Alliance study the company co-authored. 10 This digest covered that study, its 2,596 participants and its 42,759 cycles in May, and the September post introduces no new dataset. The company's other posts in the window were a healthspan podcast episode on September 16 and a battery explainer on September 14.
Oura's newest post, on September 17, announced an expansion of its partnership with the primary care company Counsel through the CMS ACCESS model, and carries no sleep research. 11 Its September 11 data analysis of tennis, heart rate and recovery was covered in last week's issue. Eight Sleep's most recent post remains the September 3 Biological Age announcement, which this digest also covered. 12
From the researchers
Matthew Walker's podcast returned on September 14 with episode 151, on geography and sleep. The episode sets wall-clock time against biological time: latitude changes the seasonal pattern of daylight and melatonin, artificial light flattens that pattern, and wide time zones manufacture a social jetlag that is worst at a zone's western edge, where sunset arrives latest and sleep loss accumulates. Warm climates cut into deep sleep, daylight saving disrupts the clock twice a year, and people born in autumn or winter are likelier to be morning types. His three suggestions were morning light in winter, protecting bedtime if you live at the western edge of a time zone, and keeping the bedroom cool. 13 Walker states in each episode that he is not a medical doctor and that the content is not medical advice.
One seven-day action
Establish the error of your own device before you read anything into a night. Seven nights, one parameter, no changes.
- Choose one parameter, not the composite score: sleep efficiency (time asleep divided by time in bed) or wake after sleep onset. Write down which one and do not switch mid-week.
- Change nothing for seven nights. Keep your usual bedtime, caffeine and training. This week measures the instrument, not you.
- Record the value each morning and at the end of the week note the highest and the lowest.
- Compute the range by subtracting the lowest from the highest. That number is the noise floor of your setup at this time of year, and it will be larger than most of the effects you are trying to detect.
- Read the following weeks against it. Treat a change as signal only when it exceeds that range, and keep recording for at least another week before acting. A single night above or below the range is what an unchanged person looks like.
The procedure is the practical form of this week's evidence: the limits of agreement between two ways of measuring one night of sleep run to hours, so a number is worth acting on only once you know how much it moves on its own.
References
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10
- 11
- 12
- 13#151 - The Geography of Sleep
themattwalkerpodcast.buzzsprout.com
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
