A sleep-tracker study found duration transferred better than sleep stages
Researchers compared 752 paired home nights from 74 older adults across a watch, nearable sensor, research actigraphy and diary. Most single-night measures correlated weakly, and a stage readout is not a diagnosis.

A night of sleep can return in the morning as a tidy architecture: a start time, a total, an efficiency score and coloured blocks labelled as stages. The precision is persuasive. It can also conceal an awkward question. Would another device describe the same night in the same way?
A paper published online in the journal SLEEP on 22 August tested that question across four ways of tracking sleep at home. The researchers collected 752 paired nights from 74 older adults, including 20 people living with dementia. Participants simultaneously used research-grade actigraphy, a consumer wristwatch, an under-mattress sleep analyser and a sleep diary, then completed one assessment with in-lab polysomnography.
Most measures did not travel cleanly from one modality to another. Single-night correlations between devices, and correlations with polysomnography, were weak for most measures. Duration and timing measures did better, reaching moderate correlations. Detailed sleep-stage durations showed poor agreement.
That does not make every wearable useless. It changes what the morning dashboard can reasonably mean. A device may be a consistent lens on a person's routine while still disagreeing with another lens about how the night should be divided.
The study's clearest cross-device result was broad rather than granular. After the researchers filtered measures for reliability and reduced them into larger components, duration was the only sleep aspect with consistent moderate associations between devices.
The more detailed categories were less portable. In comparisons with polysomnography, the tested modalities systematically overestimated total sleep time and sleep efficiency, underestimated wake after sleep onset, and agreed poorly on sleep-stage durations. Those directions describe this study's tested methods and sample. They are not a ranking of every current watch, ring, phone app or bedside sensor.
Correlation also needs careful reading. Two systems can rise and fall together across nights while still producing different absolute values. The researchers used Bland-Altman analysis alongside correlations precisely because moving in the same direction is not the same as agreeing on the number. The systematic differences from polysomnography show why a seemingly similar trend cannot make the outputs interchangeable.
This is where the visual polish of a sleep dashboard can outrun the measurement. Minutes assigned to light, deep or REM sleep look like shared biological units. The new findings suggest that, for the tested technologies, some of that apparent precision remained tied to the modality producing it.
One of the study's useful complications is that repeatability improved with repetition. The number of nights needed to reach the researchers' threshold for acceptable reliability varied by measure, modality and participant group. Within seven nights, 71% of the measures reached that threshold.
Reliability here does not mean clinical accuracy or agreement between brands. It means a measure became stable enough to characterise differences between participants within the study's repeated-measure framework. A bathroom scale that is always offset can still be repeatable; consistency and correctness answer different questions.
For a consumer, the distinction favours trends over isolated verdicts. If a tracker is being used to notice bedtime drift, changes in total duration or a recurring interruption, several nights on the same system may be more interpretable than one unusually bad score. Switching devices and then comparing exact stage minutes as though the categories were identical is harder to justify from this evidence.
That is an inference for reading personal data, not a clinical protocol. The study did not test every product, prescribe a minimum tracking period for consumers or establish that seven nights is a universal rule. Its sample was limited to 74 older adults, and one group was living with dementia. Results may differ in younger people, people with other sleep conditions and newer devices.
The disagreement is easier to understand once the instruments are separated. The US National Heart, Lung, and Blood Institute describes polysomnography as a sleep study that records signals including brain waves, heart rate, breathing and blood oxygen. Sensors may also track eye and muscle movements. Those signals allow trained scorers and algorithms to classify sleep using a much richer physiological record.
Many wearables work from a narrower view. A 2024 scoping review in npj Digital Medicine examined 35 studies and 62 wearable setups. It found that accelerometer-only devices were effective for sleep-versus-wake detection but fell short when identifying multiple stages. Combining movement with photoplethysmography, the optical pulse signal used by many wearables, was the growing approach, but validation methods, participant groups and proprietary algorithms still varied.
A watch is therefore not a miniature sleep laboratory. It infers sleep states from the signals available at the wrist and from the rules encoded in its software. An under-mattress sensor observes another slice of the night. A diary records remembered experience. Each can be useful, but they do not begin with the same raw evidence.
That also explains why two dashboards can disagree without either one capturing a simple lie. They may define, sense and smooth the night differently. The mistake is treating their polished categories as if they were direct, device-independent observations.
The American Academy of Sleep Medicine has long advised that consumer sleep technology should not replace validated diagnostic testing or a comprehensive sleep evaluation. Its position statement also recognises a constructive role: tracker data can support a conversation between a patient and clinician when interpreted in the context of symptoms and proper evaluation.
A practical reading order follows from the evidence:
- Start with the question. Bedtime, wake time and broad duration are different claims from exact minutes in REM or deep sleep.
- Compare like with like. A repeated pattern from one device is not automatically comparable with the same label from another device.
- Keep the display's precision in proportion. A coloured segment is an algorithmic classification, not a direct view of the brain.
- Put persistent symptoms ahead of a score. Ongoing sleep problems, excessive daytime sleepiness or fatigue deserve medical attention regardless of whether a wearable labels the night good or bad.
The new paper does not ask people to throw away their trackers. It asks for a narrower confidence. Across the tested modalities, duration was the part of sleep that transferred best. The internal map of the night remained much more dependent on who, or what, was doing the measuring.
Editorial note. This article reports consumer sleep-technology research for general information. It does not diagnose a sleep disorder, interpret an individual's tracker data or provide medical treatment advice.
Sources
- **Ravindran et al., SLEEP, Crossref record and abstract:** Verifies the 22 August 2026 publication, 74-person sample including 20 people living with dementia, 752 paired nights, four simultaneous home modalities, later polysomnography, correlation and agreement results, seven-night reliability result, and duration conclusion
- **Birrer et al., npj Digital Medicine scoping review:** Verifies the review of 35 studies and 62 wearable setups, the distinction between accelerometer-only and multi-sensor devices, the weaker performance of movement-only devices for multiple stages, and variation in algorithms and validation
- **US National Heart, Lung, and Blood Institute, Sleep Studies:** Verifies what polysomnography records, how sensors are used and the role of sleep studies in diagnosing sleep disorders
- **American Academy of Sleep Medicine position statement:** Verifies that consumer sleep technology should not replace validated diagnostic testing or comprehensive evaluation, while patient-generated data may support a clinician conversation
Help us improve
Was this article useful?
One anonymous tap helps Sona improve future reporting, headlines and source context.
Up next

The analysis covered 204,907 adults, but it measured whether relationships felt satisfying, not friend counts, and one survey wave cannot show which came first.
Continue reading

