Culture & Society
A 2025 Oxford-published study in SLEEP Advances tested six consumer wearables against clinical polysomnography, revealing consistent over- and underestimation of sleep stages—and U.S. medical guidelines explicitly advise against relying on wrist-based devices for blood pressure measurement.

Consumer smartwatches may display precise-looking metrics—“1 hour 12 minutes of deep sleep,” “94% blood oxygen,” “118 bpm”—but those numbers are not direct physiological measurements. Instead, they are algorithmic estimates derived from optical and motion sensors, not clinical-grade instruments. A growing body of peer-reviewed evidence confirms that high price does not equate to medical accuracy.
Most smartwatches and fitness bands rely on photoplethysmography (PPG), a non-invasive optical method that emits light into the skin and detects changes in reflected light caused by blood flow. Algorithms then combine PPG signals with accelerometer data and other inputs to infer physiological states. While pulse rate correlates closely with the raw PPG signal, more complex parameters—including blood pressure and sleep staging—require indirect estimation, not direct sensing.
Conventional sphygmomanometers measure arterial pressure directly via an inflatable cuff that occludes and then gradually releases pressure on the brachial artery. In contrast, cuffless wrist devices estimate blood pressure using PPG waveforms, pulse transit time, and machine learning models. According to a scientific statement issued by the American Heart Association (AHA), such devices typically report relative changes in blood pressure compared to a prior calibration reading—not independent, absolute values.
The 2025 AHA/American College of Cardiology (ACC) Hypertension Guideline explicitly recommends against using cuffless devices—including smartwatches—for clinical blood pressure assessment until higher levels of accuracy and reliability are demonstrated. The U.S. Food and Drug Administration (FDA) has also warned consumers against relying on unapproved smartwatch features or wearable rings that claim to measure blood pressure, noting that inaccurate readings could lead to missed diagnoses of hypertension or hypotension—or inappropriate treatment decisions.
Clinical gold-standard sleep assessment is polysomnography (PSG), which records multiple physiological signals: electroencephalography (EEG) for brain activity, electrooculography (EOG) for eye movement, and electromyography (EMG) for muscle tone. Consumer wearables lack EEG, EOG, and EMG capability. Instead, they infer sleep stages solely from wrist motion, heart rate variability, and peripheral PPG-derived signals.
A 2025 study published in SLEEP Advances—a journal affiliated with the University of Oxford—directly compared six widely used consumer devices—Apple Watch Series 8, Fitbit Sense, Fitbit Charge 5, Whoop 4.0, Garmin Vivosmart 4, and Withings ScanWatch—against PSG in 62 adult participants.
All six devices achieved over 90% sensitivity in detecting overall sleep periods. However, specificity—the ability to correctly identify wakefulness—varied sharply: Garmin Vivosmart 4 correctly identified wake periods only 29% of the time; Apple Watch Series 8 reached 52%; Fitbit Sense scored 49%; and Fitbit Charge 5 scored 48%. Researchers attributed this limitation to the fundamental ambiguity between immobile wakefulness and sleep—both produce similar motion and heart-rate signatures at the wrist.
Differences extended beyond wake/sleep classification. Compared to PSG, Whoop 4.0 overestimated deep sleep by an average of 31 minutes and REM sleep by 15 minutes. Apple Watch Series 8 underestimated deep sleep by 25 minutes and REM sleep by 13 minutes. For light sleep, Apple Watch Series 8 overestimated duration by 59 minutes, while Withings ScanWatch underestimated it by more than 33 minutes.
These deviations confirm that displayed sleep-stage durations remain algorithmic approximations—not objective physiological measurements. No device in the study was labeled categorically “inaccurate”; rather, all produced estimates subject to systematic bias relative to the clinical reference standard.
A separate 2026 SLEEP Advances study evaluated consumer sleep trackers—including Fitbit and Oura—across age groups. It found persistent limitations in estimating sleep onset timing, particularly among older adults. Devices consistently estimated sleep onset earlier than PSG-determined onset in older participants—a tendency that inflates total reported sleep duration. The findings underscore a core constraint: algorithms trained on generalized population data cannot fully account for inter-individual variation in physiology, movement patterns, or health status.
Smartwatch pricing reflects engineering complexity—not sensor fidelity alone. Costs include high-resolution displays, multi-core processors, GPS modules, wireless connectivity stacks, water resistance, premium materials, operating system development, and years of algorithm refinement. Companies invest millions in improving both hardware and software. Yet no amount of investment overcomes a foundational physical limitation: algorithms cannot infer signals that sensors do not capture.
If a watch lacks EEG, EOG, and EMG, sleep staging remains an inference—not a measurement. If it lacks cuff-based arterial occlusion, blood pressure remains an estimate—not a reading.
The same 2025 SLEEP Advances study concluded that certain devices—including Apple Watch Series 8, Fitbit Sense, and Fitbit Charge 5—can reliably detect large-scale, longitudinal shifts in sleep behavior. For example, a sustained decline from a habitual seven-hour average to five hours over several weeks may signal meaningful change—even if nightly deep-sleep estimates vary by dozens of minutes.
This pattern-tracking utility applies across biometric domains. Smartwatches serve best as trend-monitoring tools and early-alert systems—not diagnostic replacements for clinical equipment. Their value lies not in absolute precision, but in consistency of relative change over time.



