What "validated against PSG" should tell you.
PSG — polysomnography — records multiple physiological signals for professional interpretation. Saying a product was compared with PSG is incomplete unless the claim also names the output, reference annotation, study design, dataset, metric, and limits.
Here is a practical checklist for reading SomniSense's evidence and other published comparisons.
The word that's doing all the work: "paired"
A strong comparison design can use paired-night data: the device and PSG recording the same person, on the same night, at the same time. The evaluation question must still match the claimed output. SomniSense's published production benchmark covers labeled 200-second window classification, not confirmation of each clinical event.
A comparison between unrelated group averages does not answer the same question as simultaneous paired recordings. Cohort size, setting, inclusion criteria, and underrepresented groups also matter, so these details should be available beside the result.
Reference scoring and blinding
Someone has to review the PSG and create the clinical reference annotations — that is "scoring," and it is done by trained technicians. The question is: did the scorer know what the app already said?
Blinded scoring helps reduce the risk that knowledge of the app output influences the reference annotations. A useful report states whether scoring was blinded and how the reference labels were produced.
"Agreement" is a spectrum, not a checkbox
Even with paired nights and blinded scoring, “agreement” is not a yes-or-no property. The right analysis depends on whether the task is segment classification, window classification, event detection, or measurement agreement.
For SomniSense, the public production benchmark reported here is classification accuracy and F1 on labeled 200-second breathing windows. It is not a Bland–Altman claim and must not be presented as validation of every event boundary, nightly BRI, or clinical AHI agreement.
How ours was actually run
Briefly, so you can hold it against the checklist above: 80 paired nights across 40 participants — some in a lab, most at home with a portable sleep-study rig — adults with and without known sleep-breathing issues, scored by AASM-trained technicians who were blinded to what SomniSense said. The full design, the cohort we didn't represent well, and the preprints behind it are on the research page; the number-by-number results are on the accuracy page.
Read next
- → The research behind SomniSense — the full validation study and the model it tested
- → How accurate is a phone at this, really? — why one metric cannot describe every task
- → The full accuracy breakdown — every metric, with its caveats
If you'd rather see your own night than read about someone else's —
It's free on both stores. Put the phone on the nightstand tonight; the report is there when you wake up.
Free basics stay available: core first-night candidate evidence, 5 audio plays per night, and a recent 7-day trend.