How accurate is a phone at this, really?
It's a fair thing to be skeptical about. A clinical sleep study straps a dozen sensors to you — airflow, oxygen, chest movement, brain waves. SomniSense uses one microphone, sitting on the nightstand, two feet away. So when someone asks me "how accurate can that possibly be," I don't think they're being rude. I think they're asking the right question.
The honest answer is task-specific. A microphone model can be evaluated for a defined acoustic classification task, while PSG measures physiological signals for clinical interpretation. Those are different jobs, so one number cannot describe the whole product.
Why an accuracy number needs a named task
A large percentage is incomplete unless it says accurate at what. Classifying a one-second snore segment, classifying a labeled 200-second breathing window, and producing a whole-night product summary are different tasks. Evidence for one must not be transferred to another.
The two numbers that actually matter pull in opposite directions:
- Sensitivity — among the reference annotations in the evaluation set, how many did the model detect?
- Precision — among the model's detections in that evaluation set, how many matched a reference annotation?
Sensitivity and precision can move in different directions. Publishing both helps readers understand the tradeoff instead of treating a single blended metric as a complete description.
What our numbers actually are
So, plainly. For 1-second snore-event detection, the five-seed means are 91.67% sensitivity and 89.01% precision. For 200-second breathing-event windows, the compressed model reports 88.49% accuracy and 88.06% F1. Those are research-cohort classification benchmarks, not a personal accuracy guarantee for tonight's SRI or BRI.
I'm deliberately not going to unpack every one of those here, because that's a longer, more careful read, and I'd rather you have it in full than in a blog-sized summary. The number-by-number version lives on the accuracy page, and the study design and the model behind it — how a phone gets to those numbers at all — is on the research page.
The honest part: where it's weaker
Production performance can vary with phone placement, room noise, microphone, partner sounds, and patterns near a decision boundary. The published 200-second window benchmark does not establish a one-sided guarantee for an individual's BRI.
Noise, another sleeper, a fan near the microphone, and unsupported ages can all affect the fit between a benchmark and a personal night. The published cohort was adults; SomniSense is not for users under 18.
Read next
- → The research behind SomniSense — the validation study and the on-device model, in plain language
- → What "validated against PSG" should tell you — why the testing method matters as much as the number
- → The full accuracy breakdown — every metric with its task and limits
If you'd rather see your own night than read about someone else's —
It's free on both stores. Put the phone on the nightstand tonight; the report is there when you wake up.
Free basics stay available: core first-night candidate evidence, 5 audio plays per night, and a recent 7-day trend.