Foundational · Jun 9, 2026 · 6 min read ·

How accurate is a phone at this, really?

A phone resting on a nightstand at night, its microphone facing a sleeping person — the only sensor SomniSense uses to estimate breathing irregularities.

It's a fair thing to be skeptical about. A clinical sleep study straps a dozen sensors to you — airflow, oxygen, chest movement, brain waves. SomniSense uses one microphone, sitting on the nightstand, two feet away. So when someone asks me "how accurate can that possibly be," I don't think they're being rude. I think they're asking the right question.

The honest answer is task-specific. A microphone model can be evaluated for a defined acoustic classification task, while PSG measures physiological signals for clinical interpretation. Those are different jobs, so one number cannot describe the whole product.

Why an accuracy number needs a named task

A large percentage is incomplete unless it says accurate at what. Classifying a one-second snore segment, classifying a labeled 200-second breathing window, and producing a whole-night product summary are different tasks. Evidence for one must not be transferred to another.

The two numbers that actually matter pull in opposite directions:

  • Sensitivity — among the reference annotations in the evaluation set, how many did the model detect?
  • Precision — among the model's detections in that evaluation set, how many matched a reference annotation?

Sensitivity and precision can move in different directions. Publishing both helps readers understand the tradeoff instead of treating a single blended metric as a complete description.

What our numbers actually are

So, plainly. For 1-second snore-event detection, the five-seed means are 91.67% sensitivity and 89.01% precision. For 200-second breathing-event windows, the compressed model reports 88.49% accuracy and 88.06% F1. Those are research-cohort classification benchmarks, not a personal accuracy guarantee for tonight's SRI or BRI.

I'm deliberately not going to unpack every one of those here, because that's a longer, more careful read, and I'd rather you have it in full than in a blog-sized summary. The number-by-number version lives on the accuracy page, and the study design and the model behind it — how a phone gets to those numbers at all — is on the research page.

The honest part: where it's weaker

Production performance can vary with phone placement, room noise, microphone, partner sounds, and patterns near a decision boundary. The published 200-second window benchmark does not establish a one-sided guarantee for an individual's BRI.

Noise, another sleeper, a fan near the microphone, and unsupported ages can all affect the fit between a benchmark and a personal night. The published cohort was adults; SomniSense is not for users under 18.

If you'd rather see your own night than read about someone else's —

It's free on both stores. Put the phone on the nightstand tonight; the report is there when you wake up.

Free basics stay available: core first-night candidate evidence, 5 audio plays per night, and a recent 7-day trend.