VALIDATED ACCURACY · METHODOLOGY OPEN
How accurate is SomniSense — really?
A single rounded accuracy number hides the task. We publish separate results for labeled 1-second snore segments and labeled 200-second breathing windows, plus the on-device model size and latency. Production event boundaries, nightly indices, and personal trends do not inherit those benchmark results.
WHAT SOMNISENSE ADDS
Connect the benchmark to the product output it does — and does not — support.
The real SomniSense app screen shows the production evidence unit: a model-estimated candidate type, time and duration, marked waveform, available audio, and whole-night position. The published 200-second window result supports its own classification task; it does not independently validate each production event boundary, duration, personal BRI, or diagnosis.
- Candidate pause-like or shallow-breathing-like type, time, and model-estimated duration
- The marked waveform region and locally available audio for review
- Whole-night position, followed across nights in recent comparisons and structured reports
Phone audio is a limited wellness signal. Candidate regions are not clinical apnea events, AHI, oxygen measurements, airway-source findings, or a diagnosis.
Why the task definition comes first
The current research is available as open preprints and has not yet completed peer review. Each headline number belongs to a specific labeled task, dataset, and evaluation procedure.
We therefore state the task beside the metric and avoid carrying a benchmark result over to event boundaries, durations, product labels, individual bedrooms, or clinical diagnosis. If later peer review or expanded validation changes a result, this page will be updated.
The single-number problem
A single rounded “accurate” percentage is incomplete until you ask: accurate at what, on what population, and against which reference labels? Snore-segment classification, breathing-window classification, production event boundaries, and a whole-night index are different questions with different evidence.
SomniSense breaks accuracy into several numbers, each tied to a specific question. No one benchmark tells the whole product story.
Two indices, one app — and why the validation numbers split that way
SomniSense produces two product-defined indices each night. Their underlying model benchmarks evaluate different tasks:
- BRI (Breathing Irregularity Index) — candidate breathing-irregularity acoustic patterns per estimated sleep hour. The underlying production benchmark reports 88.49% accuracy and 88.06% F1 for 200-second window classification, plus 56.4 KB / 0.064 ms inference on Apple M2 Neural Engine.
- SRI (Snoring Rate Index) — snore events per hour. The underlying 1-second snore-event benchmark reports 91.67% sensitivity and 89.01% precision.
Both BRI and SRI are product-defined wellness indices. BRI is not AHI, and the window-classification benchmark does not independently validate every event used to assemble a nightly index.
What our research measures, and what BRI actually does
Three different things, often mixed up. We try not to mix them:
- Window classification (what Paper E validates). Given a labeled 200-second window of bedside audio, the model classifies it as normal or apnea-or-hypopnea. Production benchmark: 88.49% accuracy and 88.06% F1.
- Per-night BRI (what the app reports each morning). The app aggregates candidate production events across the session:
BRI = candidate breathing-irregularity events / estimated sleep hours. This is a product-defined acoustic rate, not a clinical AHI. - OSA (Obstructive Sleep Apnea) diagnosis — what a sleep specialist does, not what SomniSense does. A clinical diagnosis based on polysomnography, daytime symptom assessment, and a physician's judgment. BRI is data; OSA diagnosis is a doctor's call.
Event start, event end, duration, product label, nightly BRI, and personal trends are production outputs. They should not be described as independently validated by the 200-second window result unless a separate published analysis evaluates that exact output.
The benchmark numbers we actually publish
| The question | The number | What that means |
|---|---|---|
| Of labeled snore segments, how many did the detector catch? | 91.67% | Snore-event sensitivity (recall) for 1-second segments; five-seed mean. |
| When the detector flagged a snore segment, how often was it right? | 89.01% | Snore-event precision for 1-second segments; five-seed mean. |
| How often did the compressed breathing model classify a labeled window correctly? | 88.49% | Production benchmark accuracy for normal vs apnea-or-hypopnea classification on 200-second windows. |
| How balanced was breathing-window performance across classes? | 88.06% | F1 for the same compressed production benchmark. Paper E does not publish a production sensitivity/precision pair, so we do not infer one. |
The research corpus includes 80 paired PSG nights across 40 participants (10 in-lab + 70 ambulatory PSG with nasal-airflow cannula). See the linked preprints for which subset, labels, and evaluation procedure apply to each reported metric.
That last part — "didn't know what we said" — is what "blinded scoring" means. We don't get to pre-train our scorers on our own answers. Otherwise the test would be circular.
How we tested it (the methodology, plain)
Detection performance was measured against PSG reference recordings from both in-lab and ambulatory nights, with audio annotations by certified sleep technicians. Every audio segment they scored was blinded to SomniSense output. That separation keeps the evaluation independent.
Specifically:
- Sample: 80 paired nights / 40 participants (smartphone + PSG simultaneously) — 10 in-lab PSG + 70 ambulatory PSG with nasal-airflow cannula. Adults with and without diagnosed sleep breathing concerns. We state who's underrepresented (mostly: under-18, severe BMI extremes, certain ethnic groups) openly, rather than hide it behind an aggregate number. I want to be specific about who's not in the cohort because aggregate sample size without context can be misleading.
- Recording medium: bedside smartphone (varied iPhone & Android models from 2018 onward), 50–90 cm from participant's head.
- Ground truth: PSG with synchronized audio channel; manual scoring by AASM-trained sleep technicians, blinded to SomniSense output.
- Published comparisons: task-specific snore-segment and breathing-window classification metrics. Production event boundaries and nightly indices require their own evaluation and are not silently inherited from those results.
The breathing-event detection algorithm builds on years of sleep apnea research by our founder. The version powering SomniSense was retrained from scratch and rebuilt for SomniAI LLC to handle smartphone audio specifically — different microphone, different distance, different acoustic context than clinical hardware. Three companion preprints document the methodology in full: the cascaded-baselines preprint (multi-seed bootstrap), the Coordinate-Attention 1D architecture preprint (93.2% parameter reduction), and the on-device compression preprint (0.064 ms inference on Apple Neural Engine). All are published openly on Zenodo with citable DOIs (cs.LG, eess.AS). For how the two-stage system fits together and the full preprint portfolio, see the research program. The full technical hub — the architecture in depth, the SDK, and licensing for hardware and clinical partners — lives at apneasense.com/research.
Honest limitations
Here's what we don't know yet, and what I'd want to know if I were the user:
- The cohort does not represent everyone. Performance for an individual or an underrepresented population may differ; we do not assume the published aggregate is conservative for any particular person.
- Acoustic environment matters. If your bedroom has unusual acoustic properties — hard surfaces, partner snoring louder than you, a fan blowing directly at the phone — the model may catch fewer of your events. The methodology paper documents the conditions we tested under.
- Production outputs need output-specific validation. A window benchmark cannot establish the accuracy of every event boundary, duration, label, BRI value, or trend shown in the app.
- Not validated for under-18. The cohort was adults only.
- Preprints, not yet peer-reviewed. Three companion preprints are published openly on Zenodo with citable DOIs; peer-reviewed journal publication is a separate process and will be noted on this page when complete.
What this isn't
- Not a diagnostic claim. Even at these numbers, SomniSense is not a medical device and doesn't diagnose obstructive sleep apnea (OSA). OSA diagnosis requires polysomnography, symptom assessment, and a sleep specialist's judgment. SomniSense isn't validated for users under 18.
- Not a personal guarantee. Your specific results may differ from population averages. Read the methodology paper to know whether your scenario is in or out of distribution.
- Not a replacement for a sleep study. Do not apply clinical AHI thresholds to BRI. Symptoms or concerns deserve professional evaluation regardless of the app's value.
Not sure if this is your problem? Start from a symptom
If you came here to check the evidence before trusting the app, the other way in is whatever you're actually feeling. Each of these walks through the breathing pattern behind the symptom and what to do about it:
If this is the level of evidence that satisfies you
Free keeps core first-night candidate evidence, 5 audio plays per night, and a recent 7-day trend. Pro adds continuity, not the basic right to inspect the first night. Purchase details stay on the pricing page.
Start with Free Review each benchmark and its boundary →