A Sleep App's Cough Data Showed a One-Week Lead on Flu and COVID-19 PCR Positivity in Retrospective Analysis: What UKHSA and Sleep Cycle Published on 3 September 2026
Quentir Medicine Monitor
Evidence-based insights for quantum medicine. Published by Quentir Systems LLC · September 5, 2026.

On 3 September 2026 the UK Health Security Agency published the results of a joint study with Sleep Cycle, a Swedish sleep-technology company, reporting that increases in coughing at night were often seen about a week before increases in influenza and COVID-19 activity in England. The underlying data came from a consumer smartphone app that listens while people sleep, and the preprint states that no raw audio is transmitted to Sleep Cycle's servers.
The claim worth reading closely is the smaller one. Across three years of weekly data the nocturnal cough measures moved almost in step with NHS 111 triage calls for acute respiratory infection, and only secondarily did they show a one-week lead in retrospective analysis over PCR positivity for influenza and COVID-19. The agency and the authors are careful about what that supports. The study looked backward at three respiratory seasons that have already happened, so what it establishes is a relationship in recorded data, and the authors write that operational usefulness would need a prospective evaluation against forecasting benchmarks in a live setting.
What UKHSA and Sleep Cycle Published on 3 September 2026, and What the medRxiv Preprint of 21 July 2026 Contains
The agency's announcement of 3 September 2026 sits on top of a longer document. The study itself, posted to medRxiv on 21 July 2026 as "Nocturnal cough as a syndromic surveillance signal for respiratory illness in England", is written by seven researchers at the UK Health Security Agency in London and two at Sleep Cycle AB in Gothenburg: Tommy Irons, Emil Carlsson, Maria Tang, Jonathon Mellor, Conrad Rubin, Alex Allen, Alex J. Elliot, Mikael Kagebaeck and Josef Packham.
The design is a comparison, week by week, from January 2023 to January 2026. On one side sit three cough measures from the app: total cough count, coughs per user, and coughs per hour of sleep. On the other sit the agency's own surveillance indicators: NHS 111 triage calls for acute respiratory infection, PCR positivity for influenza and for COVID-19, and hospital admissions for influenza, COVID-19 and respiratory syncytial virus. Cough events are identified by a machine-learning audio model that runs locally on the user's phone, drawing inferences from overlapping ten-second clips during sleep sessions the user starts. No raw audio goes to Sleep Cycle's servers. Records are anonymized before transmission by removing personal identifiers and perturbing geographic coordinates, then aggregated to the seven NHS England regions.
The strongest relationships were with the NHS 111 triage calls. Population-normalized cough measures showed raw national correlations of about 0.95 with those calls and held prewhitened correlations above 0.55 at lag zero, which means the cough signal followed short-term movements in an established syndromic indicator beyond the shared seasonality and long-run trend that make almost any two winter series look alike. The preprint reports a one-week lead for coughs per hour of sleep over influenza PCR positivity, and the same lead for both coughs per user and coughs per hour of sleep over COVID-19 PCR positivity. That lead is the lag at which the prewhitened cross-correlation is largest, which is a different statement from saying the epidemic peaks themselves fell a week apart, and the paper's methods section says prewhitened results are the primary basis for interpretation. Hospital-based indicators tracked less well.
Quantum pillar: not applicable. Technology readiness: TRL 4 of 9. That is this Monitor's own assessment of the proposed surveillance application, not a rating the parties publish, and it is deliberately separate from the cough detector inside the app, which is shipping consumer software and sits far higher. On the shared ladder both Evidence Registers use, rung four covers retrospective validation on real recorded data, which is what has happened here: three years of already-collected weekly data compared against surveillance the agency had already banked. Running the measure live alongside a season, against a stated forecasting benchmark, is the next rung and has not been done. The detector itself uses a machine-learning audio model with no quantum technology of its own.
Why the 0.95 Correlation With NHS 111 Calls Matters More Than the One-Week Lead
A one-week lead is the headline, and the paper's own numbers make it the weaker half of the result. For COVID-19 the leading relationship is real and modest: coughs per user and the hourly average both show statistically significant peak correlations at a lag of minus one week, at rho 0.30 and 0.31. The same-week relationship with NHS 111 calls is roughly twice as strong, at prewhitened correlations of 0.55 and 0.56 at lag zero. Prewhitening is what makes the comparison meaningful, since it removes the shared seasonality, trend and autocorrelation that would otherwise make almost any two winter series look related, and the authors use the prewhitened results as their primary basis for interpretation.
What makes that correlation useful is a property the app has and the existing systems do not. The healthcare-based surveillance indicators compared here all begin with a person deciding to seek care, and the agency notes that this decision is shaped by public awareness, by how available services are, and by demographic and socioeconomic differences, before laboratory processing and reporting add their own delay. The cough measure is generated automatically during ordinary sleep and refreshes daily. The preprint puts the reporting lag at under a day, against a normal lag of several days for the agency's existing sources. Whether a faster indicator is also a good enough one is the question a prospective evaluation would answer; the retrospective association reported here does not settle it.
The choice of measure turns out to carry most of the analytic weight. Unnormalized total cough counts produced lag structures the authors describe as less stable and often uninterpretable, which they attribute to sensitivity to changes in observation volume: the number of active users and the recorded sleep duration both move for reasons that have nothing to do with respiratory illness. The preprint's recommendation is to use population-normalized measures rather than raw counts for any surveillance application. That is a small sentence with a large operational consequence, since a dashboard built on the raw count would drift with the app's commercial fortunes.
What the Study Does Not Show: a 37.9-Year-Old Urban User Base, a Backward-Looking Design, and an Inconclusive RSV Result
The authors are explicit about three limits, and each one bites in a different place. The user population skews younger and urban. The dataset averaged 3,482 daily users in the South West of England against 11,427 in London, and among users who volunteered demographic information the mean age was 37.9 years, with 59.3 percent men and 40.2 percent women. Severe respiratory illness concentrates in infants and in older adults, so the people the app hears most are not the people who fill respiratory wards.
The second limit is the design. This is a retrospective study of temporal association, not a test of predictive performance, and the authors say demonstrating operational utility would require prospective evaluation against forecasting benchmarks in a real-time setting. The third is respiratory syncytial virus, where the results were inconclusive. The authors attribute that to a shorter available time span and to the distance between the app's user base and the groups that suffer severe RSV disease, which are infants and the very old. The study therefore has not established usefulness for RSV, which is a different statement from having shown the signal to be absent there, and a dashboard carrying this measure should say which of the three viruses it has evidence for.
There is also a dependency worth stating plainly. The signal exists because a commercial company has a large voluntary user base and chose to make anonymized aggregates available. Under the collaboration the agency announced on 28 January 2026 as a twelve-week project, no UKHSA data went to Sleep Cycle; the analysis ran on the agency's own secure systems with a dedicated research team. That is a sound arrangement for a study. It is a thinner foundation for a standing national indicator, whose continuity would depend on one company's product decisions.
What a Public Health Agency or Hospital Should Watch Between Now and the Next Respiratory Season
Four things will resolve on their own schedule. Whether the prospective evaluation the preprint names as the next step is actually run, and against which forecasting benchmark. Whether the signal is integrated into the agency's existing surveillance dashboards, since integration is where a research correlation either earns operational trust or does not. Whether the age and geographic skew is corrected by weighting or simply disclosed. And whether the arrangement acquires a durable form, because a national indicator resting on a single vendor's voluntary contribution is a different asset from one resting on an instrument the agency owns.
How Quentir Reads It
The pattern is familiar from this year's other health-data arrivals. When OpenAI connected ChatGPT to Epic charts on 1 September 2026 with read-only access, the limit that made the deployment discussable was written into the access itself rather than promised alongside it. The equivalent limit here is architectural: inference runs on the phone, the preprint states that raw audio is not transmitted, records are anonymized and coordinates perturbed before aggregation to seven regions. Whether that architecture holds in practice depends on implementation, on consent and on the governance around the app, none of which this study examines, so it is a design worth reading carefully rather than a settled privacy result.
The gap the study leaves open is the one this Monitor has been following across several instruments, where a requirement gets named before the test that would satisfy it exists. The agency now has a signal with a published retrospective association and no prospective trial behind it. Another winter of data becomes that trial only if someone writes the evaluation protocol first, with a stated forecasting benchmark and a decision rule for what would count as a failure. Without that, the next season produces a fourth year of correlations and the same open question.
Sources
Primary source: Tommy Irons, Emil Carlsson, Maria Tang, Jonathon Mellor, Conrad Rubin, Alex Allen, Alex J. Elliot, Mikael Kagebaeck and Josef Packham, "Nocturnal cough as a syndromic surveillance signal for respiratory illness in England," medRxiv preprint posted 21 July 2026, DOI 10.64898/2026.07.20.26357937, for the study design, the January 2023 to January 2026 window, the three cough measures, the comparison indicators, the correlation of about 0.95 with NHS 111 acute respiratory infection triage calls, the prewhitened correlations of 0.55 and 0.56 at lag zero, the one-week lead over influenza and COVID-19 PCR positivity together with the peak correlations of rho 0.30 and 0.31 at a lag of minus one week for COVID-19, the seasonal ARIMA prewhitening and the statement that prewhitened cross-correlations are the primary basis for interpretation, the on-device inference on overlapping ten-second clips, the anonymization and regional aggregation, the daily user counts of 3,482 in the South West and 11,427 in London, the mean age of 37.9 years and the 59.3 and 40.2 percent split, the sub-day reporting lag, the instability of unnormalized counts, and the stated limitations including the inconclusive RSV result. The UK Health Security Agency's announcement of 3 September 2026 supplies the public framing of the one-week early warning; its earlier announcement of 28 January 2026 supplies the twelve-week project structure and the arrangement under which no agency data was shared with Sleep Cycle. The judgments are this Monitor's own: the TRL 4 placement, the reading that the lag-zero correlation is sturdier than the retrospective lead time, and the four items a buyer can check.