Skip to content
All work

Healthcare

Canary Speech: Reading a Voice for What It Gives Away

Prepared by Jithin

Photograph by Pawel Czerwinski on Unsplash

Ask someone how they are and they will say fine. Listen to how they say it — the pace, the pauses, the pitch, the effort behind the words — and you learn something the answer withheld.

That is the premise Canary Speech is built on. Their technology reads vocal biomarkers in real time, surfacing signals of stress, fatigue, depression and cognitive impairment for healthcare providers, clinical researchers, and organisations working on population health and workforce wellbeing.

iLeaf built the applications that collect the voice. We were part of their early team — years before TIME named them to the first edition of the World's Top HealthTech Companies, one of 400 recognised out of thousands reviewed.

The Challenge

An application that gathers clinical-grade audio is not a recording app with a nicer button. Almost everything difficult about it sits between the microphone and the model.

The signal is fragile and the environment is not controlled. These recordings happen in a clinic room, a care home, an office — wherever the person is. Room noise, distance from the device, a hand moving across the microphone, the automatic gain control most consumer devices apply by default: each of them alters the very features the analysis depends on. A recording that sounds perfectly clear to a human can be useless for biomarker extraction.

Consistency matters more than quality. Vocal biomarkers are read as change over time — this person, compared against this person previously. That makes consistency between sessions worth more than absolute fidelity. Two recordings of the same person on the same device must differ because they differed, not because the device did something different.

The people using it are not technicians. A clinician has minutes, not patience for audio settings. A participant may be unwell, elderly, or cognitively impaired — the group whose data matters most is the group least able to tolerate a fiddly interface. Anything requiring a retake costs more than the retake.

And it is health data from the first second. A voice recording is biometric and identifiable. Everything about how it is captured, held on the device and transmitted has to assume that from the outset rather than have it added after a review.

What We Built

Canary Concussion, and the iPad and mobile applications the platform was first delivered on.

Capture built for the analysis, not for the ear. Audio configured for what the models need rather than for what sounds pleasant — consistent sample rate and format, device processing that would otherwise flatten the signal disabled, and the recording path kept identical from session to session so that a difference in the data is a difference in the person.

Guided sessions. A protocol a clinician can run without training and a participant can complete without instruction: prompts in a fixed order, timing handled by the application, and the awkward moments — a false start, background noise, someone stopping to cough — handled rather than left to a person to notice.

Checks at the point of capture. A problem found back at the desk means the participant has gone home. So the quality of the recording is assessed while they are still in the room and still able to repeat it, which is the only moment a retake is cheap.

The tablet as the right instrument. An iPad is stable on a table, has a screen a clinician and a participant can both see, and is unremarkable in a clinical setting in a way that a laptop or a microphone rig is not. The interface was built for one to be held out and spoken to, not operated.

The Approach

The recording is the product. Everything downstream — the models, the biomarkers, the clinical value — is limited by what the microphone captured. That inverts the usual priority order, and it is why the unglamorous audio path took the effort the interface would normally get.

Design for the least able user, not the average one. Cognitive impairment is one of the things this measures, which means people experiencing it are among the intended users. An interface that assumes a well, unhurried, technically confident person excludes the population the product exists to serve.

Assume health data throughout. Voice is biometric. Capture, on-device storage and transmission were built on that assumption from the start, rather than retrofitted when someone asked.

Build for repeat, not for demonstration. A single impressive recording proves nothing here. The value appears across sessions over months, so the thing worth engineering is that session forty is directly comparable with session one.

The Result

Canary Speech was named to TIME's World's Top HealthTech Companies in 2025, in the first edition of that list — one of 400 recognised from thousands reviewed.

We are on this page because of when we were there, not because of that award. iLeaf was part of the early team, building the applications the platform ran on well before the recognition arrived. Being the engineering partner a healthtech company had before it was one of the world's top healthtech companies is a different claim from having worked with one afterwards.


More on how we build for healthcare — including a clinical AI platform running entirely on a hospital's own GPUs — or tell us what you are building.

Have a system that needs to do this?

500+ systems shipped since 2011, and we still maintain most of them. Tell us what you are trying to move.