All articles
Applied Research8 July 2026 · 5 min read

Project Sonus: Real-Time Conversational Phonetics

A voice-native tutor that listens, speaks and coaches in real time: sub-second speech-to-speech conversation with phoneme-level pronunciation feedback.

D
Dr Abhishek Kumar
AI Research, ILM AI
Project Sonus: Real-Time Conversational Phonetics

Learning out loud

Most tutoring still happens through a keyboard. That works for some learners, but it quietly excludes a great many others: younger children who type slowly, language learners who think faster than they can spell, and anyone trying to revise while walking to the bus stop. Speaking is often the more natural channel, and explaining an idea aloud is one of the most reliable ways to actually understand it.

Project Sonus is our research into closing that gap. It is an ILM AI programme exploring what we call conversational phonetics: a spoken exchange fast enough to feel human, paired with live coaching on the sounds, rhythm and clarity of speech itself. The ambition is a tutor you talk with rather than type at, and one that listens closely enough to help you say things better.

This is an active research effort, not a finished product. We are sharing where the work is heading and the questions we are trying to answer.

The problem with talking to AI today

Voice is not new to educational technology, but the experience rarely feels good. Two failings show up again and again.

  • It is slow. Many voice tools stitch together separate steps for listening, thinking and speaking, and the pauses add up. A delay of a second or two breaks the rhythm of conversation and makes the whole thing feel like issuing commands rather than having a chat.
  • It is brittle. Recognition often assumes tidy, textbook speech. Real students hesitate, restart sentences, speak with regional accents and sit in noisy classrooms. Tools that stumble on this quickly become frustrating.

The result is that voice gets treated as a novelty rather than a serious mode of learning. We think it can be much more.

What conversational phonetics means

The name joins two ideas.

The first is conversation that keeps pace with a person. Our target is roughly 300 milliseconds of perceived speech-to-speech latency, the point at which a reply feels immediate rather than awaited. Hitting that consistently is a hard engineering problem, and it shapes almost every design decision in the project.

The second is phonetics, the science of speech sounds. Rather than treating a learner's voice only as a route to text, Sonus attends to how words are actually produced: the individual phonemes, the stress patterns, the intonation and the pace. That is what allows feedback to move beyond "what did you say" towards "how clearly did you say it".

What we are building

At its heart, Sonus is a voice-first tutor built as a low-latency speech-to-speech pipeline. A learner holds a natural back-and-forth, practises reading and explaining aloud, and receives gentle, specific feedback on pronunciation, fluency and delivery, right down to the individual phoneme.

The capabilities we are researching include the following.

  • Sub-second conversation. A spoken exchange quick enough to feel like talking to a person, not waiting on a machine.
  • Accent-robust recognition. Accuracy that holds up across regional accents, hesitations, false starts and everyday classroom audio, rather than only on clean, careful speech.
  • Speak to learn. Learners explain a concept aloud and get coached on their reasoning, using the act of teaching-back as a way to deepen understanding.
  • Phoneme-level feedback. Coaching on pronunciation, rhythm, intonation and pace, so a learner can hear exactly which sound or stress pattern to work on.
  • Hands-free and accessibility-first. Revision on the move, and a genuine lifeline for learners who find typing slow or difficult.

The research questions

Three areas occupy most of our attention.

Low-latency architecture

Reaching a conversational feel means rethinking how the pipeline is put together, so that listening, understanding and speaking overlap rather than queue. Shaving latency without sacrificing quality is the central technical challenge, and it runs through everything else.

Recognition on messy, spontaneous speech

Students rarely speak in neat sentences. They pause, correct themselves, mumble and trail off. We are researching recognition that stays accurate on this spontaneous, imperfect speech, because a tutor that only understands rehearsed lines is not much use in a real lesson.

Prosody-aware assessment

Words alone do not capture how someone speaks. Rhythm and intonation carry a great deal of meaning, and they are often where a learner most needs help. We are developing assessment that reads these prosodic features, and, importantly, that encourages rather than corrects. The aim is coaching that builds confidence, not a red pen for the voice.

Where it fits in ilmino

Sonus is being explored as a spoken mode for the ilmino AI tutor, sitting alongside the text-based experience rather than replacing it. Several uses look especially promising:

  • revising a topic by talking it through instead of reading;
  • rehearsing for oral exams and viva-style questioning;
  • practising spoken languages with immediate, sound-level feedback;
  • and accessibility-first learning for students who struggle with typing.

In each case the spoken mode is an addition to the tutor's range, giving learners another way in.

What it means for an institution

For schools, colleges and trusts, the appeal of a voice-native tutor is practical.

  • Inclusive access. A spoken interface reaches learners for whom typing is a barrier, widening who can benefit from independent study.
  • Oral practice at scale. Every student can rehearse speaking, reading aloud and explaining their reasoning as often as they like, without adding to timetabled staff time.
  • Staff oversight throughout. As with everything in ilmino, teachers stay in control. Sonus is designed to support and inform staff, not to sit unsupervised between a student and their learning. Human judgement remains central to how the tool is set up and used.

The point is not to replace conversation with a teacher, but to give learners far more low-stakes practice between those conversations.

Talk to us

Project Sonus is early, and we are being deliberate about it, because voice is intimate and getting the tone of feedback right matters as much as getting the latency down. If your institution cares about inclusive access and oral practice, we would welcome the conversation.

#Applied Research#Voice#Speech#Accessibility