Source: The Conversation (Au and NZ)

A few months ago, a friend posted something that caught my eye on social media. She struggled listening to what she termed “(literally) breathless AI voices for audiobooks”.
When I asked her what part she found the most off-putting, she said she had to stop listening because there was no natural breathing or variation – it’s just too formulaic.
This might seem like a small complaint. But it points to a much bigger question about language and artificial intelligence: what happens when we design systems to produce language without paying enough attention to the people who listen to it or read it?
Communication does not begin and end with the speaker. Listeners work just as hard: attending not just to what someone says, but how they say it, what has just happened and what we expect might happen next. Even what we regard as imperfections in speech can carry information.
Disfluencies have meaning
Linguists often use the word “disfluencies” to refer to hesitations, repetitions, false starts and other departures from “fluent” speech.
The label can make them sound like defects that ought to be removed. But in human conversation, they are part of the message.
A synthetic voice that strips out disfluency may be grammatically accurate and acoustically flawless, yet feel impoverished to the listener, who simply has less to work with.
We are not simply decoding signs
Language does not exist in our minds merely as a catalogue of words and how they are pronounced and used. It is entangled with memory, experience and our knowledge of the world, shaped by where we grew up, with whom, doing what and for what purpose.
Award-winning author Arundhati Roy captures something of this in her memoir Mother Mary Comes To Me when she describes most of us as “a living, breathing soup of memory and imagination”.
Roy’s phrase is useful for thinking about language, because human communication is not a clean transfer of information from one cognitive container to another. Words encounter another person’s memories, expectations and assumptions. They also encounter a listener’s other languages as well as less tractable issues such as how they’re feeling on the day.
This is one reason two people can hear the same sentence and understand it differently, and why humour is so nuanced. It is also why the social characteristics of a voice matter.
Berkeley linguist Nicole Holliday’s research shows how listeners make social judgements about voices, including AI voices, and how those judgements interact with perceptions of personality and identity. The listener is crucial to communication, and that creates an interesting problem for generative AI.
We have been building AI around the speaker
Much of the development of generative AI has focused on the system’s ability to produce language. Far less attention has been paid to those listening to it.
This matters because generative AI systems have become new entities we communicate with.
If we spend hours every day asking an AI platform questions, drafting emails or talking to voice assistants, we are participating in a new kind of communicative environment, and humans have a knack for adapting to new environments.
The future of AI voices
Many technology companies already try to make synthetic voices sound less mechanistic, introducing disfluencies and conversational markers drawn from human interaction.
As robotics develops and interaction with machines becomes routine, synthetic voices may acquire more convincing rhythm and prosody. They may recognise when a listener needs clarification, when a joke needs a pause, or when a speaker’s uncertainty should show in their delivery.
Alternatively, mechanistic features – my friend’s “breathless” voice, the perfectly structured answer, the absence of hesitation – may come to define the sound of a generative AI model: a linguistic signature that functions as a kind of watermark. We may stop seeing these as flaws and start treating them as a register.
AI-to-AI language
A particularly intriguing development is what happens when artificial systems talk to each other. In September 2026, research found autonomous AI agents developed unusual vocabulary, metaphors and coded expressions in experimental environments. Researchers described the resulting language as difficult for humans to interpret, comparing it to the emergence of innovative language in human communities.
Language has historically been shaped by interaction among humans. Communities develop new ways of saying or signing things, people borrow expressions from one another, misunderstand them, reshape them and pass them on.
Now there is another kind of participant in that environment. Synthetic systems can generate language, consume it and interact with other synthetic systems at a scale and speed humans cannot match.
Should AI sound human – or like AI?
Returning to my friend’s complaint: should AI sound human at all?
If synthetic voices retain identifiable characteristics, listeners know they are dealing with a machine – and knowing that may have value. But making a voice obviously artificial doesn’t solve every problem.
A synthetic voice can sound recognisably synthetic while still reproducing social stereotypes or narrow assumptions about what a particular given kind of speaker should sound like.
The question then is not just whether AI sounds human. It is also what kind of human language the machine is trained to represent, whose language counts as the model, and what happens to everything that falls outside it, including the voices of minoritised groups. The listener needs to be at the centre of that design.
The future of AI – and of language – will depend on how humans respond to the language machines produce, and how the interactions between humans and machines continue to shape communication. Only time will tell.
![]()
Celeste Rodriguez Louro receives funding from Google.
Original source: https://analysis1.mil-osi.com/2026/10/07/what-happens-to-language-when-machines-start-talking-back/
