I Built a Digital Multi-Cast Audiobook. Here's Why I Still Prefer Human Narration.
Human voices, digital voices, creative control, distribution and what I learned producing
Daughter of Vengeance
Audiobooks are no longer a small corner of publishing. In 2025, U.S. audiobook publisher revenue reached $2.43 billion, up 9% from the previous year. According to the Audio Publishers Association’s 2026 research, 58% of American adults have listened to an audiobook, representing an estimated 157 million people. Fiction accounted for 71% of audiobook sales.
As an author and independent publisher, those numbers matter to me. But there’s another number that caught my attention: only 16% of audiobook listeners surveyed have listened to an AI-voiced audiobook. Willingness to try AI narration also dropped from 70% in 2025 to 61% in 2026, while AI-narrated titles accounted for only 0.03% of audiobook sales revenue in 2025.
That’s a massive difference between what technology is capable of doing and what listeners are currently buying. I know both sides of this conversation because I’m not watching the digital-audiobook revolution from the sidelines. I’ve been producing a digital multi-cast edition of Daughter of Vengeance while also opening auditions for a human narrator through ACX. After working with both approaches, I’ve learned something important: digital narration can be impressive, but having complete control over an audiobook doesn’t mean much if people don’t want to listen to it.

Digital Voices Are Better Than Some People Think
I don’t believe digital narration deserves to be dismissed. While producing the multi-cast version of Daughter of Vengeance, I’ve been able to give individual characters their own voices instead of having one voice perform the entire book. That creates character separation and can make conversations feel much more alive.
The technology can also perform emotion. The difference is that I’ve learned you have to direct it. If Amara needs to deliver a line angrily, the digital voice may need that emotional direction. If she’s joking, being sarcastic, hurting or becoming aggressive, that change may need to be communicated to the system.
Punctuation becomes important, too. An exclamation point can affect the energy of a performance. Question marks affect inflection. Pauses, sentence structure and emphasis can change how a line comes out. Pronunciation is another part of the job. With an international story like Daughter of Vengeance, names, locations and cultural terms have to sound right. If the system doesn’t understand a pronunciation, I have to teach it.
That doesn’t make digital narration bad. It makes me the director. And when you have an entire cast of characters, that’s a lot of directing.

The Words Aren’t the Entire Performance
One of the biggest things I’ve noticed is that a digital voice can say the words correctly without necessarily capturing what those words mean in that particular moment. Slang needs to sound like slang. A joke shouldn’t sound like someone reading a joke from a piece of paper. Sarcasm needs sarcasm behind it. Aggression needs an edge, and sadness shouldn’t sound identical to happiness.
That became particularly noticeable with Amara. Her digital voice sounded good, but without additional direction, her emotional delivery could remain too consistent from one mood to another. For this character, that’s important. I don’t only need people to hear Amara. I need them to hear her grow.
I want listeners to recognize humor when she’s teasing someone. I want them to hear vulnerability when something hurts her. When she becomes aggressive, I want that transformation to come through the headphones. A digital performance can get closer to those things when I give it the appropriate instructions, but across a full novel, that can mean directing emotional changes repeatedly for multiple characters.
A talented human actor can read the surrounding scene and make some of those decisions instinctively. The words tell you what the character said. The performance tells you what the character meant.
Different Voices Still Have a Major Advantage
One thing producing a multi-cast audiobook reinforced for me is how much I like character separation. That’s something digital multi-cast does extremely well. Instead of one voice trying to become everybody, I can cast distinct voices based on who those people are.
Ghost Routes is a perfect example of why that matters. Ruckus needs a little street in his delivery. His humor, attitude and rhythm should reflect who he is. Kaito is different. He’s more polished, controlled and professional. Simply changing the pitch isn’t enough. They need different personalities behind those voices.
That’s true whether I’m working with digital voices or human actors. Great audiobook casting isn’t simply about finding voices that sound different. It’s about finding performances that make the characters feel different.
What I Like About Digital Multi-Cast
Tremendous creative control over casting and production.
Different characters can have genuinely distinct voices.
Corrections can be made without scheduling another recording session.
Ambitious multi-character productions become more accessible to independent publishers.
I can direct pronunciation, pacing, emotion and character identity myself.
Production can move considerably faster once the workflow is established.
Where Digital Narration Becomes Challenging
Emotional changes may require explicit direction.
Humor, slang and sarcasm can require additional attention.
Punctuation can affect performance more than expected.
Unfamiliar pronunciations may need to be taught to the system.
A producer may have to direct many emotional shifts individually.
Every regeneration still needs to be listened to and quality-checked.
Listener acceptance and distribution remain smaller than the established human-narrated market.
The Numbers Still Favor Human Narration
AI-narrated audiobooks are growing, but the 2026 consumer numbers show how early we still are. Only 16% of audiobook listeners have listened to an AI-voiced audiobook, according to the APA. While 61% say they’re willing to try one, that’s willingness, not actual consumption. AI-voiced titles represented only 0.03% of 2025 audiobook sales revenue.
That doesn’t mean digital narration has no future. I wouldn’t be investing my own time into it if I believed that. It means publishers have to separate what is technologically possible today from what the market has adopted today. I can have complete control over a digital audiobook, make unlimited creative decisions and obsess over every character, but ultimately, somebody has to press Play.
Why I Still Prefer Human Emotion
My preference remains human narration because a strong narrator isn’t simply reading my words. That narrator is interpreting the scene. A human performer can understand that Amara is angry but doesn’t want someone to know she’s angry. That’s different from simply reading a line angrily.
A narrator can recognize when a joke should be understated rather than exaggerated. They can hear sarcasm in the context. They can understand when silence after a sentence is more powerful than immediately moving to the next one. Most importantly, they can develop alongside a character.
The Amara listeners meet early in Daughter of Vengeance shouldn’t feel emotionally identical to the woman she becomes after everything she experiences. I need to hear that journey. That’s what I’m listening for in these auditions. I’m not simply searching for a beautiful voice. I’m looking for somebody who can act the journey.
There’s Also the Way I Like to Experience Books
Personally, one of my favorite experiences is reading a book on Kindle while simultaneously listening to the audiobook. I want the words in front of me while the performance brings them to life. For me, that’s one of the strongest combinations in modern publishing: reading and listening working together instead of competing with each other.
That’s also part of the reason the Audible ecosystem matters to me. I’ve already experienced what that ecosystem can do. Daughter of Vengeance II became a bestseller during its previous audiobook run on Audible. That experience matters when I’m deciding where I want to invest my time and where I believe my audiobooks have the best opportunity to find listeners.
What ACX Actually Costs an Independent Author
Human narration comes with tradeoffs, too. ACX currently offers three marketplace production arrangements: Pay-for-Production, Royalty Share and Royalty Share Plus. Rights holders who already have completed audio can also use the DIY route. Under Audible’s new royalty model, which applies to new marketplace offers sent after May 26, 2026, exclusive distribution pays a 50% royalty while non-exclusive distribution pays 30%.
Pay-for-Production: I negotiate a per-finished-hour price with the producer and pay the production cost when the audiobook is completed. After that, I can choose exclusive distribution at the new 50% royalty rate or non-exclusive distribution at 30%.
Royalty Share: I don’t pay the narrator’s production fee upfront. Instead, the exclusive 50% royalty is divided equally between the rights holder and producer, 25% each under the new model.
Royalty Share Plus: I pay the producer a reduced per-finished-hour amount and we still divide the royalties, 25% to the rights holder and 25% to the producer.
DIY: If I supply finished audio myself and ACX accepts it, there’s no ACX marketplace producer to pay or split royalties with. Under the new model, the rights holder can receive 50% with exclusive distribution or 30% with non-exclusive distribution.
Royalty Share can be extremely attractive when you’re an independent publisher trying to control upfront costs, but nothing is free. Instead of paying the full production cost upfront, you’re sharing future audiobook revenue and agreeing to exclusivity.
ACX’s current agreement establishes a seven-year initial distribution term, and Royalty Share requires exclusive distribution so the producer’s royalty interest can be protected. That’s a serious business decision, especially when you’re building a catalog rather than publishing one audiobook.
Spotify Gives Me More Freedom, But the Money Feels Less Obvious
Spotify for Authors has some significant advantages for independent publishers. Uploading an audiobook directly is free, distribution is non-exclusive, and the author retains ownership of the audiobook and its rights. Spotify also accepts digital-voice narration, although Spotify currently says digitally narrated audiobooks uploaded through Spotify for Authors are distributed on Spotify only rather than through its wider referral distribution network.
The payment side is where Spotify can feel less intuitive from an independent author’s perspective. Spotify does explain how authors are paid. Audiobook royalties can come from direct à-la-carte purchases and listening through eligible Spotify Premium plans, and Spotify provides royalty reports and payout infrastructure.
So it wouldn’t be accurate to say Spotify doesn’t explain its payment system. What isn’t nearly as intuitive is looking at an individual audiobook listen and immediately understanding what that listen is worth to me. As an independent publisher, that’s an important distinction. I can see the audience and listening activity, but understanding exactly how those listening patterns translate into money isn’t always as immediately obvious as I’d like.
Human vs. Digital: What Do I Actually Choose?
After producing digital multi-cast audio and working with human-narrated audiobooks, I don’t think this needs to become a war between technology and actors. Both approaches have legitimate strengths.
Human Narration Gives Me:
Natural emotional interpretation and acting instincts.
Better contextual understanding of humor, sarcasm, slang and subtext.
A performer capable of developing a character emotionally throughout an entire book.
Access to an established audiobook marketplace and listening culture.
Less need to manually direct every change in mood.
Digital Multi-Cast Gives Me:
Greater hands-on creative control.
The ability to cast many distinctive voices without hiring an entire human ensemble.
Faster corrections and revisions.
Direct control over pronunciation and character identity.
Lower barriers to creating an ambitious multi-character production.
The freedom to experiment with exactly how I believe the story should sound.
That’s why I don’t regret producing the digital version of Daughter of Vengeance. Quite the opposite. Producing it taught me more about what I want from the human version.
I started paying closer attention to emotional transitions. I noticed how important punctuation becomes when words are spoken aloud. I heard pronunciations that needed correction. I discovered places where sarcasm needed more bite and humor needed better timing. Working on the audiobook even helped me catch small things in the manuscript that were easier to hear than see. The digital production became another form of quality control.

So Why Am I Auditioning Human Narrators?
After experiencing both, I still prefer the emotional interpretation of a talented human performer. And publishing is ultimately about reaching people. Technology can give an independent publisher an extraordinary amount of control. I believe digital voices will continue improving, and I believe multi-cast digital production has tremendous potential.
But I’m also running a business. Being in complete control of an audiobook is great. Being in complete control of an audiobook nobody listens to isn’t nearly as valuable. Right now, the audience numbers, distribution infrastructure and my own experience still give human narration an advantage.
That’s why Be Ike-Conic Publishing House has officially opened narrator auditions for Daughter of Vengeance through ACX. The digital experiment isn’t being erased. It helped get us here.
Now I’m looking for the human voice capable of taking Amara somewhere technology hasn’t taken her yet.
The story is already written. Now I want to hear somebody truly perform it.
— Isaac Hill III
Amazon Best-Selling Author
Be Ike-Conic Publishing House






Comments