I paid for a year of ElevenLabs before I worked out I didn't need it.
In my defence, the maths is not obvious until you do it. I write novels, and I wanted them as audiobooks without paying a narrator several hundred pounds per finished hour, or signing away a slice of every future sale to Audible's production arm. AI narration looked like the answer, so I bought the annual subscription to the best-known service and started feeding it a book. Then I noticed the meter. A metered cloud voice bills you per character, forever — every render, every re-render after a typo fix, every "let me just hear that chapter again." One of my novels is about a hundred and sixty-five thousand characters. The plan I'd bought came with around a hundred and thirty thousand characters a month. A single book would not fit inside a single month's allowance. A series was out of the question.
So I built my own studio instead. It runs on hardware I already owned, it costs nothing per book, and in a blind-ish listening test it beat the paid cloud voice I'd just subscribed to. This is how it works, what went wrong along the way, and the one thing it can do that I decided it shouldn't.
The studio in the spare room
It is two computers on the home network, neither of them new. A small always-on server runs the studio software and the pipeline. A second box with a modest graphics card runs the voice engines. Nothing leaves the house: it's a private, internal studio that happens to produce things I can sell.
The narrator is an open, commercially-licensed text-to-speech model. That licence is the entire game, and it is the thing most people miss. The best-sounding open voice-cloning models are licensed for non-commercial use only — wonderful toys, completely useless the moment you want to put a book up for sale. The model I narrate with is plainer and boring in exactly the right way: it is permissively licensed, so its output is mine to sell. Boring and legal beats brilliant and forbidden every time money is involved.
A finished audiobook goes through more steps than you'd think. The prose is cleaned and run through a pronunciation lexicon so the narrator doesn't mangle invented names. It's chunked into synthesizable pieces. A voice is cast — and, critically, approved by ear on this book's own text before a single chapter is committed to. Then it renders, chapter by chapter. The raw audio is normalised in two passes to the exact retail broadcast specification the audiobook stores enforce — for the technically curious, that is −23 LUFS integrated loudness, a true-peak ceiling of −3 dB, 44.1kHz, mono. It's stitched into a single file with chapter markers, an embedded cover, and synthesized opening and closing credits. A full novel renders in about six hours — which is to say it renders faster than I could read it aloud, while I do something else — and the marginal cost of producing it is zero. The only running cost is the electricity, which is no more than leaving a game running for an evening.
The part that genuinely surprised me: on the day I stood it up, I ran the home-made synthetic narrator against the metered cloud voice on the same passage and scored them blind. The local voice took it, eight to seven. The cheap one didn't just save money. It sounded better. The narrator on my novels now — a synthetic voice I'll call Isabella, because that is the stage name she goes by — is one a paying stranger has since bought a book to listen to, and liked.
The bugs that lied to me
If there is a single theme to running software that AI builds and AI narrates, it is this: the thing that tells you it worked is not the thing that did the work, and you must never confuse the two. The studio taught me that three times.
The first was the narrator quietly eating words. The voice-cloning engine — the one I'll come back to — would occasionally skip a clause mid-sentence and carry straight on, perfectly fluent, with no sign anything was missing. You cannot hear an omission; fluent speech with a hole in it just sounds like speech. I only found it because a manual listen through an early render turned up several dropped passages in the first few scenes. The fix is a guardrail I'm rather proud of: after the audio is generated, I run it back through a speech-to-text model, then compare the transcript against the script I sent in, word for word, and measure how much got through. Below a threshold, the chunk is automatically re-recorded. The narrator is now graded against its own script, every single time, by a second model whose only job is to disbelieve the first. A tool that drops words cannot be trusted to tell you it dropped them.
The second is my favourite, because it lied with a green light. One of the voice services slowly leaked graphics memory — creeping up over three days of uptime toward the card's ceiling, until it couldn't grab the scratch space to synthesize a single second of audio. Every render failed. And the whole time, the service's health check reported green, because the health check only ever answered "is the web server running?" — never "can it actually do the work?" Inference was stone dead behind a cheerful green light for who knows how long.
It gets better. My infrastructure engineer — Scotty, who keeps the engines running — shipped a fix for the leak, and fat-fingered a single configuration value, writing a setting's name with no value attached. The result: the model now failed to load onto the graphics card at all. Every render broke again, in a completely different way — and the health check stayed green through that failure too. Two unrelated faults, one indicator lying about both. The fix wasn't only correcting the typo. It was building a guardrail that refuses to believe the green light: before any multi-hour render now begins, the studio fires one throwaway one-word synthesis at the engine — it makes it say "Sound check" — and confirms real audio comes back. If it doesn't, the whole job aborts immediately with a clear message, instead of discovering six hours later that the card was dead the entire time. I don't trust the light. I make it speak a word first.
The third was subtler and is a good lesson in why "close enough" isn't. The stores enforce that true-peak ceiling of −3 dB, so I mastered to exactly −3. One chapter came back measured at −2.89 — eleven hundredths of a decibel over the line, and an automatic rejection. The cause is a thing you only learn by being bitten: the loudness limiter and the lossy encode each nudge the peak up by a hair after you've measured it, so mastering to the ceiling pushes you straight through it. You have to leave headroom. I now target −3.5, which lands comfortably under the limit, and the quality check flags any peak over −3 — which, embarrassingly, it had not been doing before, because it had been measuring loudness and not peak, and would have happily waved the bad chapter through.
Three bugs, one shape. The transcript that wasn't checked, the health endpoint that checked the wrong thing, the quality gate that measured loudness instead of peak. Every one of them was a green checkmark attached to nothing. The studio works now not because the tools became trustworthy, but because I stopped trusting them and started making them prove it.
The capability I switched off
I've kept referring to a voice-cloning engine, and here is where it comes in, because it is the most important part of this whole story and it is not a technical one.
The studio can clone a voice from almost nothing. Feed it a reference clip well under a minute long, and it returns a voice you cannot distinguish from the original — the cadence, the warmth, the little catches in the breath. We tested it. I sat in that spare room and listened to instantly recognisable, world-famous voices reading my prose, perfectly, generated from a sample shorter than this paragraph takes to read aloud. It is thrilling for about ten seconds and deeply unsettling for a good while after. And — this is the part I have to be careful about — in our internal tests, a cloned voice scored higher than the synthetic narrator I actually ship. The best-sounding option was the one I refuse to use.
I refuse to use it, and switched that capability off for anything that leaves the house, for reasons that took me a while to say plainly. A clip under a minute is all it takes — that is not a power you sit comfortably beside. The voice isn't mine to cast: it belongs to a person who never sat in my booth and never said yes. And there is a quiet killer underneath both of those. Once a cloned voice is inside a finished product, a year later nobody remembers which engine made which voice. The provenance simply evaporates. The only moment you can hold the line is before, not after.
So the studio runs two lanes that never cross. Private experiments at home can do as they please; that's nobody's business but mine. I cloned myself, which is harmless and faintly unnerving — being able to make yourself say things you'd never say is a novelty that wears off in about a minute. I also cloned my wife, with her express permission and under terms she renegotiates roughly daily. I can report that there is a specific category of marital peril in keeping a flawless copy of your spouse's voice on a hard drive, that the danger is entirely domestic rather than legal, and that the wisest thing I have ever done with that particular file is absolutely nothing. The point underneath the joke is the serious one: those two clones exist because the people they belong to said yes. That is the whole difference, and it is not a technical setting.
Because anything I sell uses only clean, synthetic, properly-licensed voices, and never the voice of a real living person who didn't sign up for it. The rule I landed on, and the sentence I'd want you to take from all of this, is simple: a voice has to have business being in the room. Isabella earns her place — she's a synthetic voice, she's licensed, she's mine, and a paying stranger bought the book and loved her. A perfect clone of a famous stranger earns nothing but a lawsuit and a guilty conscience.
That's the whole thing, really. The interesting part of building with this technology is almost never whether you can. The tools are astonishing and getting more so by the month; "can you" stopped being the question a while ago. The interesting part — the part that's actually engineering judgement rather than a demo — is what you do in the moment after it works, when the impressive thing and the right thing turn out to be different things. I built a studio that could convincingly clone anyone alive. That is precisely why I switched that part of it off. The technology was never in doubt. The answer was still no.