- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .gitignore | ||
| audiobook.py | ||
| README.md | ||
| requirements.txt | ||
EPUB to OpenAI audiobook
This tool converts an EPUB into cleaned, chapter-aware text; chunks it safely for OpenAI text-to-speech; generates resumable audio; and optionally creates both per-chapter .m4a files and a chaptered .m4b with cover art and metadata.
It was built for The Spirits Book, but it is reusable with EPUB 2 (NCX) and EPUB 3 (navigation-document) books. It follows TOC anchors and spine order together, so it handles chapters that span several XHTML files and several chapters stored inside one XHTML file.
Recommended workflow
Use gpt-4o-mini-tts with the Speech API. OpenAI currently describes it as its newest and most reliable TTS model; it supports delivery instructions for pacing, tone, intonation, and similar narration choices. OpenAI recommends the built-in marin or cedar voices for best quality. The model page states a 2,000-input-token maximum, so this tool defaults to a conservative 1,800 tokens per request.
Official references:
OpenAI requires clear disclosure to listeners that the voice is AI-generated. By default, this tool prepends the spoken notice “This audiobook is narrated by an AI-generated voice” to the first chapter. Customize it with --disclosure "...", or pass an empty string only if the disclosure is clear elsewhere.
1. Install
Python 3.10 or later is recommended. ffmpeg and ffprobe are only needed for chapter/M4B assembly.
cd epub-audiobook-openai
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
On macOS, if FFmpeg is not already installed:
brew install ffmpeg
2. Inspect before spending anything
python audiobook.py inspect "/path/to/book.epub"
This is read-only. Confirm that the selected audiobook chapters look right.
3. Extract, clean, and chunk
python audiobook.py prepare "/path/to/book.epub" --output audiobook-build
Review audiobook-build/plan.json and a few files under audiobook-build/text/. Preparation excludes common publisher-ad, copyright, contents, and index entries by default. Add another exclusion when needed:
python audiobook.py prepare "/path/to/book.epub" \
--output audiobook-build \
--exclude "newsletter|mailing list"
4. Make a short voice sample first
The full book is long, so listen to one or two chunks before committing to all of it. Copy a prepared chunk into a tiny test build, or temporarily keep only one chunk in a copy of plan.json, then run synthesis. The recommended starting voice is marin; cedar is the other quality-oriented choice.
Set the API key in the current terminal session:
export OPENAI_API_KEY="your-api-key"
Generate the prepared chunks:
python audiobook.py synthesize audiobook-build \
--model gpt-4o-mini-tts \
--voice marin \
--format aac \
--jobs 2
The default narration direction is:
Read as a polished, unhurried audiobook narrator. Preserve the distinction between questions, quoted answers, and commentary. Use natural pauses at section headings. Do not add or omit words.
Override it with --instructions "...". Keep the same model, voice, instructions, and format for a consistent book. AAC is the default because it can normally be placed into M4A/M4B without another lossy encode.
Resumability
Every request has a hash in state.json. If a run stops, use the exact same command again. Completed matching chunks are skipped. A changed voice, model, instruction, format, or chunk text changes the hash and regenerates only affected chunks. --force deliberately regenerates everything.
5. Assemble per-chapter files and M4B
python audiobook.py assemble audiobook-build \
--m4b "The Spirits Book - Allan Kardec.m4b"
Outputs:
audiobook-build/chapter_audio/*.m4a— one file per useful chapter- the requested
.m4b— cover art, author/title metadata, and seekable chapters
If chunks were generated as MP3/WAV/FLAC/Opus, assembly transcodes each chapter once to mono AAC at 64 kbps by default. Override with --bitrate 80k. AAC source chunks are stream-copied instead.
One-command run
After inspecting the EPUB and testing the voice:
python audiobook.py all "/path/to/book.epub" \
--output audiobook-build \
--model gpt-4o-mini-tts \
--voice marin \
--format aac \
--jobs 2 \
--m4b "The Spirits Book - Allan Kardec.m4b"
Practical cautions
- Make sure you have the right to create and use an audio adaptation of the edition/translation, especially before sharing it.
- Listen to samples containing dialogue, quotations, numbers, and unusual names before generating the full book.
- Keep the build directory. It contains the cleaned source, plan, request state, and raw chunks needed to resume or rebuild the M4B.
- The script never writes the API key to disk.
- API availability, rate limits, and pricing depend on the account and can change; check the current OpenAI dashboard and official model page before a full run.