No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-21 17:37:44 +02:00
.gitignore Initial commit 2026-09-21 17:37:44 +02:00
audiobook.py Initial commit 2026-09-21 17:37:44 +02:00
README.md Initial commit 2026-09-21 17:37:44 +02:00
requirements.txt Initial commit 2026-09-21 17:37:44 +02:00

EPUB to OpenAI audiobook

This tool converts an EPUB into cleaned, chapter-aware text; chunks it safely for OpenAI text-to-speech; generates resumable audio; and optionally creates both per-chapter .m4a files and a chaptered .m4b with cover art and metadata.

It was built for The Spirits Book, but it is reusable with EPUB 2 (NCX) and EPUB 3 (navigation-document) books. It follows TOC anchors and spine order together, so it handles chapters that span several XHTML files and several chapters stored inside one XHTML file.

Use gpt-4o-mini-tts with the Speech API. OpenAI currently describes it as its newest and most reliable TTS model; it supports delivery instructions for pacing, tone, intonation, and similar narration choices. OpenAI recommends the built-in marin or cedar voices for best quality. The model page states a 2,000-input-token maximum, so this tool defaults to a conservative 1,800 tokens per request.

Official references:

OpenAI requires clear disclosure to listeners that the voice is AI-generated. By default, this tool prepends the spoken notice “This audiobook is narrated by an AI-generated voice” to the first chapter. Customize it with --disclosure "...", or pass an empty string only if the disclosure is clear elsewhere.

1. Install

Python 3.10 or later is recommended. ffmpeg and ffprobe are only needed for chapter/M4B assembly.

cd epub-audiobook-openai
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

On macOS, if FFmpeg is not already installed:

brew install ffmpeg

2. Inspect before spending anything

python audiobook.py inspect "/path/to/book.epub"

This is read-only. Confirm that the selected audiobook chapters look right.

3. Extract, clean, and chunk

python audiobook.py prepare "/path/to/book.epub" --output audiobook-build

Review audiobook-build/plan.json and a few files under audiobook-build/text/. Preparation excludes common publisher-ad, copyright, contents, and index entries by default. Add another exclusion when needed:

python audiobook.py prepare "/path/to/book.epub" \
  --output audiobook-build \
  --exclude "newsletter|mailing list"

4. Make a short voice sample first

The full book is long, so listen to one or two chunks before committing to all of it. Copy a prepared chunk into a tiny test build, or temporarily keep only one chunk in a copy of plan.json, then run synthesis. The recommended starting voice is marin; cedar is the other quality-oriented choice.

Set the API key in the current terminal session:

export OPENAI_API_KEY="your-api-key"

Generate the prepared chunks:

python audiobook.py synthesize audiobook-build \
  --model gpt-4o-mini-tts \
  --voice marin \
  --format aac \
  --jobs 2

The default narration direction is:

Read as a polished, unhurried audiobook narrator. Preserve the distinction between questions, quoted answers, and commentary. Use natural pauses at section headings. Do not add or omit words.

Override it with --instructions "...". Keep the same model, voice, instructions, and format for a consistent book. AAC is the default because it can normally be placed into M4A/M4B without another lossy encode.

Resumability

Every request has a hash in state.json. If a run stops, use the exact same command again. Completed matching chunks are skipped. A changed voice, model, instruction, format, or chunk text changes the hash and regenerates only affected chunks. --force deliberately regenerates everything.

5. Assemble per-chapter files and M4B

python audiobook.py assemble audiobook-build \
  --m4b "The Spirits Book - Allan Kardec.m4b"

Outputs:

  • audiobook-build/chapter_audio/*.m4a — one file per useful chapter
  • the requested .m4b — cover art, author/title metadata, and seekable chapters

If chunks were generated as MP3/WAV/FLAC/Opus, assembly transcodes each chapter once to mono AAC at 64 kbps by default. Override with --bitrate 80k. AAC source chunks are stream-copied instead.

One-command run

After inspecting the EPUB and testing the voice:

python audiobook.py all "/path/to/book.epub" \
  --output audiobook-build \
  --model gpt-4o-mini-tts \
  --voice marin \
  --format aac \
  --jobs 2 \
  --m4b "The Spirits Book - Allan Kardec.m4b"

Practical cautions

  • Make sure you have the right to create and use an audio adaptation of the edition/translation, especially before sharing it.
  • Listen to samples containing dialogue, quotations, numbers, and unusual names before generating the full book.
  • Keep the build directory. It contains the cleaned source, plan, request state, and raw chunks needed to resume or rebuild the M4B.
  • The script never writes the API key to disk.
  • API availability, rate limits, and pricing depend on the account and can change; check the current OpenAI dashboard and official model page before a full run.