Sign inSign up

athomasson2/ebook2audiobookxtts

By athomasson2

•Updated almost 2 years ago

This is a Docker version of this github repo https://github.com/DrewThomasson/ebook2audiobookXTTS

Image
Data science
6

50K+

athomasson2/ebook2audiobookxtts repository overview

⁠🚨 Important Message 🚨

This DockerHub repository is now legacy and will no longer be updated.

The project has moved to a new repository at ebook2audiobook.
Please visit the new repository for the latest updates.

⁠New DockerHub Repository:

https://hub.docker.com/repository/docker/athomasson2/ebook2audiobook/general⁠

⁠New GitHub Repository:

https://github.com/DrewThomasson/ebook2audiobook⁠

Ko-Fi

⁠📚 ebook2audiobook

Convert eBooks to audiobooks with chapters and metadata using Calibre and Coqui XTTS. Supports optional voice cloning and multiple languages!

For more details, check the GitHub repository: ebook2audiobookXTTS⁠.

⁠🌟 Features

  • 📖 Converts eBooks to text format with Calibre.
  • 📚 Splits eBook into chapters for organized audio.
  • 🎙️ High-quality text-to-speech with Coqui XTTS.
  • 🗣️ Optional voice cloning with your own voice file.
  • 🌍 Supports multiple languages (English by default).
  • 🖥️ Designed to run on 4GB RAM.

⁠🐳 Using Docker

You can use Docker to run the eBook to Audiobook converter. This method ensures consistency across different environments and simplifies setup.

⁠🚀 Running the Docker Container

To run the Docker container and start the Gradio interface, use the following command:

To run with a gpu

docker run -it --rm --gpus all -p 7860:7860 --platform=linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py

To run with only cpu

docker run -it --rm -p 7860:7860 --platform=linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py

This command will start the Gradio interface on port 7860.(localhost:7860)

⁠🎥 Demos

https://github.com/user-attachments/assets/8486603c-38b1-43ce-9639-73757dfb1031⁠

⁠🖥️ GUI

demo_web_gui

Click to see images image image image
⁠🛠️ For Custom Xtts Models

Models built to be better at a specific voice. Check out my Hugging Face page here⁠.

To use a custom model, paste the link of the Finished_model_files.zip file like this:

David Attenborough fine tuned Finished_model_files.zip⁠

⁠Example of using docker in headless mode + Full guide

first for a docker pull of the latest with

docker pull athomasson2/ebook2audiobookxtts:huggingface
  • Before you do run this you need to create a dir named "input-folder" in your current dir which will be linked, This is where you can put your input files for the docker image to see
mkdir input-folder && mkdir Audiobooks
  • In the command below swap out YOUR_INPUT_FILE.TXT with the name of your input file
docker run -it --rm \
    -v $(pwd)/input-folder:/home/user/app/input_folder \
    -v $(pwd)/Audiobooks:/home/user/app/Audiobooks \
    --platform linux/amd64 \
    athomasson2/ebook2audiobookxtts:huggingface \
    python app.py --headless True --ebook /home/user/app/input_folder/YOUR_INPUT_FILE.TXT
  • And that should be it!

  • The output Audiobooks will be found in the Audiobook folder which will also be located in your local dir you ran this docker command in

⁠To get the help command for the other parameters this program has you can run this

docker run -it --rm \
    --platform linux/amd64 \
    athomasson2/ebook2audiobookxtts:huggingface \
    python app.py -h

and that will output this

user/app/ebook2audiobookXTTS/input-folder -v $(pwd)/Audiobooks:/home/user/app/ebook2audiobookXTTS/Audiobooks --memory="4g" --network none --platform linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py -h
starting...
usage: app.py [-h] [--share SHARE] [--headless HEADLESS] [--ebook EBOOK] [--voice VOICE]
              [--language LANGUAGE] [--use_custom_model USE_CUSTOM_MODEL]
              [--custom_model CUSTOM_MODEL] [--custom_config CUSTOM_CONFIG]
              [--custom_vocab CUSTOM_VOCAB] [--custom_model_url CUSTOM_MODEL_URL]
              [--temperature TEMPERATURE] [--length_penalty LENGTH_PENALTY]
              [--repetition_penalty REPETITION_PENALTY] [--top_k TOP_K] [--top_p TOP_P]
              [--speed SPEED] [--enable_text_splitting ENABLE_TEXT_SPLITTING]

Convert eBooks to Audiobooks using a Text-to-Speech model. You can either launch the
Gradio interface or run the script in headless mode for direct conversion.

options:
  -h, --help            show this help message and exit
  --share SHARE         Set to True to enable a public shareable Gradio link. Defaults
                        to False.
  --headless HEADLESS   Set to True to run in headless mode without the Gradio
                        interface. Defaults to False.
  --ebook EBOOK         Path to the ebook file for conversion. Required in headless
                        mode.
  --voice VOICE         Path to the target voice file for TTS. Optional, uses a default
                        voice if not provided.
  --language LANGUAGE   Language for the audiobook conversion. Options: en, es, fr, de,
                        it, pt, pl, tr, ru, nl, cs, ar, zh-cn, ja, hu, ko. Defaults to
                        English (en).
  --use_custom_model USE_CUSTOM_MODEL
                        Set to True to use a custom TTS model. Defaults to False. Must
                        be True to use custom models, otherwise you'll get an error.
  --custom_model CUSTOM_MODEL
                        Path to the custom model file (.pth). Required if using a custom
                        model.
  --custom_config CUSTOM_CONFIG
                        Path to the custom config file (config.json). Required if using
                        a custom model.
  --custom_vocab CUSTOM_VOCAB
                        Path to the custom vocab file (vocab.json). Required if using a
                        custom model.
  --custom_model_url CUSTOM_MODEL_URL
                        URL to download the custom model as a zip file. Optional, but
                        will be used if provided. Examples include David Attenborough's
                        model: 'https://huggingface.co/drewThomasson/xtts_David_Attenbor
                        ough_fine_tune/resolve/main/Finished_model_files.zip?download=tr
                        ue'. More XTTS fine-tunes can be found on my Hugging Face at
                        'https://huggingface.co/drewThomasson'.
  --temperature TEMPERATURE
                        Temperature for the model. Defaults to 0.65. Higher Tempatures
                        will lead to more creative outputs IE: more Hallucinations.
                        Lower Tempatures will be more monotone outputs IE: less
                        Hallucinations.
  --length_penalty LENGTH_PENALTY
                        A length penalty applied to the autoregressive decoder. Defaults
                        to 1.0. Not applied to custom models.
  --repetition_penalty REPETITION_PENALTY
                        A penalty that prevents the autoregressive decoder from
                        repeating itself. Defaults to 2.0.
  --top_k TOP_K         Top-k sampling. Lower values mean more likely outputs and
                        increased audio generation speed. Defaults to 50.
  --top_p TOP_P         Top-p sampling. Lower values mean more likely outputs and
                        increased audio generation speed. Defaults to 0.8.
  --speed SPEED         Speed factor for the speech generation. IE: How fast the
                        Narrerator will speak. Defaults to 1.0.
  --enable_text_splitting ENABLE_TEXT_SPLITTING
                        Enable splitting text into sentences. Defaults to True.

Example: python script.py --headless --ebook path_to_ebook --voice path_to_voice
--language en --use_custom_model True --custom_model model.pth --custom_config
config.json --custom_vocab vocab.json
⁠🔨 Building the Docker Image

To build the Docker image, use the following command:

docker build -t athomasson2/ebook2audiobookxtts:latest .
⁠📄 Dockerfile Overview
# Use an official Python 3.10 image
FROM python:3.10-slim-buster

# Set non-interactive installation to avoid timezone and other prompts
ENV DEBIAN_FRONTEND=noninteractive

# Install necessary packages
RUN apt-get update && apt-get install -y --no-install-recommends \
    git \
    calibre \
    espeak \
    espeak-ng \
    ffmpeg \
    wget \
    tk \
    mecab \
    libmecab-dev \
    mecab-ipadic-utf8 \
    build-essential \
    && rm -rf /var/lib/apt/lists/*

# Set the working directory in the container
WORKDIR /ebook2audiobookXTTS

# Clone the ebook2audiobookXTTS repository and install dependencies
RUN git clone https://github.com/DrewThomasson/ebook2audiobookXTTS.git .

# Install Python dependencies
RUN pip install --upgrade pip
RUN pip install bs4 pydub nltk beautifulsoup4 ebooklib tqdm mecab-python3 tts==0.21.3 unidic

# Download unidic
RUN python -m unidic download

# Copy test audio file
COPY 1.wav /ebook2audiobookXTTS/

# Run a test to set up XTTS
RUN echo "import torch" > /tmp/script1.py && \
    echo "from TTS.api import TTS" >> /tmp/script1.py && \
    echo "device = 'cuda' if torch.cuda.is_available() else 'cpu'" >> /tmp/script1.py && \
    echo "print(TTS().list_models())" >> /tmp/script1.py && \
    echo "tts = TTS('tts_models/multilingual/multi-dataset/xtts_v2').to(device)" >> /tmp/script1.py && \
    echo "wav = tts.tts(text='Hello world!', speaker_wav='1.wav', language='en')" >> /tmp/script1.py && \
    echo "tts.tts_to_file(text='Hello world!', speaker_wav='1.wav', language='en', file_path='output.wav')" >> /tmp/script1.py && \
    yes | python3 /tmp/script1.py

# Remove the test audio file
RUN rm -f /ebook2audiobookXTTS/output.wav

# Verify that the script exists and has the correct permissions
RUN ls -la /ebook2audiobookXTTS/

# Check if the script exists and log its presence
RUN if [ -f /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py ]; then echo "Script found."; else echo "Script not found."; exit 1; fi

RUN pip install gradio

# Modify the Python script to set share=True
RUN sed -i 's/demo.launch(share=False)/demo.launch(share=True)/' /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py

# Install nltk
RUN pip install nltk

# Download the punkt package for nltk
RUN python -m nltk.downloader punkt

# Set the command to run your GUI application
CMD ["python", "/ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py"]
⁠📄 NVIDIA GPU Dockerfile Overview
# Use an official NVIDIA CUDA image with cudnn8 and Ubuntu 20.04 as the base
FROM nvidia/cuda:11.8.0-cudnn8-runtime-ubuntu20.04

# Set non-interactive installation to avoid timezone and other prompts
ENV DEBIAN_FRONTEND=noninteractive

# Install necessary packages including Miniconda
RUN apt-get update && apt-get install -y --no-install-recommends \
    wget \
    git \
    espeak \
    espeak-ng \
    ffmpeg \
    tk \
    mecab \
    libmecab-dev \
    mecab-ipadic-utf8 \
    build-essential \
    calibre \
    && rm -rf /var/lib/apt/lists/*

RUN ebook-convert --version

# Install Miniconda
RUN wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -O ~/miniconda.sh && \
    bash ~/miniconda.sh -b -p /opt/conda && \
    rm ~/miniconda.sh



# Set PATH to include conda
ENV PATH=/opt/conda/bin:$PATH

# Create a conda environment with Python 3.10
RUN conda create -n ebookenv python=3.10 -y

# Activate the conda environment
SHELL ["conda", "run", "-n", "ebookenv", "/bin/bash", "-c"]

# Install Python dependencies using conda and pip
RUN conda install -n ebookenv -c conda-forge \
    pydub \
    nltk \
    mecab-python3 \
    && pip install --no-cache-dir \
    bs4 \
    beautifulsoup4 \
    ebooklib \
    tqdm \
    tts==0.21.3 \
    unidic \
    gradio

# Download unidic
RUN python -m unidic download

# Set the working directory in the container
WORKDIR /ebook2audiobookXTTS

# Clone the ebook2audiobookXTTS repository
RUN git clone https://github.com/DrewThomasson/ebook2audiobookXTTS.git .

# Copy test audio file
COPY 1.wav /ebook2audiobookXTTS/

# Run a test to set up XTTS
RUN echo "import torch" > /tmp/script1.py && \
    echo "from TTS.api import TTS" >> /tmp/script1.py && \
    echo "device = 'cuda' if torch.cuda.is_available() else 'cpu'" >> /tmp/script1.py && \
    echo "print(TTS().list_models())" >> /tmp/script1.py && \
    echo "tts = TTS('tts_models/multilingual/multi-dataset/xtts_v2').to(device)" >> /tmp/script1.py && \
    echo "wav = tts.tts(text='Hello world!', speaker_wav='1.wav', language='en')" >> /tmp/script1.py && \
    echo "tts.tts_to_file(text='Hello world!', speaker_wav='1.wav', language='en', file_path='output.wav')" >> /tmp/script1.py && \
    yes | python /tmp/script1.py

# Remove the test audio file
RUN rm -f /ebook2audiobookXTTS/output.wav

# Verify that the script exists and has the correct permissions
RUN ls -la /ebook2audiobookXTTS/

# Check if the script exists and log its presence
RUN if [ -f /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py ]; then echo "Script found."; else echo "Script not found."; exit 1; fi

# Modify the Python script to set share=True
RUN sed -i 's/demo.launch(share=False)/demo.launch(share=True)/' /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py

# Download the punkt package for nltk
RUN python -m nltk.downloader punkt

# Set the command to run your GUI application using the conda environment
CMD ["conda", "run", "--no-capture-output", "-n", "ebookenv", "python", "/ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py"]

⁠🌐 Supported Languages

  • English (en)
  • Spanish (es)
  • French (fr)
  • German (de)
  • Italian (it)
  • Portuguese (pt)
  • Polish (pl)
  • Turkish (tr)
  • Russian (ru)
  • Dutch (nl)
  • Czech (cs)
  • Arabic (ar)
  • Chinese (zh-cn)
  • Japanese (ja)
  • Hungarian (hu)
  • Korean (ko)

⁠📚 Supported eBook Formats

  • .epub, .pdf, .mobi, .txt, .html, .rtf, .chm, .lit, .pdb, .fb2, .odt, .cbr, .cbz, .prc, .lrf, .pml, .snb, .cbc, .rb, .tcr
  • Best results: .epub or .mobi for automatic chapter detection

⁠📂 Output

  • Creates an .m4b file with metadata and chapters.
  • Example Output: Example

⁠🙏 Special Thanks

For more details, check the GitHub repository: ebook2audiobookXTTS⁠.

Tag summary

Content type

Image

Digest

sha256:d3514e84a…

Size

8.3 GB

Last updated

almost 2 years ago

docker pull athomasson2/ebook2audiobookxtts