This is a Docker version of this github repo https://github.com/DrewThomasson/ebook2audiobookXTTS
50K+
This DockerHub repository is now legacy and will no longer be updated.
The project has moved to a new repository at ebook2audiobook.
Please visit the new repository for the latest updates.
https://hub.docker.com/repository/docker/athomasson2/ebook2audiobook/general
https://github.com/DrewThomasson/ebook2audiobook
Convert eBooks to audiobooks with chapters and metadata using Calibre and Coqui XTTS. Supports optional voice cloning and multiple languages!
For more details, check the GitHub repository: ebook2audiobookXTTS.
You can use Docker to run the eBook to Audiobook converter. This method ensures consistency across different environments and simplifies setup.
To run the Docker container and start the Gradio interface, use the following command:
To run with a gpu
docker run -it --rm --gpus all -p 7860:7860 --platform=linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py
To run with only cpu
docker run -it --rm -p 7860:7860 --platform=linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py
This command will start the Gradio interface on port 7860.(localhost:7860)
https://github.com/user-attachments/assets/8486603c-38b1-43ce-9639-73757dfb1031
Models built to be better at a specific voice. Check out my Hugging Face page here.
To use a custom model, paste the link of the Finished_model_files.zip file like this:
David Attenborough fine tuned Finished_model_files.zip
first for a docker pull of the latest with
docker pull athomasson2/ebook2audiobookxtts:huggingface
mkdir input-folder && mkdir Audiobooks
docker run -it --rm \
-v $(pwd)/input-folder:/home/user/app/input_folder \
-v $(pwd)/Audiobooks:/home/user/app/Audiobooks \
--platform linux/amd64 \
athomasson2/ebook2audiobookxtts:huggingface \
python app.py --headless True --ebook /home/user/app/input_folder/YOUR_INPUT_FILE.TXT
And that should be it!
The output Audiobooks will be found in the Audiobook folder which will also be located in your local dir you ran this docker command in
docker run -it --rm \
--platform linux/amd64 \
athomasson2/ebook2audiobookxtts:huggingface \
python app.py -h
and that will output this
user/app/ebook2audiobookXTTS/input-folder -v $(pwd)/Audiobooks:/home/user/app/ebook2audiobookXTTS/Audiobooks --memory="4g" --network none --platform linux/amd64 athomasson2/ebook2audiobookxtts:huggingface python app.py -h
starting...
usage: app.py [-h] [--share SHARE] [--headless HEADLESS] [--ebook EBOOK] [--voice VOICE]
[--language LANGUAGE] [--use_custom_model USE_CUSTOM_MODEL]
[--custom_model CUSTOM_MODEL] [--custom_config CUSTOM_CONFIG]
[--custom_vocab CUSTOM_VOCAB] [--custom_model_url CUSTOM_MODEL_URL]
[--temperature TEMPERATURE] [--length_penalty LENGTH_PENALTY]
[--repetition_penalty REPETITION_PENALTY] [--top_k TOP_K] [--top_p TOP_P]
[--speed SPEED] [--enable_text_splitting ENABLE_TEXT_SPLITTING]
Convert eBooks to Audiobooks using a Text-to-Speech model. You can either launch the
Gradio interface or run the script in headless mode for direct conversion.
options:
-h, --help show this help message and exit
--share SHARE Set to True to enable a public shareable Gradio link. Defaults
to False.
--headless HEADLESS Set to True to run in headless mode without the Gradio
interface. Defaults to False.
--ebook EBOOK Path to the ebook file for conversion. Required in headless
mode.
--voice VOICE Path to the target voice file for TTS. Optional, uses a default
voice if not provided.
--language LANGUAGE Language for the audiobook conversion. Options: en, es, fr, de,
it, pt, pl, tr, ru, nl, cs, ar, zh-cn, ja, hu, ko. Defaults to
English (en).
--use_custom_model USE_CUSTOM_MODEL
Set to True to use a custom TTS model. Defaults to False. Must
be True to use custom models, otherwise you'll get an error.
--custom_model CUSTOM_MODEL
Path to the custom model file (.pth). Required if using a custom
model.
--custom_config CUSTOM_CONFIG
Path to the custom config file (config.json). Required if using
a custom model.
--custom_vocab CUSTOM_VOCAB
Path to the custom vocab file (vocab.json). Required if using a
custom model.
--custom_model_url CUSTOM_MODEL_URL
URL to download the custom model as a zip file. Optional, but
will be used if provided. Examples include David Attenborough's
model: 'https://huggingface.co/drewThomasson/xtts_David_Attenbor
ough_fine_tune/resolve/main/Finished_model_files.zip?download=tr
ue'. More XTTS fine-tunes can be found on my Hugging Face at
'https://huggingface.co/drewThomasson'.
--temperature TEMPERATURE
Temperature for the model. Defaults to 0.65. Higher Tempatures
will lead to more creative outputs IE: more Hallucinations.
Lower Tempatures will be more monotone outputs IE: less
Hallucinations.
--length_penalty LENGTH_PENALTY
A length penalty applied to the autoregressive decoder. Defaults
to 1.0. Not applied to custom models.
--repetition_penalty REPETITION_PENALTY
A penalty that prevents the autoregressive decoder from
repeating itself. Defaults to 2.0.
--top_k TOP_K Top-k sampling. Lower values mean more likely outputs and
increased audio generation speed. Defaults to 50.
--top_p TOP_P Top-p sampling. Lower values mean more likely outputs and
increased audio generation speed. Defaults to 0.8.
--speed SPEED Speed factor for the speech generation. IE: How fast the
Narrerator will speak. Defaults to 1.0.
--enable_text_splitting ENABLE_TEXT_SPLITTING
Enable splitting text into sentences. Defaults to True.
Example: python script.py --headless --ebook path_to_ebook --voice path_to_voice
--language en --use_custom_model True --custom_model model.pth --custom_config
config.json --custom_vocab vocab.json
To build the Docker image, use the following command:
docker build -t athomasson2/ebook2audiobookxtts:latest .
# Use an official Python 3.10 image
FROM python:3.10-slim-buster
# Set non-interactive installation to avoid timezone and other prompts
ENV DEBIAN_FRONTEND=noninteractive
# Install necessary packages
RUN apt-get update && apt-get install -y --no-install-recommends \
git \
calibre \
espeak \
espeak-ng \
ffmpeg \
wget \
tk \
mecab \
libmecab-dev \
mecab-ipadic-utf8 \
build-essential \
&& rm -rf /var/lib/apt/lists/*
# Set the working directory in the container
WORKDIR /ebook2audiobookXTTS
# Clone the ebook2audiobookXTTS repository and install dependencies
RUN git clone https://github.com/DrewThomasson/ebook2audiobookXTTS.git .
# Install Python dependencies
RUN pip install --upgrade pip
RUN pip install bs4 pydub nltk beautifulsoup4 ebooklib tqdm mecab-python3 tts==0.21.3 unidic
# Download unidic
RUN python -m unidic download
# Copy test audio file
COPY 1.wav /ebook2audiobookXTTS/
# Run a test to set up XTTS
RUN echo "import torch" > /tmp/script1.py && \
echo "from TTS.api import TTS" >> /tmp/script1.py && \
echo "device = 'cuda' if torch.cuda.is_available() else 'cpu'" >> /tmp/script1.py && \
echo "print(TTS().list_models())" >> /tmp/script1.py && \
echo "tts = TTS('tts_models/multilingual/multi-dataset/xtts_v2').to(device)" >> /tmp/script1.py && \
echo "wav = tts.tts(text='Hello world!', speaker_wav='1.wav', language='en')" >> /tmp/script1.py && \
echo "tts.tts_to_file(text='Hello world!', speaker_wav='1.wav', language='en', file_path='output.wav')" >> /tmp/script1.py && \
yes | python3 /tmp/script1.py
# Remove the test audio file
RUN rm -f /ebook2audiobookXTTS/output.wav
# Verify that the script exists and has the correct permissions
RUN ls -la /ebook2audiobookXTTS/
# Check if the script exists and log its presence
RUN if [ -f /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py ]; then echo "Script found."; else echo "Script not found."; exit 1; fi
RUN pip install gradio
# Modify the Python script to set share=True
RUN sed -i 's/demo.launch(share=False)/demo.launch(share=True)/' /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py
# Install nltk
RUN pip install nltk
# Download the punkt package for nltk
RUN python -m nltk.downloader punkt
# Set the command to run your GUI application
CMD ["python", "/ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_gradio.py"]
# Use an official NVIDIA CUDA image with cudnn8 and Ubuntu 20.04 as the base
FROM nvidia/cuda:11.8.0-cudnn8-runtime-ubuntu20.04
# Set non-interactive installation to avoid timezone and other prompts
ENV DEBIAN_FRONTEND=noninteractive
# Install necessary packages including Miniconda
RUN apt-get update && apt-get install -y --no-install-recommends \
wget \
git \
espeak \
espeak-ng \
ffmpeg \
tk \
mecab \
libmecab-dev \
mecab-ipadic-utf8 \
build-essential \
calibre \
&& rm -rf /var/lib/apt/lists/*
RUN ebook-convert --version
# Install Miniconda
RUN wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -O ~/miniconda.sh && \
bash ~/miniconda.sh -b -p /opt/conda && \
rm ~/miniconda.sh
# Set PATH to include conda
ENV PATH=/opt/conda/bin:$PATH
# Create a conda environment with Python 3.10
RUN conda create -n ebookenv python=3.10 -y
# Activate the conda environment
SHELL ["conda", "run", "-n", "ebookenv", "/bin/bash", "-c"]
# Install Python dependencies using conda and pip
RUN conda install -n ebookenv -c conda-forge \
pydub \
nltk \
mecab-python3 \
&& pip install --no-cache-dir \
bs4 \
beautifulsoup4 \
ebooklib \
tqdm \
tts==0.21.3 \
unidic \
gradio
# Download unidic
RUN python -m unidic download
# Set the working directory in the container
WORKDIR /ebook2audiobookXTTS
# Clone the ebook2audiobookXTTS repository
RUN git clone https://github.com/DrewThomasson/ebook2audiobookXTTS.git .
# Copy test audio file
COPY 1.wav /ebook2audiobookXTTS/
# Run a test to set up XTTS
RUN echo "import torch" > /tmp/script1.py && \
echo "from TTS.api import TTS" >> /tmp/script1.py && \
echo "device = 'cuda' if torch.cuda.is_available() else 'cpu'" >> /tmp/script1.py && \
echo "print(TTS().list_models())" >> /tmp/script1.py && \
echo "tts = TTS('tts_models/multilingual/multi-dataset/xtts_v2').to(device)" >> /tmp/script1.py && \
echo "wav = tts.tts(text='Hello world!', speaker_wav='1.wav', language='en')" >> /tmp/script1.py && \
echo "tts.tts_to_file(text='Hello world!', speaker_wav='1.wav', language='en', file_path='output.wav')" >> /tmp/script1.py && \
yes | python /tmp/script1.py
# Remove the test audio file
RUN rm -f /ebook2audiobookXTTS/output.wav
# Verify that the script exists and has the correct permissions
RUN ls -la /ebook2audiobookXTTS/
# Check if the script exists and log its presence
RUN if [ -f /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py ]; then echo "Script found."; else echo "Script not found."; exit 1; fi
# Modify the Python script to set share=True
RUN sed -i 's/demo.launch(share=False)/demo.launch(share=True)/' /ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py
# Download the punkt package for nltk
RUN python -m nltk.downloader punkt
# Set the command to run your GUI application using the conda environment
CMD ["conda", "run", "--no-capture-output", "-n", "ebookenv", "python", "/ebook2audiobookXTTS/custom_model_ebook2audiobookXTTS_with_link_gradio.py"]

For more details, check the GitHub repository: ebook2audiobookXTTS.
Content type
Image
Digest
sha256:d3514e84a…
Size
8.3 GB
Last updated
almost 2 years ago
docker pull athomasson2/ebook2audiobookxtts