A simple automated build of the ebook-tools Dockerfile
10K+
This is a collection of bash shell scripts for automated and semi-automated organization and management of large ebook collections. It contains the following tools:
organize-ebooks.sh is used to automatically organize folders with potentially huge amounts of unorganized ebooks. This is done by renaming the files with proper names and moving them to other folders:
.epub, .mobi, .azw, .pdf, .djvu, .chm, .cbr, .cbz, .txt, .lit, .rtf, .doc, .docx, .pdb, .html, .fb2, .lrf, .odt, .prc and potentially others. Even compressed ebooks in arbitrary archive files are supported. For example a .zip, .rar or other archive file that contains the .pdf or .html chapters of an ebook can be organized without a problem..pdf, .djvu and image files when no ISBNs were found in them by the fast and straightforward conversion to .txt. This is very useful for scanned ebooks that only contain images or were badly OCR-ed in the first place.interactive-organizer.sh can be used to interactively and manually organize ebook files quickly. A good use case is the organization of the files that could not be automatically organized by the organize-ebooks.sh script. It can also be used to semi-automatically verify the organized files by the above script and potentially reorganize some of them:
organize-ebooks.sh was called with --keep-metadata, the interactive organizer compares the old filename with the new one and shows suspicious differences between the two. Wrongly renamed files can be interactively renamed with this script..txt and shown with less directly in the current terminal or they can be opened with an external viewer without exiting from the interactive organization.find-isbns.sh tries to find valid ISBNs inside a file or in stdin if no file was specified. Searching for ISBNs in files uses progressively more resource-intensive methods until some ISBNs are found, see the documentation below for more details.
convert-to-txt.sh converts the supplied file to a text file. It can optionally also use OCR for .pdf, .djvu and image files.
rename-calibre-library.sh traverses a calibre library folder and renames all the book files in it by reading their metadata from calibre's metadata.opf files.
split-into-folders.sh splits the supplied ebook files (and the accompanying metadata files if present) into folders with consecutive names that each contain the specified number of files.
All of the tools use a library file lib.sh that has useful functions for building other ebook management scripts. More details for the different script options and parameters can be found in the Usage, options and configuration section.
There are two ways you can install and use the tools in this repository - directly or via docker images.
Since all of the tools are shell scripts, you should be able to use them directly from source in most up-to-date GNU/Linux distributions, as long as you have the needed dependencies installed. They should also be usable on other *nix systems like OS X and *BSD if you have the GNU versions of the dependencies installed or in the Windows Subsystem for Linux.
However, since non-linux systems are officially unsupported and may have unexpected issues, Docker containers are the preferred way to use the scripts in those systems. The docker images may also be easier to use than the bare scripts on non-GNU linux distributions or on older linux distributions like some LTS releases.
To install and use the bare shell scripts, follow these steps:
PATH environment variable.You need recent versions of:
file, less, bash 4.3+ and GNU coreutils, awk, sed and grep.ISBN_METADATA_FETCH_ORDER and ORGANIZE_WITHOUT_ISBN_SOURCES to empty strings..pdf, .doc and .djvu files respectively to .txt.The scripts are only tested on linux, though they should work on any *nix system that has the needed dependencies. You can install everything needed with this command in Arch Linux:
pacman -S file less bash coreutils gawk sed grep calibre p7zip tesseract tesseract-data-eng python2-lxml poppler catdoc djvulibre
Note: you can probably get much better OCR results by using the unstable 4.0 version of Tesseract. It is present in the AUR or you can easily make a package like this yourself.
Here is how to install the packages on Debian (and Debian-based distributions like Ubuntu):
apt-get install file less bash coreutils gawk sed grep calibre p7zip-full tesseract-ocr tesseract-ocr-osd tesseract-ocr-eng python-lxml poppler-utils catdoc djvulibre-bin
Keep in mind that a lot of debian-based distributions do not have up-to-date packages and the scripts work best when calibre's version is at least 2.84. For earlier versions you have to set ISBN_METADATA_FETCH_ORDER and ORGANIZE_WITHOUT_ISBN_SOURCES to empty strings.
The docker image includes all of the needed dependencies, even the extra calibre plugins. There is an automatically built docker image in the Docker Hub. You can pull it locally with docker pull ebooktools/scripts. You can also easily build the docker image yourself: simply clone this repository (or download the latest release archive and extract it) and then run docker build -t ebooktools/scripts:latest . in the folder.
Here are some Docker-specific usage details:
docker run -it -v /some/host/folder:/unorganized-books ebooktools/scripts:latest. This will run a bash prompt that has all of the dependencies installed and all of the scripts already in the PATH so all the usage instructions bellow should apply. The contents of the host folder /some/host/folder (the path to the folder on your machine that you want to organize) will be mounted as the /unorganized-books folder in the container.-v option of docker run multiple times to mount several host folders in the container.--rm option of docker run to clean up your containers after you are done with them.--user option of docker run or by editing the Dockerfile and rebuilding it yourself.docker run -it [other-docker-run-options] ebooktools/scripts:latest organize-ebooks.sh [ebook-tools-script-options]For more Docker details, read the docker documentation and docker run reference specifically.
Scripts that work with multiple files recursively scan the supplied folders and assume that one file is one ebook. Ebooks that consist of multiple files should be compressed in a single file archive. The archive type does not matter, it can be .zip, .rar, .tar, .7z and others - all supported archive types by 7zip are fine.
All of the options documented below can either be passed to the scripts via command-line parameters or via environment variables. Command-line parameters supersede environment variables. Most parameters are not required and if nothing is specified, the default value will be used.
All of these options are part of the common library and may affect some or all of the scripts.
-v, --verbose; env. variable VERBOSE; default value false
Whether debug messages will be displayed on stderr. Passing the parameter or changing VERBOSE to true and piping the stderr to a file is useful for debugging or keeping a record of exactly what happens without cluttering and overwhelming the normal execution output.
-d, --dry-run; env. variable DRY_RUN; default value false
If this is enabled, no file rename/move/symlink/etc. operations will actually be executed.
-sl, --symlink-only; env. variable SYMLINK_ONLY; default value false
Instead of moving the ebook files, create symbolic links to them.
-km, --keep-metadata; env. variable KEEP_METADATA; default value false
Do not delete the gathered metadata for the organized ebooks, instead save it in an accompanying file together with each renamed book. It is very useful for semi-automatic verification of the organized files with interactive-organizer.sh or for additional verification, indexing or processing at a later date.
-i=<value>, --isbn-regex=<value>; env. variable ISBN_REGEX; see default value in lib.sh
This is the regular expression used to match ISBN-like numbers in the supplied books. It is matched with grep -P, so look-ahead and look-behind can be used. Also it is purposefully a bit loose (i.e. it can match some non-ISBN numbers), since the found numbers will be checked for validity. Due to unicode handling, the default value is too long for the README, you can find it in lib.sh.
--isbn-blacklist-regex=<value>; env. variable ISBN_BLACKLIST_REGEX; default value ^(0123456789|([0-9xX])\2{9})$
Any ISBNs that were matched by the ISBN_REGEX above and pass the ISBN validation algorithm are normalized and passed through this regular expression. Any ISBNs that successfully match against it are discarded. The idea is to ignore technically valid but probably wrong numbers like 0123456789, 0000000000, 1111111111, etc.
--isbn-direct-grep-files=<value>; env. variable ISBN_DIRECT_GREP_FILES; default value ^text/(plain|xml|html)$
This is a regular expression that is matched against the MIME type of the searched files. Matching files are searched directly for ISBNs, without converting or OCR-ing them to .txt first.
--isbn-ignored-files=<value>; env. variable ISBN_IGNORED_FILES; see default value in lib.sh
This is a regular expression that is matched against the MIME type of the searched files. Matching files are not searched for ISBNs beyond their filename. The default value is a bit long because it tries to make the scripts ignore .gif and .svg images, audio, video and executable files and fonts, you can find it in lib.sh.
--reorder-files-for-grep=<value>; env. variables ISBN_GREP_REORDER_FILES, ISBN_GREP_RF_SCAN_FIRST, ISBN_GREP_RF_REVERSE_LAST; default values true, 400, 50
These options specify if and how we should reorder the ebook text before searching for ISBNs in it. By default, the first 400 lines of the text are searched as they are, then the last 50 are searched in reverse and finally the remainder in the middle. This reordering is done to improve the odds that the first found ISBNs in a book text actually belong to that book (ex. from the copyright section or the back cover), instead of being random ISBNs mentioned in the middle of the book. No part of the text is searched twice, even if these regions overlap. If you use the command-line option, the format for <value> is false to disable the functionality or first_lines,last_lines to enable it with the specified values.
-mfo=<value>, --metadata-fetch-order=<value>; env. variables ISBN_METADATA_FETCH_ORDER; default value Goodreads,Amazon.com,Google,ISBNDB,WorldCat xISBN,OZON.ru
This option allows you to specify the online metadata sources and order in which the scripts will try searching in them for books by their ISBN. The actual search is done by calibre's fetch-ebook-metadata command-line application, so any custom calibre metadata plugins can also be used. To see the currently available options, run fetch-ebook-metadata --help and check the description for the --allowed-plugin option.
If you use Calibre versions that are older than 2.84, it's required to manually set this option to an empty string.
-ocr=<value>, --ocr-enabled=<value>; env. variable OCR_ENABLED; default value false
Whether to enable OCR for .pdf, .djvu and image files. It is disabled by default and can be used differently in two scripts:
organize-ebooks.sh can use OCR for finding ISBNs in scanned books. Setting the value to true will cause it to use OCR for books that failed to be converted to .txt or were converted to empty files by the simple conversion tools (ebook-convert, pdftotext, djvutxt). Setting the value to always will cause it to use OCR even when the simple tools produced a non-empty result, if there were no ISBNs in it.convert-to-txt.sh can use OCR for the conversion to .txt. Setting the value to true will cause it to use OCR for books that failed to be converted to .txt or were converted to empty files by the simple conversion tools. Setting it to always will cause it to first try OCR-ing the books before trying the simple conversion tools.-ocrop=<value>, --ocr-only-first-last-pages=<value>; env. variable OCR_ONLY_FIRST_LAST_PAGES; default value 7,3 (except for convert-to-txt.sh where it's false)
Value n,m instructs the scripts to convert only the first n and last m pages when OCR-ing ebooks. This is done because OCR is a slow resource-intensive process and ISBN numbers are usually at the beginning or at the end of books. Setting the value to false disables this optimization and is the default for convert-to-txt.sh, where we probably want the whole book to be converted.
-ocrc=<value>, --ocr-command=<value>; env. variable OCR_COMMAND; default value tesseract_wrapper
This allows us to define a hook for using custom OCR settings or software. The default value is just a wrapper that allows us to use both tesseract 3 and 4 with some predefined settings. You can use a custom bash function or shell script - the first argument is the input image (books are OCR-ed page by page) and the second argument is the file you have to write the output text to.
--token-min-length=<value>; env. variable TOKEN_MIN_LENGTH; default value 3
When files and file metadata are parsed, they are split into words (or more precisely, either alpha or numeric tokens) and ones shorter than this value are ignored. By default, single and two character number and words are ignored.
--tokens-to-ignore=<value>; env. variable TOKENS_TO_IGNORE; complex default value
A regular expression that is matched against the filename/author/title tokens and matching tokens are ignored. The default regular expression includes common words that probably hinder online metadata searching like book, novel, series, volume and others, as well as probable publication years like (so 1999 is ignored while 2033 is not). You can see it in lib.sh.
-owis=<value>, --organize-without-isbn-sources=<value>; env. variable ORGANIZE_WITHOUT_ISBN_SOURCES; default value Goodreads,Amazon.com,Google
This option allows you to specify the online metadata sources in which the scripts will try searching for books by non-ISBN metadata (i.e. author and title). The actual search is done by calibre's fetch-ebook-metadata command-line application, so any custom calibre metadata plugins can also be used. To see the currently available options, run fetch-ebook-metadata --help and check the description for the --allowed-plugin option. Because Calibre versions older than 2.84 don't support the --allowed-plugin option, if you want to use such an old Calibre version you should manually set ORGANIZE_WITHOUT_ISBN_SOURCES to an empty string.
In contrast to searching by ISBNs, searching by author and title is done concurrently in all of the allowed online metadata sources. The number of sources is smaller because some metadata sources can be searched only by ISBN or return many false-positives when searching by title and author.
-oft=<value>, --output-filename-template=<value>; env. variable OUTPUT_FILENAME_TEMPLATE; default value:
"${d[AUTHORS]// & /, } - ${d[SERIES]:+[${d[SERIES]}] - }${d[TITLE]/:/ -}${d[PUBLISHED]:+ (${d[PUBLISHED]%%-*})}${d[ISBN]:+ [${d[ISBN]}]}.${d[EXT]}"
This specifies how the filenames of the organized files will look. It is a bash string that is evaluated so it can be very flexible (and also potentially unsafe). The book metadata is present in a hashmap with name d and uppercase keys. When changing this parameter, keep in mind that you have to either escape the $ symbols or wrap everything in single quotes like so:
-oft='"${d[TITLE]} by ${d[AUTHORS]}.${d[EXT]}"'
By default the organized files start with the comma-separated author name(s), followed by the book series name and number in square brackets (if present), followed by the book title, the year of publication (if present), the ISBN(s) (if present) and the original extension. Here are are how output filenames using the default template look:
Cory Doctorow - [Little Brother #1] - Little Brother (2008) [0765319853].pdf
Cory Doctorow - [Little Brother #2] - Homeland (2013) [9780765333698].epub
Eliezer Yudkowsky - Harry Potter and the Methods of Rationality (2015).epub
Lawrence Lessig - Remix - Making Art and Commerce Thrive in the Hybrid Economy (2008) [9781594201721].djvu
Rick Falkvinge - Swarmwise (2013) [1463533152].pdf
-ome=<value>, --output-metadata-extension=<value>; env. variable OUTPUT_METADATA_EXTENSION; default value meta
If KEEP_METADATA is enabled, this is the extension of the additional metadata file that is saved next to each newly renamed file.
-fsf=<value>, --file-sort-flags=<value>; env. variable FILE_SORT_FLAGS; default value () (an empty bash array)
A list with the sort options that will be used every time multiple files are processed (i.e. in every script except convert-to-txt.sh).
--debug-prefix-length=<value>; env. variable DEBUG_PREFIX_LENGTH; default value 40
The length of the debug prefix used by the some scripts in the output when VERBOSE mode is enabled.
--lib-hook=<path-fo-bash-file>; not configurable with environment variables for security reasons
Warning: for advanced users with bash knowledge only! This option allows you to inject (i.e. source) bash files into all of the scripts in this collection. You can use this as a way to customize a script's internals, add functionality, make "plugins", etc. This isn't configurable from an environment variable for obvious security reasons, it explicitly has to be set via the CLI flag.
organize-ebooks.sh [<OPTIONS>] folder-to-organize [...]This is probably the most versatile script in the repository. It can automatically organize folders with huge quantities of unorganized ebook files. This is done by extracting ISBNs and/or metadata from the ebook files, downloading their full and hopefully correct metadata from online sources and auto-renaming the unorganized files with full and correct names and moving them to specified folders. Is supports virtually all ebook types, including ebooks in arbitrary or even nested archives (like the other scripts, it assumes that one file is one ebook, even if it's a huge archive). OCR can be used for scanned ebooks and corrupt ebooks and non-ebook documents (pamphlets) can be separated in specified folders. Most of the general options and flags above affect how this script operates, but there are also some specific options for it.
-cco, --corruption-check-only; env. variable CORRUPTION_CHECK_ONLY; default value false
Do not organize or rename files, just check them for corruption (ex. zero-filled files, corrupt archives or broken .pdf files). Useful with the OUTPUT_FOLDER_CORRUPT option.
--tested-archive-extensions=<value>; env. variable TESTED_ARCHIVE_EXTENSIONS; default value ^(7z|bz2|chm|arj|cab|gz|tgz|gzip|zip|rar|xz|tar|epub|docx|odt|ods|cbr|cbz|maff|iso)$
A regular expression that specifies which file extensions will be tested with 7z t for corruption.
-owi, --organize-without-isbn; env. variable ORGANIZE_WITHOUT_ISBN; default value false
Specify whether the script will try to organize ebooks if there were no ISBN found in the book or if no metadata was found online with the retrieved ISBNs. If enabled, the script will first try to use calibre's ebook-meta command-line tool to extract the author and title metadata from the ebook file. The s
Content type
Image
Digest
Size
212.2 MB
Last updated
about 8 years ago
docker pull ebooktools/scripts