Feed subtitle files into a Markov chain generator
610
Generate screenshots of videos with fake subtitles, using real subtitles as training data.
ffmpeg built with libass support
ffmpeg -buildconf to see if --enable-libass is presentSubtitle files in .srt or .ass format
.mkv video with subtitles, you can use ffmpeg to extract themSome subtitles files may have fancy typesetting (karaoke, signs, etc) which you might not want as training data. The program attempts to sanitize some of these cases, but I recommend removing problematic lines manually (you can do this in a text editor).
To start, we need to feed subtitle files to a Markov model, which will be used
to generate text in future steps. If we want to save the model to model.yaml:
subkatsu train -o model.yaml subtitle1.srt subtitle2.srt subtitle3.srt
You can also use the -r flag to recursively find subtitles in a directory:
subkatsu train -o model.yaml -r /path/to/subtitles/
By default, it will create a Markov model with order 2.
You can use the --order flag to adjust:
subkatsu train -o model.yaml --order 1 -r /path/to/subtitles/
To check that our model works, we can try generating some text:
subkatsu generate -n 10 model.yaml
This will generate 10 lines to stdout.
Given an input .mkv file that has embedded subtitles, we can generate some
screenshots as follows:
subkatsu screenshots \
--model model.yaml \
--video video.mkv \
--output-dir /path/to/screenshots/ \
-n 10
This will generate some text, using video.mkv as reference (for subtitle
timing/typesetting) and export 10 jpg files to /path/to/screenshots/.
Some additional flags are available:
--min-length 10: Ensures each line has at least 10 characters--subtitles-out /path/to/subs.ass: If you want to save the generated subtitles file--all: Save a screenshot for every subtitle line--resolution 30s: Save at most one screenshot per 30 seconds--format %H%M%S%f_%t: Screenshot filename format. See --help for more info.Content type
Image
Digest
Size
50.2 MB
Last updated
over 7 years ago
docker pull walfie/subkatsu