Sign inSign up

dadoonet/fscrawler

By dadoonet

•Updated 3 days ago

Docker images for FSCrawler project

Image
Content management system
Databases & storage
5

100K+

dadoonet/fscrawler repository overview

Welcome to the FS Crawler for Elasticsearch⁠.

This crawler helps to index binary documents such as PDF, Open Office, MS Office.

⁠Main features

  • Local file system (or a mounted drive) crawling and index new files, update existing ones and removes old ones.
  • Remote file system over SSH or FTP crawling.
  • REST interface to let you “upload” your binary documents to elasticsearch.

⁠Installation guide

The default image contains Tesseract⁠ and all the trained language data⁠ which adds more than 500mb of data.

docker pull dadoonet/fscrawler

If you don't want to use OCR at all, you can use a smaller image by using instead the noocr images.

docker pull dadoonet/fscrawler:noocr

Then run:

docker run -it --rm \
     -v ~/.fscrawler:/root/.fscrawler \
     -v ~/tmp:/tmp/es:ro \
     dadoonet/fscrawler

Note:

  • ~/tmp contains the documents you want to index
  • Job settings will be stored in ~/.fscrawler/job_name/_settings.yaml

You can change the log level using the FS_JAVA_OPTS env variable:

docker run -it --rm \
     -v ~/.fscrawler:/root/.fscrawler \
     -v ~/tmp:/tmp/es:ro \
     -v ~/logs:/root/logs \
     -e FS_JAVA_OPTS="-DLOG_LEVEL=debug -DDOC_LEVEL=debug" \
     dadoonet/fscrawler

Read the documentation⁠ and specifically the "Using Docker"⁠ section for more details on how to use it.

Tag summary

Content type

Image

Digest

sha256:307784a8c…

Size

611.2 MB

Last updated

about 1 month ago

docker pull dadoonet/fscrawler