Upload texts from a file share to PlagAware user library for plagiarism checking.
1.8K
PlagAware is an award-winning anti-plagiarism software suite. PlagAware is used by numerous educational institutions, publishers and individuals across the globe. In order to detect plagiarism and duplicate content, PlagAware compares texts against Internet sources and texts in user libraries. Texts in user libraries can either be uploaded manually, extracted from previous plagiarism checks or uploaded from existing databases or file shares.
LibCrawler provides you with an easy-to-use minimal script environment that searches for text documents in a file directory and uploads the texts into the PlagAware user library.
LibCrawler is built from a minimal Ubuntu 24.04 LTS image running PHP 8.3. It consists of two small PHP scripts located in /usr/lib/libcrawler/scripts. libcrawler.php is the main script invoking helper functions in libtools.php. Text extraction is perfomed by Apache Tika located in /usr/lib/libcrawler/tools.
LibCrawler is typically executed by the startup script /usr/local/sbin/libcrawler.sh, which is also the entry point of the Docker image.
After startup, LibCrawler scans the given root directory for any supported text documents. All text from the document will be extracted and normalized. Optionally, the text is transformed into a so called Fingerprint, which is a strongly compressed encrypted version of the text. The text fingerprint cannot be back-translated into the originating text and thus can be securely stored without violating copyrights.
Optionally, you can provide a JSON encrypted index file containing a file list and associated metadata to further describe your local documents.
After text normalization and transformation, the text or its fingerprint will be uploaded to your PlagAware user library. Using the /cache directory, LibCrawler keeps track of uploaded and skipped documents to prevent duplicate uploads.
After the first pass, LibCrawler will watch for modifications in the specified root folder or index file and upload newly identified documents.
/files- Root directory to search for files (mandatory)/cache- Root directory for caching (recommended). Specify a directory with 0777 permissions to persist cache between subsequent container runs.LC_API_KEY (mandatory) - API key as created within user profile at PlagAwareLC_API_ENDPOINT (optional) - Endpoint of PlagAware API, defaults to https://www.plagaware.com/api/LC_FINGERPRINT_ONLY (optional) - Convert texts to fingerprint (1, default) or upload full text (0)LC_PROJECT_ID (optional) - Id of project in PlagAware to be associated with the uploaded filesLC_INDEX_FILE (optional) - Filename of a JSON index file. If specified, LibCrawler expects the index file inside the /files volume. Instead of crawling the /files volume, only files specified in the index file will be considered.LC_SCAN_INTERVAL (optional) - Sleep for n seconds after re-checking for updated file in the /files directory or in the specified index file. Defaults to 300 (5 minutes).LC_INDEX_FILE (optional) - JSON index file path relative to /files directory specifiying the file list to be uploaded and additional metadata for the files.LC_DELETE_PROCESSED_FILES (optional) - Set to DELETE in order to delete files from /files volume after successful API call to PlagAware (either added to PlagAware library or skipped because it existed already). Use with caution, this will PERMENENTLY DELETE all files from your crawl directory! Please mount /files volume writable(rw) for this to work and confirm with LC_CONFIRM_DELETE_PROCESSED_FILES.LC_CONFIRM_DELETE_PROCESSED_FILES (opotional) - Set to CONFIRMED in order to reconfirm permanent file deletion after upload (see LC_DELETE_PROCESSED_FILES).version: "3.7"
services:
libcrawler:
image: plagaware/libcrawler:latest
volumes:
- /my/path/to/files:/files:ro # root folder for crawling files (read only)
- /my/path/to/cache:/cache:rw # root folder for caching files (read/write, 0777 permissions)
environment:
LC_API_KEY: api_key # API user code, create at www.plagaware.com
restart: unless-stopped
/cache directoryLC_MAX_FILES removed/files directoryLC_INDEX_FILE to specifiy metadata of files to be uploaded/cache/libcrawler_fileslog.json to track uploaded filesFor further support, please reach out to our PlagAware Support Team.
Content type
Image
Digest
sha256:1db732010…
Size
422.8 MB
Last updated
over 1 year ago
docker pull plagaware/libcrawler