Sign inSign up

plagaware/libcrawler

By plagaware

•Updated over 1 year ago

Upload texts from a file share to PlagAware user library for plagiarism checking.

Image
1

1.8K

plagaware/libcrawler repository overview

⁠PlagAware LibCrawler

⁠What is PlagAware LibCrawler?

PlagAware is an award-winning anti-plagiarism software suite⁠. PlagAware is used by numerous educational institutions, publishers and individuals across the globe. In order to detect plagiarism and duplicate content, PlagAware compares texts against Internet sources and texts in user libraries. Texts in user libraries can either be uploaded manually, extracted from previous plagiarism checks or uploaded from existing databases or file shares.

LibCrawler provides you with an easy-to-use minimal script environment that searches for text documents in a file directory and uploads the texts into the PlagAware user library.

⁠Prerequisites

  • In order to use LibCrawler, you need to create a PlagAware user account at PlagAware⁠.
  • Docker engine installed on your client machine.

⁠What's in the package?

LibCrawler is built from a minimal Ubuntu 24.04 LTS image running PHP 8.3. It consists of two small PHP scripts located in /usr/lib/libcrawler/scripts. libcrawler.php is the main script invoking helper functions in libtools.php. Text extraction is perfomed by Apache Tika⁠ located in /usr/lib/libcrawler/tools.

LibCrawler is typically executed by the startup script /usr/local/sbin/libcrawler.sh, which is also the entry point of the Docker image.

⁠How does it work?

After startup, LibCrawler scans the given root directory for any supported text documents. All text from the document will be extracted and normalized. Optionally, the text is transformed into a so called Fingerprint, which is a strongly compressed encrypted version of the text. The text fingerprint cannot be back-translated into the originating text and thus can be securely stored without violating copyrights.

Optionally, you can provide a JSON encrypted index file containing a file list and associated metadata to further describe your local documents.

After text normalization and transformation, the text or its fingerprint will be uploaded to your PlagAware user library. Using the /cache directory, LibCrawler keeps track of uploaded and skipped documents to prevent duplicate uploads.

After the first pass, LibCrawler will watch for modifications in the specified root folder or index file and upload newly identified documents.

⁠Setup

⁠Volumes
  • /files- Root directory to search for files (mandatory)
  • /cache- Root directory for caching (recommended). Specify a directory with 0777 permissions to persist cache between subsequent container runs.
⁠Environment Variables
  • LC_API_KEY (mandatory) - API key as created within user profile at PlagAware
  • LC_API_ENDPOINT (optional) - Endpoint of PlagAware API, defaults to https://www.plagaware.com/api/
  • LC_FINGERPRINT_ONLY (optional) - Convert texts to fingerprint (1, default) or upload full text (0)
  • LC_PROJECT_ID (optional) - Id of project in PlagAware to be associated with the uploaded files
  • LC_INDEX_FILE (optional) - Filename of a JSON index file. If specified, LibCrawler expects the index file inside the /files volume. Instead of crawling the /files volume, only files specified in the index file will be considered.
  • LC_SCAN_INTERVAL (optional) - Sleep for n seconds after re-checking for updated file in the /files directory or in the specified index file. Defaults to 300 (5 minutes).
  • LC_INDEX_FILE (optional) - JSON index file path relative to /files directory specifiying the file list to be uploaded and additional metadata for the files.
  • LC_DELETE_PROCESSED_FILES (optional) - Set to DELETE in order to delete files from /files volume after successful API call to PlagAware (either added to PlagAware library or skipped because it existed already). Use with caution, this will PERMENENTLY DELETE all files from your crawl directory! Please mount /files volume writable(rw) for this to work and confirm with LC_CONFIRM_DELETE_PROCESSED_FILES.
  • LC_CONFIRM_DELETE_PROCESSED_FILES (opotional) - Set to CONFIRMED in order to reconfirm permanent file deletion after upload (see LC_DELETE_PROCESSED_FILES).
⁠Minimal Docker Compose file
version: "3.7"
services:
 libcrawler:
  image: plagaware/libcrawler:latest
  volumes:
  - /my/path/to/files:/files:ro        # root folder for crawling files (read only)
  - /my/path/to/cache:/cache:rw        # root folder for caching files (read/write, 0777 permissions)
  environment:
    LC_API_KEY: api_key                                 # API user code, create at www.plagaware.com
  restart: unless-stopped

⁠Hints and Remarks

  • Crawl directory can and should be mounted read only.
  • For organizations, a technical user should be used to upload files to the library.
  • Remove the /cache/lastscan file to reindex all files
  • PlagAware offers a Moodle Plugin which allows the upload of documents stored in Moodle to PlagAware

⁠Release Notes

⁠Changes in 1.1
  • Fix: Support for German Umlaute in file conversion
  • Change: Caching method changed from caching file database to text files in /cache directory
  • Change: Optional parameter LC_MAX_FILES removed
  • Added: Support for file deletion from /files directory
  • Added: Support for JSON index file LC_INDEX_FILE to specifiy metadata of files to be uploaded
  • Added: Creation of JSON log file /cache/libcrawler_fileslog.json to track uploaded files

⁠Support

For further support, please reach out to our PlagAware Support Team⁠.

Tag summary

Content type

Image

Digest

sha256:1db732010…

Size

422.8 MB

Last updated

over 1 year ago

docker pull plagaware/libcrawler