Sign inSign up

skilja/extractionservice

By skilja

•Updated 17 days ago

Tegra ExtractionService from Skilja GmbH

Image
Machine learning & AI
3

10K+

skilja/extractionservice repository overview

⁠Quick reference

⁠Supported tags and Versioning

Image tags adhere to <major>.<minor>.<patch>.<revision> format.

<major>.<minor>.<patch>.<revision> points to a specific version. <Major>.<Minor> always points to the latest version. This version is compatible with all previous images of the same <Major>.<Minor> version. <latest> always points to the latest version, but such a version might require service and project database schema updates. For loading the latest but compatible version, we recommend to pull a <Major>.<Minor> with

docker pull skilja/extractionservice:<Major>.<Minor>

Note: Tags contain the optional suffix -llm or -llm-cuda.

-llm contains the LLM models LaBERTa small and large, which can be used with LLM-based extraction. The execution is purely CPU-based and will be somewhat slower compared to the hardware-assisted versions in llm-cuda.

-llm-cuda contains the LLM models and the NVIDIA CUDA toolkit to use LLM-based extraction in an extraction project. The LLM will be executed using NVIDIA CUDA GPU-based hardware acceleration and therefore will be faster (10x vs. CPU). However this requires a dedicated NVIDIA CUDA compatible GPU to run. You also have to run the container using the --gpus all parameter or equivalent with the deploy/devices keyword ([https://docs.docker.com/compose/how-tos/gpu-support/⁠])

Regular images, e.g. tagged with 8.0, will also contain one LLM model LaBERTa small in order to use LLM-based extraction in an extraction project.

The latest image tags are :

Warning : Breaking Changes with 8.0

  • Change of used service user (more info)⁠
  • Mandatory inital usage of the Site Management Tool (more info)⁠
  • Change of used Microsoft Read OCR Method: Document Intelligence Service Instead of Computer Vision

(More Information be found within the provided release notes)

logo drawing

⁠What is Tegra Extraction Service?

Extraction Service allows you to extract documents contents using an extraction project. For example, to get an invoice date or an invoice number from an invoice.

The service provides a REST API that can is typically used to integrate the cognitive features to your custom application.

Extraction Service is a part of the Tegra product suite, which provides services to conduct OCR and to perform classification or extraction tasks. Projects can be created using Designer applications which are available separately. You can monitor the document workflow through the monitor applications exposed by the container.

Extraction Service requires a license to run. After the deployment is completed, you need to obtain a license from Skilja and import it into your system using the monitor web application.

⁠How to use this image

The following sample shows how to run the service using PostgreSQL:

1. Optionally for using a PostgreSQL database server, install and start a PostgreSQL container:

docker pull postgres:15.1
docker run -v <localpath>:/var/lib/postgresql/data -e POSTGRES_PASSWORD=MyPassword POSTGRES_USER=tegra  -p 5432:5432 -d postgres:15.1 

2. For the Extraction Service, configure the environment file to use the user and password from above

EXTRSERVICE_SERVICEENDPOINT=http://+:8044
EXTRSERVICE_EXTRSDATABASESERVERTYPE=2     (2 - PostgreSQL)
EXTRSERVICE_EXTRSDATABASESERVER=<database-server-name> (the host computer name in case of local docker installation) EXTRSERVICE_EXTRSUSEINTEGRATEDSECURITY=false
EXTRSERVICE_EXTRSSQLUSER=tegra
EXTRSERVICE_EXTRSSQLPASSWORD=MyPassword

3. Pull and start the Extraction Service image:

docker pull skilja/extractionservice
docker run -p 8044:8044 -d --env-file envfile.txt skilja/extractionservice

⁠... via [docker-compose]

Example docker-compose.yml for skilja/extractionservice:

# This compose file will start a postgres db and the extraction service at once # make sure you have a local DB folder created for the postgres db files

services: 
  postgres-db: 
    image:  postgres
    environment: 
      - POSTGRES_PASSWORD=postgres
      - POSTGRES_USER=tegra
    ports: 
      - "5432:5432"
    volumes: 
      - ./DB:/var/lib/postgresql/data:z

  Tegra_Extract: 
    image:  skilja/extractionservice
    depends_on: 
      - postgres-db
    links: 
      - postgres-db
    ports: 
      - "8044:8044"
    environment: 
      EXTRSERVICE_SERVICEENDPOINT: http://+:8044
      EXTRSERVICE_EXTRSDATABASESERVERTYPE: 2
      EXTRSERVICE_EXTRSDATABASESERVER: postgres-db-1
      EXTRSERVICE_EXTRSUSESSL: false
      EXTRSERVICE_EXTRSTRUSTCERTIFICATE: true
      EXTRSERVICE_EXTRSUSEINTEGRATEDSECURITY: false
      EXTRSERVICE_EXTRSSQLUSER: tegra
      EXTRSERVICE_EXTRSSQLPASSWORD: myPassword
      
      #project database, required

      EXTRSERVICE_PROJDATABASESERVERTYPE: 2
      EXTRSERVICE_PROJDATABASESERVER: postgres-db-1
      EXTRSERVICE_PROJUSESSL: false
      EXTRSERVICE_PROJTRUSTCERTIFICATE: true
      EXTRSERVICE_PROJSQLUSER: tegra
      EXTRSERVICE_PROJSQLPASSWORD: postgres
      EXTRSERVICE_PROJUSEINTEGRATEDSECURITY: false

      #some customizations, optional
      EXTRSERVICE_PARALLELWORKITEMS: 2
      EXTRSERVICE_PARALLELOCRPROCESSES: 2
      EXTRSERVICE_AUTOSCALINGCPUUSAGE: 80
      EXTRSERVICE_AUTOSCALINGENABLED: true

      #this disables the need for an existing Authorization-Server
      EXTRSERVICE_HASAUTHENABLED: false 

Run docker compose up, wait for it to initialize completely, and visit http://localhost:8044/ or http://host-ip:8044 or http://container-ip:8044

At http://localhost:8044/extractionservice/apihelp you'll find the API Overview which explains how to interact with the service.
You will also get a more sophisticated guide and help via the http://localhost:8044/extractionservice/help page.

Warning: The shown sample compose is running without Authorization! It is strongly recommended to use authorization in order to protect the service (licenses) and data that might be stored during processing (documents). Here you can find more information about the Authorization Server⁠. The option to use API Keys will also be enabled once you connected the Authorization Server.

⁠HTTPS

You probably noticed that the described sample is not using an SSL certificate. The currently recommended way is to use a Reverse Proxy like NGIX or Traefik.

⁠Supporting Modules

  • Authorization Server⁠ to provide single sign-on (SSO) for extraction monitor.

  • In more certain scenarios, documents must be classified before being extracted. This is where Classification Service⁠ comes in handy. The setup is similar and you can start them together with a joined docker compose.
    After setting up the Classification Service you can connect both services via a feature called Extraction Bridge. It allows you to map certain Classification classes to extraction document definitions. You can find more information about this inside the Classification Service⁠ help.

⁠Service User

Beginning with version 8.0 the used service user is changed to be uniform across all Skilja containers. For .Net based containers like this one, the UID/GID 1654 is now used. This is a non root user called 'app'. This becomes important when working with Volume Mounts.

⁠Site Mangement Tool

When the Classification Service is started for the first time, users do not have any permissions assigned by default.
For the first configuration please start the Site Mangement Tool⁠. This step is only required to initially assign the Site Admin and Org Admin roles for the installation. The same initialization step is also required for the following components:

(AI Server is currently using its own builtin Site Mangement Tool and is therefore excluded)

For that reason, it is recommended to complete the full installation first and then run the Site Management Tool once to configure all components together.

⁠Environment Variables

The extractionservice image uses several environment variables, of which some are required others are optional.

Note: Due to Docker Hub size limitations, not all available environment variables are listed here. To see all available Environment Variables please visit the Online Documentation ⁠ at the Partner Portal.

⁠EXTRSERVICE_SERVICEENDPOINT

The environment variable EXTRSERVICE_SERVICEENDPOINT sets the URL where the extraction service is exposed locally on the computer. The default value if not configured is "http://+:8044".

⁠EXTRSERVICE_PATHBASE

When not empty, the environment variable EXTRSERVICE_PATHBASE (PathBase) allows the application to be hosted at the desired base path. This may be necessary when running behind a reverse proxy or in containerized environments. When the base path is configured, the application still continues to be hosted at the root path as well.

For example, to make the application accessible for clients and services at http://server.com/extractionservice⁠, set the path base to "/extractionservice".

⁠EXTRSERVICE_EXTRSDATABASESERVERTYPE

The environment variable EXTRSERVICE_EXTRSDATABASESERVERTYPE sets the database server type to connect to as integer for the extraction service database.

- 0 - SQL Server
- 1 - Oracle database
- 2 - PostgreSQL
⁠EXTRSERVICE_EXTRSDATABASESERVER

The environment variable EXTRSERVICE_EXTRSDATABASESERVER defines the database server name that hosts the extraction service database.

⁠EXTRSERVICE_EXTRSDATABASENAME

The environment variable EXTRSERVICE_EXTRSDATABASENAME sets the database name for the extraction service database.

⁠EXTRSERVICE_EXTRSUSEINTEGRATEDSECURITY

The environment variable EXTRSERVICE_EXTRSUSEINTEGRATEDSECURITY set to true, uses integrated security for database access. When set to false, it uses SQL user and password.

⁠EXTRSERVICE_EXTRSUSESSL

The environment variable EXTRSERVICE_EXTRSUSESSL defines whether to use SSL on the extraction service database connection or not. This is currently supported for MS SQL server and PostgreSQL server.

⁠EXTRSERVICE_EXTRSTRUSTCERTIFICATE

The environment variable EXTRSERVICE_EXTRSTRUSTCERTIFICATE is only relevant when UseSSL is true. If set to true, the service trusts the SSL certificate even if not publicly-signed. Otherwise, the service refuses the connection with incorrect or self-signed SSL certificates. This is currently supported for MS SQL server and PostgreSQL server.

⁠EXTRSERVICE_EXTRSSQLUSER

The environment variable EXTRSERVICE_EXTRSSQLUSER sets the SQL user name if integrated security is false.

⁠EXTRSERVICE_EXTRSSQLPASSWORD

The environment variable EXTRSERVICE_EXTRSSQLPASSWORD sets the SQL password name if integrated security is false.

⁠EXTRSERVICE_PROJBASEFOLDER

The environment variable EXTRSERVICE_PROJBASEFOLDER defines the base folder for extraction projects when working in folder based mode.

⁠EXTRSERVICE_PROJDATABASESERVERTYPE

The environment variable EXTRSERVICE_PROJDATABASESERVERTYPE defines the database server type to connect to as integer for the project database

Same values like in : EXTRSERVICE_EXTRSDATABASESERVERTYPE

⁠EXTRSERVICE_PROJDATABASESERVER

The environment variable EXTRSERVICE_PROJDATABASESERVER defines the database server name that hosts the extraction project database.

⁠EXTRSERVICE_PROJDATABASENAME

The environment variable EXTRSERVICE_PROJDATABASENAME defines the database name for the extraction project database.

⁠EXTRSERVICE_PROJUSEINTEGRATEDSECURITY

The environment variable EXTRSERVICE_PROJUSEINTEGRATEDSECURITY set to true, use integrated security for database access to the project database when false, uses SQL user and password.

⁠EXTRSERVICE_PROJUSESSL

The environment variable EXTRSERVICE_PROJUSESSL set to true uses SSL encryption on the project database connection. This is currently supported for MS SQL server and PostgreSQL server.

⁠EXTRSERVICE_PROJTRUSTCERTIFICATE

The environment variable EXTRSERVICE_PROJTRUSTCERTIFICATE is only relevant when UseSSL is true. If set to true, the service trusts the SSL certificate even if not publicly-signed. Otherwise, it refuses the connection with incorrect or self-signed SSL certificates. This is currently supported for MS SQL server and PostgreSQL server.

⁠EXTRSERVICE_PROJSQLUSER

The environment variable EXTRSERVICE_PROJSQLUSER sets the SQL user name if integrated security is false.

⁠EXTRSERVICE_PROJSQLPASSWORD

The environment variable EXTRSERVICE_PROJSQLPASSWORD sets the SQL password name if integrated security is false.

⁠EXTRSERVICE_SELFHOSTWEBSITE

The environment variable EXTRSERVICE_SELFHOSTWEBSITE set to true means that the monitor web site is self-hosted, otherwise, it's not hosted automatically.

⁠EXTRSERVICE_PARALLELWORKITEMS

The environment variable EXTRSERVICE_PARALLELWORKITEMS defines the number of worker threads that are working in parallel.

⁠EXTRSERVICE_PARALLELOCRPROCESSES

The environment variable EXTRSERVICE_PARALLELOCRPROCESSES defines the number of parallel processes started for OCR.
For stability reasons, the default is set to '2' processes. For systems with many cores, you can increase the number to '4' or '8'.

⁠EXTRSERVICE_AUTOSCALINGENABLED

The environment variable EXTRSERVICE_AUTOSCALINGENABLED set to true means that ParallelWorkItems and ParallelOCRProcesses are interpreted differently. If this option is enabled, the configuration adjusts itself automatically to the actual number of CPU cores so that you can use the same configuration file on systems of different size and CPU count. Please visit the Online Documentation⁠ for more information

⁠EXTRSERVICE_AUTOSCALINGCPUUSAGE

The environment variable EXTRSERVICE_AUTOSCALINGCPUUSAGE is used only in combination with AutoScalingEnabled = 'True'.

This option defines how much CPU usage should be consumed. AutoScalingCpuUsage 1-100.

⁠EXTRSERVICE_LESA_USEGPU

By default a GPU is used automatically if detected. This property can be used to disable GPU usage. Set the property or environement variable to 'false' to disable GPU usage for Lesa OCR.

⁠EXTRSERVICE_LESA_GPU_INDEX

This property is relevant if a system has multiple GPUs available. By default the value is set to "-1" and Lesa uses all GPUs that are detected in a system. Please visit the Online Documentation ⁠ for more information

⁠EXTRSERVICE_LESA_MAX_OCRPROC_PER_GPU

Defines the number of OCR processes running in parallel for each GPU. The default value is 1. With high end GPUs a higher throughput can be achieved by running 2-4 OCR processes per GPU. Please visit the Online Documentation ⁠ for more information

⁠Oracle connection modes

For Oracle databases, the basic connection mode and TNS connection mode are supported. Please visit the Online Documentation⁠ for more information.

⁠Other Platforms

The service is also available for Windows installation. Please visit The Partner Portal⁠ or write us a email [email protected]⁠⁠ for more information.

Tag summary

Content type

Image

Digest

sha256:7d6503f11…

Size

1.3 GB

Last updated

17 days ago

docker pull skilja/extractionservice