Sign inSign up

skilja/classificationservice

By skilja

•Updated 17 days ago

Tegra ClassificationService from Skilja GmbH

Image
Machine learning & AI
1

4.0K

skilja/classificationservice repository overview

⁠Quick reference

⁠Supported tags and Versioning

Image tags adhere to <major>.<minor>.<patch>.<revision> format.

<major>.<minor>.<patch>.<revision> points to a specific version. <Major>.<Minor> always points to the latest version. This version is compatible with all previous images of the same <Major>.<Minor> version. <latest> always points to the latest version, but such a version might require service and project database schema updates. For loading the latest but compatible version, we recommend to pull a <Major>.<Minor> with

docker pull skilja/classificationservice:<Major>.<Minor>

Note: Tags contain the optional suffix -llm or -llm-cuda.

-llm contains the LLM models LaBERTa small and large, which can be used with the content classifier (LLM supported). The execution is purely CPU-based and will be somewhat slower compared to the hardware-assisted versions in llm-cuda.

-llm-cuda contains the LLM models and the NVIDIA CUDA toolkit to use the content classifier (LLM supported) in a classification project. The LLM will be executed using NVIDIA CUDA GPU-based hardware acceleration and therefore will be faster (10x vs. CPU). However this requires a dedicated NVIDIA CUDA compatible GPU to run. You also have to run the container using the --gpus all parameter or equivalent with the deploy/devices keyword (https://docs.docker.com/compose/how-tos/gpu-support/⁠)

Regular images, e.g. tagged with 8.0, will also contain ONE LLM model LaBERTa small in order to use the content classifier (LLM supported) in a classification project.

Note: The latest tag (aka also no tag used) will always refer to the latest regular Image like 8.0

The most recent images are :

Warning : Breaking Changes with 8.0

  • Change of used service user (more info)⁠
  • Mandatory inital usage of the Site Management Tool (more info)⁠
  • Change of used Microsoft Read OCR Method: Document Intelligence Service Instead of Computer Vision

(More Information be found within the provided release notes)



logo drawing

⁠What is Tegra Classification Service?

The Classification Service allows you to classify documents for a given classification project to identify the document type of your processed documents. For example, to distinguish an invoice from a normal letter or from an accounting.

The service provides a REST API that allows implementing a custom application.

Classification Service is a part of Tegra Cognitive Services product suite. The suite provides various services to process documents and to perform classification and extraction. Furthermore, design applications are available to set up classification and extraction projects. You can monitor the document workflow using web applications that also provide license management.

Optionally, you can install the so-called Dashboard that provides a Web site with easy access to the installed Web applications.

Classification Service is license-protected. After the installation is completed, you need to get a license to register your system so that you are able to process documents.

⁠How to use this image

The following sample shows how to run the service using PostgreSQL:

1. Optionally for using a PostgreSQL database server, install and start a PostgreSQL container:

docker pull postgres:15.1
docker run -v <localpath>:/var/lib/postgresql/data -e POSTGRES_PASSWORD=MyPassword POSTGRES_USER=tegra  -p 5432:5432 -d postgres:15.1 

2. For the Classification Service, configure the environment file to use the user and password from above

CLSSERVICE_SERVICEENDPOINT=http://+:8043
CLSSERVICE_CLSSDATABASESERVERTYPE=2     (2 - PostgreSQL)
CLSSERVICE_CLSSDATABASESERVER=<database-server-name> (the host computer name in case of local docker installation)
CLSSERVICE_CLSSUSEINTEGRATEDSECURITY=false
CLSSERVICE_CLSSSQLUSER=tegra
CLSSERVICE_CLSSSQLPASSWORD=MyPassword

3. Pull and start the Classification Service image:

docker pull skilja/classificationservice  
docker run -p 8043:8043 -d --env-file envfile.txt skilja/classificationservice

⁠... via [docker-compose]

Example docker-compose.yml for skilja/classificationservice:

# This compose file will start a postgres db and the classification service at once # make sure you have a local DB folder created for the postgres db files

services: 
  postgres-db: 
    image:  postgres
    environment: 
      - POSTGRES_PASSWORD=myPassword
      - POSTGRES_USER=tegra
    ports: 
      - "5432:5432"
    volumes: 
      - ./DB:/var/lib/postgresql/data:z

  Tegra_Classify: 
    image:  skilja/classificationservice
    depends_on: 
      - postgres-db
    ports: 
      - "8043:8043"
    environment: 
      CLSSERVICE_SERVICEENDPOINT: http://+:8043
      CLSSERVICE_CLSSDATABASESERVERTYPE: 2
      CLSSERVICE_CLSSDATABASESERVER: postgres-db-1
      CLSSERVICE_CLSSUSESSL: false
      CLSSERVICE_CLSSTRUSTCERTIFICATE: true
      CLSSERVICE_CLSSUSEINTEGRATEDSECURITY: false
      CLSSERVICE_CLSSSQLUSER: tegra
      CLSSERVICE_CLSSSQLPASSWORD: myPassword

      
      #project database, required

      CLSSERVICE_PROJDATABASESERVERTYPE: 2
      CLSSERVICE_PROJDATABASESERVER: postgres-db-1
      CLSSERVICE_PROJUSESSL: false
      CLSSERVICE_PROJTRUSTCERTIFICATE: true
      CLSSERVICE_PROJSQLUSER: tegra
      CLSSERVICE_PROJSQLPASSWORD: myPassword
      CLSSERVICE_SELFHOSTWEBSITE: true
      CLSSERVICE_PROJUSEINTEGRATEDSECURITY: false

      #some customizations, optional
      CLSSERVICE_PARALLELWORKITEMS: 2
      CLSSERVICE_PARALLELOCRPROCESSES: 2
      CLSSERVICE_AUTOSCALINGCPUUSAGE: 80
      CLSSERVICE_AUTOSCALINGENABLED: true

      #this disables the need for an existing Authorization-Server, more on that later.

      CLSSERVICE_HASAUTHENABLED: false 

      #online learning, optional

      CLSSERVICE_OLDATABASESERVERTYPE: 2
      CLSSERVICE_OLDATABASESERVER: postgres-db-1
      CLSSERVICE_OLUSESSL: false
      CLSSERVICE_OLTRUSTCERTIFICATE: true
      CLSSERVICE_OLSQLUSER: tegra
      CLSSERVICE_OLSQLPASSWORD: myPassword
      CLSSERVICE_OLUSEINTEGRATEDSECURITY: false

Run docker compose up, wait for it to initialize completely, and visit http://localhost:8043/ or http://host-ip:8043 or http://cotainer-ip:8043

At http://localhost:8043/apihelp you'll find the API Overview which explains how to interact with the service.
You will also get a more sophisticated guide and help via the http://localhost:8043/help page.

Warning: The shown sample compose is running without Authorization! This is suitable for local development or testing but not for production. It is strongly recommended to use authorization in order to protect the service (licenses) and data that might be stored during processing (documents). Here you can find more information about the Authorization Server⁠. The option to use API Keys will also be enabled once you connected the Authorization Server.

⁠HTTPS

You probably noticed that the described sample is not using a SSL certificate. The currently recommended way is to use a Reverse Proxy like NGINX or Traefik.

⁠Supporting Modules

  • Authorization Server⁠ to provide single sign-on (SSO) for classification monitor and API Key support.

  • In more certain scenarios, documents must be extracted after beeing classified. This is where Extraction Service⁠ comes in handy. The setup is very similar and you can start them together with a joined docker compose.
    After setting up the Extraction Service you can connect both services via a feature called Extraction Bridge. It allows you to map certain classification classes to extraction document definitions.

⁠Service User

Beginning with version 8.0 the used service user is changed to be uniform across all Skilja containers. For .Net based containers like this one, the UID/GID 1654 is now used. This is a non root user called 'app'. This becomes important when working with Volume Mounts.

⁠Site Mangement Tool

When the Classification Service is started for the first time, users do not have any permissions assigned by default.
For the first configuration please start the Site Mangement Tool⁠. This step is only required to initially assign the Site Admin and Org Admin roles for the installation. The same initialization step is also required for the following components:

(AI Server is currently using its own builtin Site Mangement Tool and is therefore excluded)

For that reason, it is recommended to complete the full installation first and then run the Site Management Tool once to configure all components together.

⁠Environment Variables

The classificationervice image uses several environment variables, of which some are required others are optional. The following environment variables are provided.

⁠CLSSERVICE_SERVICEENDPOINT (ServiceEndpoint)

The URL where the classification service is exposed locally on the computer. This parameter is ignored when the service is hosted inside the IIS.

⁠CLSSERVICE_PATHBASE (PathBase)

When not empty, allows the application to be hosted at the desired base path. This may be necessary when running behind a reverse proxy or in containerized environments. When the base path is configured, the application still continues to be hosted at the root path as well.

For example, to make the application accessible for clients and services at http://server.com/classificationservice⁠, set the path base to /classificationservice.

⁠CLSSERVICE_CLSSDATABASESERVERTYPE (ClssDatabaseServerType)

The database server type to connect to as integer for the classification service database

  • 0 - SQL Server
  • 1 - Oracle database
  • 2 - PostgreSQL
⁠CLSSERVICE_CLSSDATABASESERVER (ClssDatabaseServer)

The database server name that hosts the classification service database.

⁠CLSSERVICE_CLSSUSESSL (ClssUseSSL)

Use SSL on the classification service database connection. This is currently supported for MS SQL server and PostgreSQL server.

⁠CLSSERVICE_CLSSTRUSTCERTIFICATE (ClssTrustCertificate)

If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to True for self-signed SSL certificates that are not officially trusted.

⁠CLSSERVICE_CLSSDATABASENAME (ClssDatabaseName)

The database name for the classification service database.

⁠CLSSERVICE_CLSSUSEINTEGRATEDSECURITY (ClssUseIntegratedSecurity)

When true, use integrated security for database access when false, use SQL user and password.

⁠CLSSERVICE_CLSSSQLUSER (ClssSQLUser)

SQL user name if integrated security is false.

⁠CLSSERVICE_CLSSSQLPASSWORD (ClssSQLPassword)

SQL password if integrated security is false.

⁠CLSSERVICE_OLDATABASESERVERTYPE (OLDatabaseServerType)

The database server type to connect to as integer for the online learning database.

  • 0 - SQL Server
  • 1 - Oracle database
  • 2 - PostgreSQL
⁠CLSSERVICE_OLDATABASESERVER (OLDatabaseServer)

Database server name hosting the online learning database.

⁠CLSSERVICE_OLDATABASENAME (OLDatabaseName)

The database name for the online learning database.

⁠CLSSERVICE_OLUSEINTEGRATEDSECURITY (OLUseIntegratedSecurity)

When true, use integrated security for the database access when false, use SQL user and password.

⁠CLSSERVICE_OLSQLUSER (OLSQLUser)

SQL user name if integrated security is false.

⁠CLSSERVICE_OLSQLPASSWORD (OLSQLPassword)

SQL password if integrated security is false.

⁠CLSSERVICE_OLUSESSL (OLUseSSL)

Use SSL on the online learning database connection. This is currently supported for MS SQL server and PostgreSQL server.

⁠CLSSERVICE_OLTRUSTCERTIFICATE (OLTrustCertificate)

If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to True for self-signed SSL certificates that are not officially trusted.

⁠CLSSERVICE_PROJISFILEBASED (PROJIsFileBased)

When true, loads classification projects from a directory. When false, loads projects from a database.

⁠CLSSERVICE_PROJBASEFOLDER (PROJBaseFolder)

The base folder for classification projects when working in folder-based mode.

⁠CLSSERVICE_PROJDATABASESERVERTYPE (PROJDatabaseServerType)

the database server type to connect to as integer for the classification project database

  • 0 - SQL Server
  • 1 - Oracle database
  • 2 - PostgreSQL
⁠CLSSERVICE_PROJDATABASESERVER (PROJDatabaseServer)

Database server name hosting the classification project database.

⁠CLSSERVICE_PROJDATABASENAME (PROJDatabaseName)

The database name for the classification project database.

⁠CLSSERVICE_PROJUSESSL (PROJUseSSL)

Use SSL encryption on the project database connection. This is currently supported for MS SQL server and PostgreSQL server.

⁠CLSSERVICE_PROJTRUSTCERTIFICATE (PROJTrustCertificate)

If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to True for self-signed SSL certificates that are not officially trusted.

⁠CLSSERVICE_PROJUSEINTEGRATEDSECURITY (PROJUseIntegratedSecurity)

When true, use integrated security for the database access to the project database when false, use SQL user and password.

⁠CLSSERVICE_PROJSQLUSER (PROJSQLUser)

SQL user name if integrated security is false.

⁠CLSSERVICE_PROJSQLPASSWORD (PROJSQLPassword)

SQL password name if integrated security is false.

⁠CLSSERVICE_SELFHOSTWEBSITE (SelfHostWebsite)

When true, the monitor web site is self hosted. Otherwise, it's not hosted automatically.

⁠CLSSERVICE_MAXPROJIDLESEC (MaxProjIdleSec)

Number of idle time in seconds until a cached classification project is unloaded from memory.

⁠CLSSERVICE_MAXOCRENGINEIDLESEC (MaxOCREngineIdleSec)

Number of idle time in seconds until an OCR process is unloaded to free memory.

⁠CLSSERVICE_PARALLELWORKITEMS (ParallelWorkItems)

Number of worker threads that are working in parallel. Worker threads are locking documents for OCR, classification, online learning and export.

⁠CLSSERVICE_PARALLELOCRPROCESSES (ParallelOCRProcesses)

Number of parallel processes started for OCR. OCR processes are shared between all worker threads. A single OCR process performs OCR on a single page at a time.

For stability reasons, the default is set to '2' processes. For systems with many cores, the number can be increased to '4' or '8'.

⁠CLSSERVICE_DEFAULTSLAPERIODMINUTES (DefaultSlaPeriodMinutes)

The default SLA time for all documents where the SLA time is not specified during the upload with the slaPeriodMinutes parameter. The default value if not configured is '120' minutes. The SLA time span gets added to the current time when uploading a document. The resulting time-to-end is then later used for the priority sorting.

⁠CLSSERVICE_INSTANCENAME (InstanceName)

The name shown in the monitor web application which has processed a document last.
If empty, the instance name is set to the current computer name. For docker instances with random names, it's recommended to set this property via an environment variable.

⁠CLSSERVICE_RESERVEDPROJECTNAMES (ReservedProjectNames)

This property can contain a list of project names separated by semicolon. This property is useful when running multiple instances of the classification service against the same database. It allows you to reserve a single instance to process only documents from a specific subset of projects.

⁠CLSSERVICE_PROCESSOTHERPROJECTSONIDLE (ProcessOtherProjectsOnIdle)

This property works in combination with the previous property (ReservedProjectNames). When set to true and ReservedProjectNames is not empty, the instance processes any document if no document is currently available from the reserved project list. If ReservedProjectNames is empty, the property has no meaning.

⁠CLSSERVICE_HASAUTHENABLED (HasAuthEnabled)

If this property is True, authentication is enabled for the classification service. Users need to log in via the configured authorization server and need to have access permissions with corresponding roles. Custom applications need to provide an API key with each API call. If this property is False, no authentication and no API keys are required.

⁠CLSSERVICE_AUTHSERVERURL (AuthServerUrl)

The URL of the authorization server if authentication is enabled.

⁠CLSSERVICE_WEBAPPCLIENTID (WebAppClientId)

The client ID of the web application that is registered as a public client within the authorization server.

⁠CLSSERVICE_SERVERCLIENTID (ServerClientId)

The client ID of the server side application that is registered as a confidential client within the authorization server.

⁠CLSSERVICE_SERVERCLIENTSECRET (ServerClientSecret)

The client secret corresponding to the client ID of the server side application.

⁠CLSSERVICE_ENABLEAUTODELETE (EnableAutoDelete)

This property is used to enable automatic deletion of documents in specific workflow states. If set to True, documents in 'Classified' and 'Done' state get deleted after a certain retention time.
By default, the property is set to False.

⁠CLSSERVICE_AUTODELETECLASSIFIED (AutoDeleteClassified)

If EnableAutoDelete is set to True, this property controls if documents in the 'Classified' state get automatically deleted. If set to False, no documents in the 'Classified' state are automatically deleted.

⁠CLSSERVICE_AUTODELETEDONE (AutoDeleteDone)

If EnableAutoDelete is set to True, this property controls if documents in the 'Done' state get automatically deleted. If set to False, no documents in the 'Done' state are automatically deleted.

⁠CLSSERVICE_AUTODELETEHOURS (AutoDeleteHours)

This property defines the retention time in hours before documents get deleted in 'Classified' or 'Done' status. The retention time is measured between the current time and the last modification time of a document. The default value is 48 hours.

⁠CLSSERVICE_LESA_USEGPU (LesaUseGPU)

By default a GPU is used automatically if detected. This property can be used to disable GPU usage. Set the property or environement variable to 'false' to disable GPU usage for Lesa OCR.

⁠CLSSERVICE_LESA_GPU_INDEX (LesaGPUIndex)

This property is relevant if a system has multiple GPUs available. By default the value is set to "-1" and Lesa uses all GPUs that are detected in a system. This property allows to restrict the Lesa OCR to use only a specific GPU specified by index. The property can be set to the 0-based index of the GPU that should be used. E.g. "0" to use the first GPU, "1" to use the second GPU or a comma separated list to use multiple GPUs. E.g. "2,3" to use the third and forth GPU of a system.

⁠CLSSERVICE_LESA_MAX_OCRPROC_PER_GPU (LesaMaxOCRProcPerGPU)

Defines the number of OCR processes running in parallel for each GPU. The default value is 1. With high end GPUs a higher throughput can be achieved by running 2-4 OCR processes per GPU. In combination with multiple GPUs, the final number of OCR processes will be this value multiplied by number of GPUs. This value significantly affects the OCR performance. If the value is too low, GPU usage is suboptimal. If too high it could lead to GPU memory overload. Please check the GPU usage during processing to find optimal value for your system. A single OCR processes normally needs 4 GByte GPU memory.

⁠OCR Speed and process priority

In some newer versions of Windows operating systems, a compute-intensive task running as a background process is not fully supported by the operating system. When the same task is run as a foreground process, it has been observed that to run almost twice as fast compared to running within a service or IIS application pool. The observed reason for the performance degradation was that the operating system drastically reduced the CPU frequency even during the execution of the OCR tasks. This affects the OCR speed in both CPU and GPU mode. To overcome this performance drop, there is an environment variable that increases the process priority of the OCR process. This option causes the operating system to keep the CPU frequency at the highest level during OCR execution. To increase the OCR process priority set the global environment variable OCR_HIGHER_PRIORITY to '1'. This option is currently not available within the settings file but only as an environment variable.

Please note that this type of configuration carries a certain risk. Running too many OCR processes can overload a system that does not have enough CPU or memory resources. If too many OCR processes are executed with a higher priority, the system may become unstable and unresponsive. Only use this setting if the CPU utilization is at or below 50% when OCR is running at full capacity.

Tag summary

Content type

Image

Digest

sha256:a66e024b9…

Size

1.1 GB

Last updated

17 days ago

docker pull skilja/classificationservice