Sign inSign up

skilja/classificationdesigner

By skilja

•Updated 17 days ago

Designer for classification projects

Image
Machine learning & AI
Data science
0

2.5K

skilja/classificationdesigner repository overview

⁠Quick reference

⁠Supported tags and Versioning

Image tags adhere to <major>.<minor>.<patch>.<revision> format.

<major>.<minor>.<patch>.<revision> points to a specific version. <Major>.<Minor> always points to the latest version. This version is compatible with all previous images of the same <Major>.<Minor> version. <latest> always points to the latest version, but such a version might require service and project database schema updates. For loading the latest but compatible version, we recommend to pull a <Major>.<Minor> with

docker pull skilja/classificationdesigner:<Major>.<Minor>

Note: Tags contain the optional suffix -llm-cuda.

-llm-cuda contains the LLM models and the NVIDIA CUDA toolkit to use LLM-based classification in an classification project. The LLM will be executed using NVIDIA CUDA GPU-based hardware acceleration and therefore will be faster (10x vs. CPU). However this requires a dedicated NVIDIA CUDA compatible GPU to run. You also have to run the container using the --gpus all parameter or equivalent with the deploy/devices keyword (https://docs.docker.com/compose/how-tos/gpu-support/⁠)

Regular images, e.g. tagged with 8.0, will also contain one LLM model LaBERTa small in order to use LLM-based classification.



The latest image tags are :

Warning : Breaking Changes with 8.0

  • Change of used service user (more info)⁠
  • Mandatory inital usage of the Site Management Tool (more info)⁠
  • Change of used Microsoft Read OCR Method: Document Intelligence Service Instead of Computer Vision

(More Information be found within the provided release notes)

logo drawing

⁠What is the Classification Designer?

In order to configure an classification project you need an Classification Designer. An classification project is created with a so-called designer that comes in two flavors. Windows Designer and Web Designer. This container hosts the Web Designer. The Classification Web Designer application allows an efficient configuration of classification projects and provides a user-friendly interface.

You can use both designers in parallel and the created projects are compatible. It is absolutely feasible to start the creation of a project and its classes with Classification Web Designer and then refine it later with access to the inner workings.

The Classification Web Designer is a part of the Laera product suite, which provides services to conduct OCR and to perform classification or extraction tasks.

The Classification Web Designer requires a valid license to operate. We recommend obtaining your license in advance and specifying it during initial deployment, so you won't need to update environment variables or restart the container later. The license is valid for a specific version and until a defined date. Upgrading to a new major or minor version or exceeding the expiration date requires a new license.

The Classification Web Designer requieres also an own "Classification Projects" database. If not existing, the needed tables will be created on the first launch. If there are pre existing tables, those will be used.

Note: The web designer in comparison to the Windows application has some restrictions and simplifications and it does not provide the full range of all options yet. For more information check the in app help, once deployed.

⁠How to use this image

The following sample shows how to run the service using PostgreSQL:

1. Optionally for using a PostgreSQL database server, install and start a PostgreSQL container:

docker pull postgres:17.1
docker run -v <localpath>:/var/lib/postgresql/data -e POSTGRES_PASSWORD=MyPassword POSTGRES_USER=laera  -p 5432:5432 -d postgres:17.1 

2. For the Classification Web Designer, configure the environment file to use the user and password from above

CLSDESIGNER_SERVICEENDPOINT: http://+:8045
CLSDESIGNER_DATABASESERVERTYPE: 2         (2 - PostgreSQL)
CLSDESIGNER_DATABASESERVER: <database-server-name> (the host computer name in case of local docker installation)
CLSDESIGNER_DATABASENAME: <classification project db name>
CLSDESIGNER_MAXOCRPROCESSCOUNT: 1 (stable default, can go higher if system is capeable)
CLSDESIGNER_USEINTEGRATEDSECURITY: "false"
CLSDESIGNER_SQLUSER: "laera"
CLSDESIGNER_SQLPASSWORD: "MyPassword"
CLSDESIGNER_LICENSE: "<License Token>"
CLSDESIGNER_AUTOCREATEDATABASE : "true"   (Set if the classification project database should be created automatically)

3. Pull and start the Classification Web Designer image:

docker pull skilja/classificationdesigner
docker run -p 8045:8045 -d --env-file envfile.txt skilja/classificationdesigner

⁠... via [docker-compose]

Example docker-compose.yml for skilja/classificationdesigner:

# This compose file will start a postgres db and the Classification Web Designer at once # make sure you have a local DB folder created for the postgres db files

services: 
  postgres-db: 
    image:  postgres

    healthcheck:
        test: "pg_isready -U tegra"
        interval: 5s
        timeout: 5s
        retries: 5

    environment: 
      - POSTGRES_PASSWORD=postgres
      - POSTGRES_USER=tegra
    ports: 
      - "5432:5432"
    volumes: 
      - ./DB:/var/lib/postgresql/data:z

  Laera_ClassificationDesigner: 
    image:  skilja/classificationdesigner
    depends_on:
       postgres-db:
         condition: service_healthy
    links: 
      - postgres-db
    ports: 
      - "8045:8045"
    environment: 
      CLSDESIGNER_SERVICEENDPOINT: http://+:8045
      CLSDESIGNER_DATABASESERVERTYPE: 2
      CLSDESIGNER_DATABASESERVER: "postgres-db-1"
      CLSDESIGNER_DATABASENAME: "ClassificationProjects"
      CLSDESIGNER_USEINTEGRATEDSECURITY: "False"
      CLSDESIGNER_SQLUSER: "tegra"
      CLSDESIGNER_SQLPASSWORD: "postgres"
      CLSDESIGNER_LICENSE: "<License Token>"
      CLSDESIGNER_AUTOCREATEDATABASE : "True"
      CLSDESIGNER_MAXOCRPROCESSCOUNT: 1
      CLSDESIGNER_HASAUTHENABLED: "False" 
      #this disables the need for an existing Authorization-Server, more on that later.

Run docker compose up, wait for it to initialize completely, and visit http://localhost:8045/ or http://host-ip:8045 or http://container-ip:8045

You will also get a more sophisticated guide and help via the http://localhost:8045/classificationdesigner/help page.

Warning: The shown sample compose is running without Authorization! It is strongly recommended to use authorization in order to protect the service (licenses) and data that might be stored during processing (documents). Here you can find more information about the Authorization Server⁠.

⁠HTTPS

You probably noticed that the described sample is not using an SSL certificate. The currently recommended way is to use a Reverse Proxy like NGINX reverse proxy or Traefik.

⁠Supporting Modules

⁠Service User

Beginning with version 8.0 the used service user is changed to be uniform across all Skilja containers. For .Net based containers like this one, the UID/GID 1654 is now used. This is a non root user called 'app'. This becomes important when working with Volume Mounts.

⁠Site Mangement Tool

When the Classification Service is started for the first time, users do not have any permissions assigned by default.
For the first configuration please start the Site Mangement Tool⁠. This step is only required to initially assign the Site Admin and Org Admin roles for the installation. The same initialization step is also required for the following components:

(AI Server is currently using its own builtin Site Mangement Tool and is therefore excluded)

For that reason, it is recommended to complete the full installation first and then run the Site Management Tool once to configure all components together.

⁠Environment Variables

The Classification Web Designer image uses several environment variables, of which some are required others are optional.

⁠CLSDESIGNER_SERVICEENDPOINT (ServiceEndpoint)

The URL where the classification service is exposed locally on the computer. This parameter is ignored when the service is hosted inside the IIS.

⁠CLSDESIGNER_PATHBASE (PathBase)

When not empty, allows the application to be hosted at the desired base path. This may be necessary when running behind a reverse proxy or in containerized environments. When the base path is configured, the application still continues to be hosted at the root path as well.

For example, to make the application accessible at http://server.com/classificationdesigner/⁠, set the path base to "/classificationdesigner".

⁠CLSDESIGNER_DATABASESERVERTYPE (DatabaseServerType)

The database server type to connect to as integer for the classification project database

  • 0 - SQL Server
  • 1 - Oracle database
  • 2 - PostgreSQL
⁠CLSDESIGNER_DATABASESERVER (DatabaseServer)

The database server name that hosts the classification project database.

⁠CLSDESIGNER_USESSL (UseSSL)

Use SSL on the classification project database connection. This is currently supported for MS SQL server and PostgreSQL server.

⁠CLSDESIGNER_TRUSTCERTIFICATE (TrustCertificate)

If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to True for self-signed SSL certificates that are not officially trusted.

⁠CLSDESIGNER_DATABASENAME (DatabaseName)

The database name for the classification project database.

⁠CLSDESIGNER_USEINTEGRATEDSECURITY (UseIntegratedSecurity)

When true, use integrated security for database access when false, use SQL user and password.

⁠CLSDESIGNER_SQLUSER (SQLUser)

SQL user name if integrated security is false.

⁠CLSDESIGNER_SQLPASSWORD (SQLPassword)

SQL password if integrated security is false.

⁠CLSDESIGNER_AUTOCREATEDATABASE (AutoCreateDatabase)

When the property is true, the classification project database is created automatically. Otherwise, the database is not created.

⁠CLSDESIGNER_SELFHOSTWEBSITE (SelfHostWebsite)

When true, the Classification Web Designer site is self hosted. Otherwise, it's not hosted automatically.

⁠CLSDESIGNER_HASAUTHENABLED (HasAuthEnabled)

If this property is True, authentication is enabled for the Classification Web Designer. Users need to log in via the configured authorization server and need to have access permissions with corresponding roles. Custom applications need to provide an API key with each API call. If this property is False, no authentication and no API keys are required.

⁠CLSDESIGNER_AUTHSERVERURL (AuthServerUrl)

The URL of the authorization server if authentication is enabled.

⁠CLSDESIGNER_WEBAPPCLIENTID (WebAppClientId)

The client ID of the web application that is registered as a public client within the authorization server.

⁠CLSDESIGNER_SERVERCLIENTID (ServerClientId)

The client ID of the server side application that is registered as a confidential client within the authorization server.

⁠CLSDESIGNER_SERVERCLIENTSECRET (ServerClientSecret)

The client secret corresponding to the client ID of the server side application.

⁠CLSDESIGNER_LICENSE (License)

Running the Classification Web Designer in Docker requires a license. Copy the token from the license file into the environment variable to start the Classification Web Designer in Docker. The license is valid for a specific version and until a defined date. Upgrading to a new major or minor version or exceeding the expiration date requires a new license.

⁠CLSDESIGNER_PROJECTUNLOADIDLETIME (ProjectUnloadIdleTime)

Unload classification projects in idle state after specified time in seconds.

⁠CLSDESIGNER_READLOCKTIMEOUT (ReadLockTimeout)

Defines the maximum time to wait to get a read-lock on a single project.

⁠CLSDESIGNER_WRITELOCKTIMEOUT (WriteLockTimeout)

Defines the maximum time to wait to get a write-lock on a single project.

⁠CLSDESIGNER_MAXOCRPROCESSCOUNT (MaxOCRProcessCount)

Maximum number of parallel processes for OCR.

⁠CLSDESIGNER_MAXCONCURRENTJOBCOUNT (MaxConcurrentJobCount)

Maximum number of concurrent jobs that can run in parallel.

⁠CLSDESIGNER_OLSERVICEISINSTALLED (OLServiceIsInstalled)

When true, the Classification Monitor is installed. Otherwise, online learning is not configured.

⁠CLSDESIGNER_OLSERVICEENDPOINT (OLServiceEndpoint)

The endpoint of the Classification Monitor.

⁠`CLSDESIGNER_MAXREQUESTBODYSIZE (MaxRequestBodySize)

Defines the maximum body size in bytes for requests that are sent to the Classification Web Designer. Limits the size of projects or files that can be uploaded.

⁠CLSDESIGNER_MAXPAGESTESTSEPARATION (MaxPagesTestSeparation)
⁠CLSDESIGNER_ISMICROSOFTREADOCRDOCKERBASE (IsMicrosoftReadOCRDockerBase)

When true, the Microsoft Read OCR is running in Docker. This requires an API key.

⁠CLSDESIGNER_MICROSOFTREADOCRENDPOINTURL (MicrosoftReadOCREndpointUrl)

The endpoint of the Microsoft Read Service.

⁠CLSDESIGNER_MICROSOFTREADOCRAPIKEY (MicrosoftReadOCRApiKey)

The API key for Microsoft Read OCR when running in Docker.

⁠CLSDESIGNER_MICROSOFTREADOCRMODELVERSION (MicrosoftReadOCRModelVersion)

The version of the Microsoft Read OCR model.

⁠CLSDESIGNER_LESA_USEGPU (LesaUseGPU)

By default a GPU is used automatically if detected. This property can be used to disable GPU usage. Set the property or environement variable to 'false' to disable GPU usage for Lesa OCR.

⁠CLSDESIGNER_LESA_GPU_INDEX (LesaGPUIndex)

This property is relevant if a system has multiple GPUs available. By default the value is set to "-1" and Lesa uses all GPUs that are detected in a system. This property allows to restrict the Lesa OCR to use only a specific GPU specified by index. The property can be set to the 0-based index of the GPU that should be used. E.g. "0" to use the first GPU, "1" to use the second GPU or a comma separated list to use multiple GPUs. E.g. "2,3" to use the third and forth GPU of a system.

⁠CLSDESIGNER_LESA_MAX_OCRPROC_PER_GPU (LesaMaxOCRProcPerGPU)

Defines the number of OCR processes running in parallel for each GPU. The default value is 1. With high end GPUs a higher throughput can be achieved by running 2-4 OCR processes per GPU. In combination with multiple GPUs, the final number of OCR processes will be this value multiplied by number of GPUs. This value significantly affects the OCR performance. If the value is too low, GPU usage is suboptimal. If too high it could lead to GPU memory overload. Please check the GPU usage during processing to find optimal value for your system. A single OCR processes normally needs 4 GByte GPU memory.

⁠OCR Speed and process priority

In some newer versions of Windows operating systems, a compute-intensive task running as a background process is not fully supported by the operating system. When the same task is run as a foreground process, it has been observed that to run almost twice as fast compared to running within a service or IIS application pool. The observed reason for the performance degradation was that the operating system drastically reduced the CPU frequency even during the execution of the OCR tasks. This affects the OCR speed in both CPU and GPU mode. To overcome this performance drop, there is an environment variable that increases the process priority of the OCR process. This option causes the operating system to keep the CPU frequency at the highest level during OCR execution. To increase the OCR process priority set the global environment variable OCR_HIGHER_PRIORITY to '1'. This option is currently not available within the settings file but only as an environment variable.

Please note that this type of configuration carries a certain risk. Running too many OCR processes can overload a system that does not have enough CPU or memory resources. If too many OCR processes are executed with a higher priority, the system may become unstable and unresponsive. Only use this setting if the CPU utilization is at or below 50% when OCR is running at full capacity. Changing CLSDESIGNER_MAXOCRPROCESSCOUNT from the default 1 to something higher comes with a significant higher RAM consumption. Docker has no concept on requested ressources like K8. It will take what your host can offer. However you can limit RAM and CPU via docker parameter (https://docs.docker.com/engine/containers/resource_constraints⁠)

⁠Oracle Connection Modes

For Oracle databases, the basic connection mode and TNS connection mode are supported. In basic connection mode, the database server name is the actual computer that hosts the Oracle database. In TNS connection mode, the connection is configured in a separate file that is usually deployed to each client by the Oracle database administrator. This file is named 'TNSNAME.ORA' and there are different ways how the file gets located. The 'TNSNAME.ORA' is searched within a directory defined by the 'TNS_ADMIN' or 'ORACLE_HOME' environment variable. If a TNS configuration file exists, the value for the database server is evaluated as a network alias defined in that file. Please note that for Oracle the user name and the database name must be identical. Using different values for user name and database name is not fully supported.

The Integrated security flag is currently not supported for the connection to Oracle databases. A database user and password must always be specified.

⁠Kubernetes Support

Helm charts and related information about running skilja containers in Kubernetes can be found on the Partner Portal⁠

Tag summary

Content type

Image

Digest

sha256:e355ed542…

Size

1.2 GB

Last updated

17 days ago

docker pull skilja/classificationdesigner