Sign inSign up

skilja/aiserver

By skilja

•Updated 18 days ago

A service that hosts LLMs for Laera classification and extraction tasks.

Image
Machine learning & AI
0

808

skilja/aiserver repository overview

⁠Quick reference

⁠Supported tags and Versioning

Image tags adhere to <major>.<minor>.<patch>.<revision> format.

<major>.<minor>.<patch>.<revision> points to a specific version. <Major>.<Minor> always points to the latest version. This version is compatible with all previous images of the same <Major>.<Minor> version. <latest> always points to the latest version, but such a version might require service and project database schema updates. For loading the latest but compatible version, we recommend to pull a <Major>.<Minor> with

docker pull skilja/aiserver:<Major>.<Minor>

Note: AI Server images always contains NVIDIA CUDA drivers. The used models will be executed using NVIDIA CUDA GPU-based hardware acceleration and therefore will be faster (10x vs. CPU). However this requires a dedicated NVIDIA CUDA compatible GPU to run. You also have to run the container using the --gpus all parameter or equivalent with the deploy/devices keyword (https://docs.docker.com/compose/how-tos/gpu-support/⁠)

⁠Supported architectures

AI Server images are published as multi-arch manifests since version 8.0.1 and support the following platforms:

  • linux/amd64 (NVIDIA CUDA GPU acceleration possible, requires --gpus all)
  • linux/arm64 (NVIDIA CUDA GPU acceleration possible, requires --gpus all)



The latest image is :

Warning : Breaking Changes with 8.0

(More Information be found within the provided release notes)



logo drawing

⁠What is the AI Server?

The AI Server provides a Web application for executing large language model requests. You can use the server to perform request directly or to perform request from classification and extraction components.

Classifying or extracting data using large language models requires additional hardware, particularly in a production environment. For encoder-only transformer models (LaBERTa models), a small to medium size GPU is recommended. For decoder-only transformer models (GPT or Llama models), a medium to large size GPU is required.

While it is possible to run those models on a CPU-only machine, it is not recommended. The performance will not be satisfactory

However, when using AI Server for classification and extraction processes, the server only performs the request. The training of the LoRA layer still has to happen locally on your system.

For direct requests, you can use so-called Prompts, create a summary of the provided text, and allows extracting a couple of fields directly.

⁠How to use this image

The following sample shows how to run AI Server using PostgreSQL:

1. Optionally for using a PostgreSQL database server, install and start a PostgreSQL container:

docker pull postgres:latest
docker run -v <localpath>:/var/lib/postgresql/data -e POSTGRES_PASSWORD=MyPassword POSTGRES_USER=aiserver  -p 5432:5432 -d postgres:latest

2. For the AI Server, configure the environment file to use the DB user and password from above. ( Also remove comments)

BERNINA_SERVICEENDPOINT: http://+:8047
BERNINA_DBSERVERTYPE: 2     (2 - PostgreSQL)
BERNINA_DBSERVER: <database-server-name> (the host computer name in case of local docker installation)
BERNINA_DBNAME: AiServer (if not there, it will be created)
BERNINA_DBINTEGRATEDSECURITY: "False"
BERNINA_DBUSER: aiserver
BERNINA_DBPASSWORD: MyPassword
BERNINA_DBUSESSL: "False"
BERNINA_DBTRUSTSERVERCERTIFICATE: "True"
BERNINA_SELFHOSTWEBSITE: "True"
BERNINA_PATHBASE: "/aiserver"
BERNINA_HASAUTHENABLED: "False"
BERNINA_HF_SEARCHLIMIT: 800
BERNINA_LOG_TO_DB: "True"
BERNINA_START_LLMPROC: "True"
BERNINA_LLM_MAXCACHEDMODELS: 8
BERNINA_LLM_MAXPROCESSCOUNT: 1
BERNINA_LLAMA_QUEUETIMEOUT: 120000
BERNINA_LLAMA_MAXCACHEDMODELSGB: 8
BERNINA_LLAMA_TEMPERATURE: 0.01
BERNINA_LLAMA_MAXCONCURRENTUSERS: 10
BERNINA_QUEUE_MAXLENGTH: 100
BERNINA_QUEUE_TIMEOUTMINUTES: 20
BERNINA_QUEUE_PURGEDELETE: "False"
BERNINA_QUEUE_SOFTDELETEMINUTES: 15

3. Pull and start the AI Server image:

docker pull skilja/aiserver
docker run -p 8047:8047 -d --env-file envfile.txt skilja/aiserver

⁠... via [docker-compose]

Example docker-compose.yml for skilja/aiserver:

# This compose file will start a postgres db and the AI Server at once 
# make sure you have the local folders created for the mapped dirs

services:
  postgres-db:
    image: postgres
    
    healthcheck:
        test: "pg_isready -U aiserver"
        interval: 5s
        timeout: 5s
        retries: 5
    
    environment:
      POSTGRES_PASSWORD: postgres
      POSTGRES_USER: aiserver
    ports:
      - "5432:5432"
    volumes:
      - "D:\\DockerFiles\\AiServer\\DB:/var/lib/postgresql/data:z"
    #user: 999:999

  aiserver:
    image: skilja/aiserver:latest
    depends_on:
       postgres-db:
         condition: service_healthy
    ports:
      - "8047:8047"
    volumes:
        - "D:\\DockerFiles\\AIServer\\Models:/opt/Skilja/AIServer/Models"
    environment:
      BERNINA_SERVICEENDPOINT: http://+:8047
      BERNINA_DBSERVERTYPE: 2
      BERNINA_DBSERVER: postgres-db
      BERNINA_DBNAME: AiServer
      BERNINA_DBINTEGRATEDSECURITY: "False"
      BERNINA_DBUSER: aiserver
      BERNINA_DBPASSWORD: postgres
      BERNINA_DBUSESSL: "False"
      BERNINA_DBTRUSTSERVERCERTIFICATE: "True"
      BERNINA_SELFHOSTWEBSITE: "True"
      BERNINA_PATHBASE: "/aiserver"
      BERNINA_HASAUTHENABLED: "False"
      BERNINA_HF_SEARCHLIMIT: 800
      BERNINA_LOG_TO_DB: "True"
      BERNINA_START_LLMPROC: "True"
      BERNINA_LLM_MAXCACHEDMODELS: 8
      BERNINA_LLM_MAXPROCESSCOUNT: 1
      BERNINA_LLAMA_QUEUETIMEOUT: 120000
      BERNINA_LLAMA_MAXCACHEDMODELSGB: 8
      BERNINA_LLAMA_TEMPERATURE: 0.01
      BERNINA_LLAMA_MAXCONCURRENTUSERS: 10
      BERNINA_QUEUE_MAXLENGTH: 100
      BERNINA_QUEUE_TIMEOUTMINUTES: 20
      BERNINA_QUEUE_PURGEDELETE: "False"
      BERNINA_QUEUE_SOFTDELETEMINUTES: 15

      
    # Uncomment to enable GPU access (requires NVIDIA Container Toolkit installed on dockerhost)
    # deploy:
    #   resources:
    #     reservations:
    #       devices:
    #         - driver: nvidia
    #           count: 1
    #           capabilities: [gpu]

Run docker compose up, wait for it to initialize completely, and visit http://localhost:8047/ or http://host-ip:8047 or http://cotainer-ip:8047

You get a more sophisticated guide and help via the http://localhost:8047/help page.

Warning: The shown sample compose is running without Authorization! It is strongly recommended to use authorization in order to protect the service (licenses) and data that might be stored during processing (documents). Here you can find more information about the Authorization Server⁠. The option to use API Keys will also be enabled once you connected the Authorization Server.

⁠Configuring the Authorization Server Client

When you setup AI Server with the Authorization Server, you will need to also configure AI Server's Client in the Authorization Server. This is not done automatically. From 8.0 the AS client config changed. If you already have clients inplace, delete them and create the only needed one.

1. Log in to your Authorization Server with administrative rights.
2. Open the **Management** tab.
3. Create a new client application with the following settings:
   - **Client ID:** `BERNINA_Instance1` (example)
   - **Client Secret:** your client secret
   - **Authorization mode:** Authorization Code only
   - **Introspection: No
   - **Additional permissions:** None
   - **Scopes:** Roles and Profile
4. Add the redirect URI for your configured AI Server URL, for example:
   `https://aiserverdev.com/aiserver/signin-oidc`
5. Add the post-logout redirect URI for your configured AI Server URL, for example:
   `https://vinnadev.local/aiserver/signout-callback-oidc`
6. Add the Client ID to the EnvVar BERNINA_SERVERCLIENTID
7. Add the Client Secret to the EnvVar BERNINA_SERVERCLIENTSECRET

> Note: The redirect URIs are case-sensitive.

⁠HTTPS

You probably noticed that the described sample is not using a SSL certificate. The currently recommended way is to use a Reverse Proxy like NGIX or Traefik.

⁠Supporting Modules

⁠Service User

Beginning with version 8.0.0⁠ the used service user is changed to be uniform across all Skilja containers. For .Net based containers like this one, the UID/GID 1654 is now used. This is a non root user called 'app'. This becomes important when working with Volume Mounts.

⁠Renaming the image

With version 8.0.0⁠ the product was renamed to "AI Server". The Environment Variables changed from "LABERTASERVER_" to "BERNINA_" For compability reasons, all previously used "LABERTASERVER_*" ENV Vars ( like LABERTASERVER_SERVICEENDPOINT) will be still valid. When renaming an existing deployment please check also path base, model mounts and redirect URIs of Auth Client(s).

⁠Site Mangement Tool

When the AI Server is started for the first time, users do not have any permissions assigned by default.
("You need read permission to view this page", indicates this state)

To grant administrative access, complete the following steps:

  1. Access the running container using a tool such as docker exec or kubectl exec.

  2. Run the Site Management tool:

    dotnet Bernina.SiteManagement.dll
    
  3. Follow the interactive prompts and assign your user account the Site Administrator role.

  4. Restart the container to apply the changes.

After the restart, additional role and permission changes can be managed through the Web UI.

⁠Environment Variables

The AI Server image uses several environment variables, of which some are required others are optional.

  • BERNINA_SERVICEENDPOINT (ServiceEndpoint) - URL of the AI Server.

  • BERNINA_HASAUTHENABLED (HasAuthEnabled) - If this property is true, authentication is enabled for the AI Server. Users need to log in via the configured authorization server and need to have access permissions with corresponding roles. Custom applications need to provide an API key with each API call. If this property is false, no authentication and no API keys are required.

  • BERNINA_AUTHSERVERURL (AuthServerUrl) - The URL of the authorization server if authentication is enabled.

  • BERNINA_SERVERCLIENTID (ServerClientId) - The client ID of the server side application that is registered as a confidential client within the authorization server.

  • BERNINA_SERVERCLIENTSECRET (ServerClientSecret) - The client secret corresponding to the client ID of the server side application.

  • BERNINA_DBSERVERTYPE (Database.ServerType) - The database server type to connect to as integer for the AI Server. Supported are SQL Server and PostgreSQL.

    • 0 - SQL Server
    • 2 - PostgreSQL
  • BERNINA_DBSERVER (Database.Server) - Database server name hosting the AI Server database.

  • BERNINA_DBNAME (Database.Name) - The database name for the AI Server database.

  • BERNINA_DBUSESSL (Database.UseSSL) - Use SSL encryption on the AI Server database connection. This is currently supported for MS SQL server and PostgreSQL server.

  • BERNINA_DBTRUSTSERVERCERTIFICATE (Database.TrustCertificate) - If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to true for self-signed SSL certificates that are not officially trusted.

  • BERNINA_DBINTEGRATEDSECURITY (Database.UseIntegratedSecurity) - When true, use integrated security for the database access to the AI Server database when false, use SQL user and password.

  • BERNINA_DBUSER (Database.User) - SQL user name if the integrated security option is false.

  • BERNINA_DBPASSWORD (Database.Password) - SQL password if integrated security option is false.

  • BERNINA_QUEUE_MAXLENGTH (QueueSettings.MaxLength) - The maximum about of waiting jobs in the queue for generative AI requests.

  • BERNINA_QUEUE_TIMEOUTMINUTES (QueueSettings.TimeoutMinutes) - Defines a timeout in minutes for jobs waiting in the queue. Expired job are cancelled. Defaults to 20 minutes.

  • BERNINA_QUEUE_PURGEDELETE (QueueSettings.PurgeDelete) - When true, expired or canceled jobs are immediatelly deleted from the database. When false, the jobs are soft deleted only but remain in the database until the deletion time expires.

  • BERNINA_QUEUE_SOFTDELETEMINUTES (QueueSettings.SoftDeleteMinutes) - Defines after how many minutes soft deleted jobs are finally deleted from the database.

  • BERNINA_LOG_TO_DB (LogToDB) - When true writes log messages to the database.

  • BERNINA_MODELPATH (ModelPath) - Directory where the models are stored. When empty defaults to %APPDATA%\Skilja\AIServer\Models on Windows and opt/skilja/aiserver/models on Linux systems.

  • BERNINA_HF_SEARCHLIMIT (SearchLimit) - Defines a limit of models when searching for generative models on a remote server.

  • BERNINA_START_LLMPROC (StartLLMProcess) - When true warm up the server for faster execution of BERT models.

  • BERNINA_LLM_MAXPROCESSCOUNT (MaxLLMProcessCount) - Maximum number of LLM Processes that can execute BERT models in parallel. Adjust to 0 for Gateway / LB mode. More info in docu.

  • BERNINA_LLM_MAXCACHEDMODELS (MaxLLMCachedModels) - Maximum number of cached BERT models.

  • BERNINA_LLAMA_MAXCONCURRENTUSERS (MaxConcurrentUsers) - Maximum concurrent users for animating results from the LLama Process.

  • BERNINA_LLAMA_QUEUETIMEOUT (LLamaQueueTimeout) - Maxmimum time in milliseconds for processing a generative AI requests.

  • BERNINA_LLAMA_MAXCACHEDMODELSGB (LLamaMaxCachedModelsGB) - Maximum size in GB of cached generative AI models. Set to 0 to dynamically calculate based on GPU VRAM size.

  • BERNINA_LLAMA_TEMPERATURE - Default temperature for processing generative AI requests.

  • BERNINA_LLAMA_MAXPROCESSCOUNT - Maximum number of LLama Processes that can execute generative AI models in parallel. Processing is limited to 1 LLama Process. Set 0 to disable processing generative AI models. Adjust to 0 for Gateway / LB mode. More info in docu.

  • BERNINA_STATISTICS_MAX_RETENTION_HOURS (StatisticsMaxRetentionHours) - Maximum hours that statistics are stored in the database.

  • BERNINA_REMOTEAIAPI_PROCESSINGTIMEOUT (RemoteAiApi.ProcessingTimeout) - Maximum time in milliseconds for processing a remote AI request.

  • BERNINA_REMOTEAIAPI_MAXPARALLELREQUESTS (RemoteAiApi.MaxParallelRequests) - Maximum requests that can be processed in parallel with a remote AI. Set 0 to disable processing with remote AI.

  • BERNINA_DATAPROTECTION_PASSWORD (DataProtection.Password) - The password to protect certificates for encryption and decryption of API keys for remote generative AI models.

  • BERNINA_LLAMA_USE_MMAP - By default, the GPU memory for generative LLMs is initialized using an memory map (mmap) function call. This allows a fast loading for the models into the GPU. The disadvantage of this function call is a very short but high memory requirement in system RAM, where the same amount of system RAM like the GPU memory is required. In case of very large GPU memory usage, e.g. 80 GB, also 80 GB of RAM are only needed when loading the model into GPU. Settings this environment variable to false reduces the RAM requirements by at least 50% during LLM loading time while making the model loading a bit slower.

  • BERNINA_INSTANCENAME (InstanceName) - The name of the instance reporting service metrics via OpenTelemetry. If empty, the current computer name is used. For Docker instances with random names, it is recommended to set this manually.

  • BERNINA_OTLP_ENDPOINT (OpenTelemetry.Endpoint) - Specify the OpenTelemetry endpoint to enable metrics reporting. (Sample: http://otel-opentelemetry-collector.monitoring.svc.cluster.local:4317)

  • BERNINA_OTLP_ASPNETCORE_METRICS_ENABLED (OpenTelemetry.AspNetCoreMetricsEnabled) - Enable ASP.NET Core-specific metrics for Hosting, Routing, Kestrel Server, and SignalR. (Default is disabled).

  • BERNINA_OTLP_NETRUNTIME_METRICS_ENABLED (OpenTelemetry.NetRuntimeMetricsEnabled) - Enable .NET Runtime metrics for .NET runtime libraries. (Default is disabled).

  • BERNINA_OTLP_RESOURCEMONITORING_METRICS_ENABLED (OpenTelemetry.ResourceMonitoringMetricsEnabled) - Enable resource monitoring metrics, such as process or container CPU and memory utilization. (Default is disabled).

  • BERNINA_OTLP_REPORTING_INTERVAL_SECONDS (OpenTelemetry.ReportingIntervalSeconds) - Set the interval in seconds to regularly report service-specific metrics for generative AI request waiting times. Set to 0 to disable regular reporting.



⁠Other Platforms

The service is also available for Windows installation. Please visit The Partner Portal⁠ or write us a email [email protected]⁠⁠ for more information.

A beta version for ARM support is also available for this image.

Tag summary

Content type

Image

Digest

sha256:3bfc99bc6…

Size

7.9 GB

Last updated

18 days ago

docker pull skilja/aiserver