Sign inSign up

skilja/labertaserver

By skilja

•Updated 18 days ago

A service that hosts LLMs for Laera classification and extraction tasks.

Image
Machine learning & AI
0

1.9K

skilja/labertaserver repository overview

⁠Image renamed

With version 8.0⁠ the product was renamed to "AI Server". There will be no new versions with this image name anymore.

⁠Quick reference

⁠Supported tags and Versioning

Image tags adhere to <major>.<minor>.<patch>.<revision> format.

<major>.<minor>.<patch>.<revision> points to a specific version. <Major>.<Minor> always points to the latest version. This version is compatible with all previous images of the same <Major>.<Minor> version. <latest> always points to the latest version, but such a version might require service and project database schema updates. For loading the latest but compatible version, we recommend to pull a <Major>.<Minor> with

docker pull skilja/labertaserver:<Major>.<Minor>

Note: LaBERTa Server images always contains NVIDIA CUDA drivers. The used models will be executed using NVIDIA CUDA GPU-based hardware acceleration and therefore will be faster (10x vs. CPU). However this requires a dedicated NVIDIA CUDA compatible GPU to run. You also have to run the container using the --gpus all parameter or equivalent with the deploy/devices keyword (https://docs.docker.com/compose/how-tos/gpu-support/⁠)

The latest image is :



logo drawing

⁠What is LaBERTa Server?

The LaBERTa Server provides a Web application for executing large language model requests. You can use the server to perform request directly or to perform request from classification and extraction components.

Classifying or extracting data using large language models requires additional hardware, particularly in a production environment. For encoder-only transformer models (LaBERTa models), a small to medium size GPU is recommended. For decoder-only transformer models (GPT or Llama models), a medium to large size GPU is required.

While it is possible to run those models on a CPU-only machine, it is not recommended. The performance will not be satisfactory

However, when using LaBERTa Server for classification and extraction processes, the server only performs the request. The training of the LoRA layer still has to happen locally on your system.

For direct requests, you can use so-called Prompts, create a summary of the provided text, and allows extracting a couple of fields directly.

⁠How to use this image

The following sample shows how to run LaBERTa Server using PostgreSQL:

1. Optionally for using a PostgreSQL database server, install and start a PostgreSQL container:

docker pull postgres:latest
docker run -v <localpath>:/var/lib/postgresql/data -e POSTGRES_PASSWORD=MyPassword POSTGRES_USER=laberta  -p 5432:5432 -d postgres:latest

2. For the LaBERTa Server, configure the environment file to use the DB user and password from above. ( Also remove comments)

LABERTASERVER_SERVICEENDPOINT: http://+:8047
LABERTASERVER_DBSERVERTYPE: 2     (2 - PostgreSQL)
LABERTASERVER_DBSERVER: <database-server-name> (the host computer name in case of local docker installation)
LABERTASERVER_DBNAME: LabertaServer (if not there, it will be created)
LABERTASERVER_DBINTEGRATEDSECURITY: "False"
LABERTASERVER_DBUSER: laberta
LABERTASERVER_DBPASSWORD: MyPassword
LABERTASERVER_DBUSESSL: "False"
LABERTASERVER_DBTRUSTSERVERCERTIFICATE: "True"
LABERTASERVER_SELFHOSTWEBSITE: "True"
LABERTASERVER_HASAUTHENABLED: "False"
LABERTASERVER_HF_SEARCHLIMIT: 800
LABERTASERVER_LOG_TO_DB: "True"
LABERTASERVER_START_LLMPROC: "True"
LABERTASERVER_LLM_MAXCACHEDMODELS: 8
LABERTASERVER_LLM_MAXPROCESSCOUNT: 1
LABERTASERVER_LLAMA_QUEUETIMEOUT: 120000
LABERTASERVER_LLAMA_MAXCACHEDMODELSGB: 8
LABERTASERVER_LLAMA_TEMPERATURE: 0.01
LABERTASERVER_LLAMA_MAXCONCURRENTUSERS: 10
LABERTASERVER_QUEUE_MAXLENGTH: 100
LABERTASERVER_QUEUE_TIMEOUTMINUTES: 20
LABERTASERVER_QUEUE_PURGEDELETE: "False"
LABERTASERVER_QUEUE_SOFTDELETEMINUTES: 15

3. Pull and start the LaBERTa Server image:

docker pull skilja/labertaserver
docker run -p 8047:8047 -d --env-file envfile.txt skilja/labertaserver

⁠... via [docker-compose]

Example docker-compose.yml for skilja/labertaserver:

# This compose file will start a postgres db and the LaBERTa Server at once 
# make sure you have the local folders created for the mapped dirs

services:
  postgres-db:
    image: postgres
    
    healthcheck:
        test: "pg_isready -U laberta"
        interval: 5s
        timeout: 5s
        retries: 5
    
    environment:
      POSTGRES_PASSWORD: postgres
      POSTGRES_USER: laberta
    ports:
      - "5432:5432"
    volumes:
      - "D:\\DockerFiles\\Laberta\\DB:/var/lib/postgresql/data:z"
    #user: 999:999

  Laberta:
    image: skilja/labertaserver:latest
    depends_on:
       postgres-db:
         condition: service_healthy
    ports:
      - "8047:8047"
    volumes:
        - "D:\\DockerFiles\\Laberta\\Models:/opt/Skilja/LaBERTaServer/Models"
    environment:
      LABERTASERVER_SERVICEENDPOINT: http://+:8047
      LABERTASERVER_DBSERVERTYPE: 2
      LABERTASERVER_DBSERVER: postgres-db
      LABERTASERVER_DBNAME: LabertaServer
      LABERTASERVER_DBINTEGRATEDSECURITY: "False"
      LABERTASERVER_DBUSER: laberta
      LABERTASERVER_DBPASSWORD: postgres
      LABERTASERVER_DBUSESSL: "False"
      LABERTASERVER_DBTRUSTSERVERCERTIFICATE: "True"
      LABERTASERVER_SELFHOSTWEBSITE: "True"
      LABERTASERVER_HASAUTHENABLED: "False"
      LABERTASERVER_HF_SEARCHLIMIT: 800
      LABERTASERVER_LOG_TO_DB: "True"
      LABERTASERVER_START_LLMPROC: "True"
      LABERTASERVER_LLM_MAXCACHEDMODELS: 8
      LABERTASERVER_LLM_MAXPROCESSCOUNT: 1
      LABERTASERVER_LLAMA_QUEUETIMEOUT: 120000
      LABERTASERVER_LLAMA_MAXCACHEDMODELSGB: 8
      LABERTASERVER_LLAMA_TEMPERATURE: 0.01
      LABERTASERVER_LLAMA_MAXCONCURRENTUSERS: 10
      LABERTASERVER_QUEUE_MAXLENGTH: 100
      LABERTASERVER_QUEUE_TIMEOUTMINUTES: 20
      LABERTASERVER_QUEUE_PURGEDELETE: "False"
      LABERTASERVER_QUEUE_SOFTDELETEMINUTES: 15

      NVIDIA_VISIBLE_DEVICES: all

Run docker compose up, wait for it to initialize completely, and visit http://localhost:8047/ or http://host-ip:8047 or http://cotainer-ip:8047

You get a more sophisticated guide and help via the http://localhost:8047/help page.

Warning: The shown sample compose is running without Authorization! It is strongly recommended to use authorization in order to protect the service (licenses) and data that might be stored during processing (documents). Here you can find more information about the Authorization Server⁠. The option to use API Keys will also be enabled once you connected the Authorization Server.

⁠HTTPS

You probably noticed that the described sample is not using a SSL certificate. The currently recommended way is to use a Reverse Proxy like NGIX or Traefik.

⁠Supporting Modules

⁠Environment Variables

The LaBERTa Server image uses several environment variables, of which some are required others are optional.

  • LABERTASERVER_SERVICEENDPOINT (ServiceEndpoint) - URL of the LaBERTa Server.

  • LABERTASERVER_HASAUTHENABLED (HasAuthEnabled) - If this property is true, authentication is enabled for the LaBERTa Server. Users need to log in via the configured authorization server and need to have access permissions with corresponding roles. Custom applications need to provide an API key with each API call. If this property is false, no authentication and no API keys are required.

  • LABERTASERVER_AUTHSERVERURL (AuthServerUrl) - The URL of the authorization server if authentication is enabled.

  • LABERTASERVER_WEBAPPCLIENTID (WebAppClientId) - The client ID of the web application that is registered as a public client within the authorization server.

  • LABERTASERVER_SERVERCLIENTID (ServerClientId) - The client ID of the server side application that is registered as a confidential client within the authorization server.

  • LABERTASERVER_SERVERCLIENTSECRET (ServerClientSecret) - The client secret corresponding to the client ID of the server side application.

  • LABERTASERVER_DBSERVERTYPE (Database.ServerType) - The database server type to connect to as integer for the LaBERTa Server. Supported are SQL Server and PostgreSQL.

    • 0 - SQL Server
    • 2 - PostgreSQL
  • LABERTASERVER_DBSERVER (Database.Server) - Database server name hosting the LaBERTa Server database.

  • LABERTASERVER_DBNAME (Database.Name) - The database name for the LaBERTa Server database.

  • LABERTASERVER_DBUSESSL (Database.UseSSL) - Use SSL encryption on the LaBERTa Server database connection. This is currently supported for MS SQL server and PostgreSQL server.

  • LABERTASERVER_DBTRUSTSERVERCERTIFICATE (Database.TrustCertificate) - If SSL is enabled, by default SSL certificate must be officially signed. Set this parameter to true for self-signed SSL certificates that are not officially trusted.

  • LABERTASERVER_DBINTEGRATEDSECURITY (Database.UseIntegratedSecurity) - When true, use integrated security for the database access to the LaBERTa Server database when false, use SQL user and password.

  • LABERTASERVER_DBUSER (Database.User) - SQL user name if the integrated security option is false.

  • LABERTASERVER_DBPASSWORD (Database.Password) - SQL password if integrated security option is false.

  • LABERTASERVER_QUEUE_MAXLENGTH (QueueSettings.MaxLength) - The maximum about of waiting jobs in the queue for generative AI requests.

  • LABERTASERVER_QUEUE_TIMEOUTMINUTES (QueueSettings.TimeoutMinutes) - Defines a timeout in minutes for jobs waiting in the queue. Expired job are cancelled. Defaults to 20 minutes.

  • LABERTASERVER_QUEUE_PURGEDELETE (QueueSettings.PurgeDelete) - When true, expired or canceled jobs are immediatelly deleted from the database. When false, the jobs are soft deleted only but remain in the database until the deletion time expires.

  • LABERTASERVER_QUEUE_SOFTDELETEMINUTES (QueueSettings.SoftDeleteMinutes) - Defines after how many minutes soft deleted jobs are finally deleted from the database.

  • LABERTASERVER_LOG_TO_DB (LogToDB) - When true writes log messages to the database.

  • LABERTASERVER_MODELPATH (ModelPath) - Directory where the models are stored. When empty defaults to %APPDATA%\Skilja\LaBERTaServer\Models on Windows and opt\Skilja\LaBERTaServer\Models on Linux systems.

  • LABERTASERVER_HF_SEARCHLIMIT (SearchLimit) - Defines a limit of models when searching for generative models on a remote server.

  • LABERTASERVER_START_LLMPROC (StartLLMProcess) - When true warm up the server for faster execution of BERT models.

  • LABERTASERVER_LLM_MAXPROCESSCOUNT (MaxLLMProcessCount) - Maximum number of LLM Processes that can execute BERT models in parallel.

  • LABERTASERVER_LLM_MAXCACHEDMODELS (MaxLLMCachedModels) - Maximum number of cached BERT models.

  • LABERTASERVER_LLAMA_MAXCONCURRENTUSERS (MaxConcurrentUsers) - Maximum concurrent users for animating results from the LLama Process.

  • LABERTASERVER_LLAMA_QUEUETIMEOUT (LLamaQueueTimeout) - Maxmimum time in milliseconds for processing a generative AI requests.

  • LABERTASERVER_LLAMA_MAXCACHEDMODELSGB (LLamaMaxCachedModelsGB) - Maximum size in GB of cached generative AI models. Set to 0 to dynamically calculate based on GPU VRAM size.

  • LABERTASERVER_LLAMA_TEMPERATURE - Default temperature for processing generative AI requests.

  • LABERTASERVER_LLAMA_MAXPROCESSCOUNT - Maximum number of LLama Processes that can execute generative AI models in parallel. Processing is limited to 1 LLama Process. Set 0 to disable processing generative AI models.

  • LABERTASERVER_STATISTICS_MAX_RETENTION_HOURS (StatisticsMaxRetentionHours) - Maximum hours that statistics are stored in the database.

  • LABERTASERVER_REMOTEAIAPI_PROCESSINGTIMEOUT (RemoteAiApi.ProcessingTimeout) - Maximum time in milliseconds for processing a remote AI request.

  • LABERTASERVER_REMOTEAIAPI_MAXPARALLELREQUESTS (RemoteAiApi.MaxParallelRequests) - Maximum requests that can be processed in parallel with a remote AI. Set 0 to disable processing with remote AI.

  • LABERTASERVER_DATAPROTECTION_PASSWORD (DataProtection.Password) - The password to protect certificates for encryption and decryption of API keys for remote generative AI models.

  • LABERTASERVER_LLAMA_USE_MMAP - By default, the GPU memory for generative LLMs is initialized using an memory map (mmap) function call. This allows a fast loading for the models into the GPU. The disadvantage of this function call is a very short but high memory requirement in system RAM, where the same amount of system RAM like the GPU memory is required. In case of very large GPU memory usage, e.g. 80 GB, also 80 GB of RAM are only needed when loading the model into GPU. Settings this environment variable to false reduces the RAM requirements by at least 50% during LLM loading time while making the model loading a bit slower.



⁠Other Platforms

The service is also available for Windows installation. Please visit The Partner Portal⁠ or write us a email [email protected]⁠⁠ for more information.

Tag summary

Content type

Image

Digest

sha256:3bfc99bc6…

Size

7.9 GB

Last updated

18 days ago

docker pull skilja/labertaserver