CloverDX Server designed for GPU-accelerated machine learning workloads.
1.2K
CloverDX Data Integration Platform covers the entire lifecycle of even the most complex data pipelines. It is a complete solution for automating and managing complex data pipelines, covering development lifecycle as well as reliable operations.
Focused heavily on automation, it helps boost productivity and trust in data and the process behind it.
CloverDX empowers teams across your entire organization with data, providing access to insights that can help drive your business forward. With CloverDX, you can simplify your data integration and management processes, giving you more time to focus on what really matters - growing your business.
CloverDX operates with variety of data sources, including on-premise relational databases (Oracle, MS SQL Server, Maria, PostgreSQL and many others), NoSQL databases, cloud storage (AWS S3, Snowflake, Redshift, AWS Aurora, etc.), applications (Salesforce, HubSpot, Xero, etc.) and REST/SOAP generic APIs.
This Docker container provides an easy way to deploy a CloverDX Server instance configured specifically to allow you to run GPU-accelerated machine learning models. The container spins-up a standalone CloverDX Server with basic settings for production use and may require additional settings depending on your exact use case. The container requires a machine with at least one NVIDIA GPU to work properly. As an example, you can use AWS EC2 G4dn or AWS EC2 G5 instances.
For the source of this image and additional documentation see Dockerfile.AI file in https://github.com/cloverdx/cloverdx-server-docker.
CloverDX Server requires persistent storage and a relational database. The storage is used for your sandboxes – CloverDX projects on the Server – and is needed to ensure that your data is not lost when container is stopped or upgraded. The database stores configuration and logs of the processes that you run on your CloverDX Server instance. To start a full CloverDX Server with external storage and external database, you can use the following command:
docker run -d \
--name cloverdx \
--gpus 1 \
--memory=14g \
-p 8080:8080 \
-e "CLOVER_SERVER_HEAP_SIZE=1792" \
-e "CLOVER_WORKER_HEAP_SIZE=3584" \
-e clover.jdbc.driverClassName=org.postgresql.Driver \
-e clover.jdbc.url="jdbc:postgresql://hostname:5432/clover?charSet=UTF-8" \
-e clover.jdbc.username=user \
-e clover.jdbc.password=pass \
-e clover.jdbc.dialect=org.hibernate.dialect.PostgreSQLDialect \
--mount type=bind,source=/data/your-host-clover-home-dir,target=/var/clover \
cloverdx/cloverdx-server-ai:latest
The above will start CloverDX container with:
/data/your-host-clover-home-dir directory,The following options are used in the command-line above:
--name: name to identify the running container.--gpus: enables the container to access GPUs.--memory: limit the maximum amount of memory the container can use. You should set this to at least 8 GB. Note that AI workloads can be memory intensive. To avoid out-of-memory errors or performance issues, it's recommended to manually set the JVM maximum heap size for both the CloverDX Server Core and Worker processes.-p: maps ports from the container to the outside. In the example above, we are mapping an “in-container” port 8080 to outside port 8080. This port is the port on which the Server exposes its management web app.-e CLOVER_SERVER_HEAP_SIZE and -e CLOVER_WORKER_HEAP_SIZE: override the automatic settings of max heap memory for CloverDX Server Core and Worker processes - the optimal values for these settings depend on your specific use case - including the complexity of executed jobs, data volume, and the AI models being used.-e: multiple parameters that provide settings for the external database. The settings above configure the Server to use PostgreSQL instance as its system database. PostgreSQL is the recommended database but you can also use other databases if needed. The image contains drivers for three more databases you can use –MySQL, Oracle and Microsoft SQL Server. Please consult our documentation to learn more about supported database versions and configuration for each database. Another way of configuring the database connection is via the clover.properties file.--mount: specifies the binding between directories within the container and directories outside of the container. When running CloverDX Server, you should bind a host directory to /var/clover/ inside the container as a mounted volume.CloverDX Server Console (the management interface for CloverDX Server) will now be available at http://localhost:8080/clover. The Server will require additional configuration – user accounts, permissions, sandboxes and more. All this configuration can be changed in CloverDX Server Console. Please consult our online documentation for additional details. Note that once you start the container, you will need to provide CloverDX Server license to activate it. By default, the container will search for a license file in the conf/license.dat path. The license file is a text file that you can download from CloverDX Customer Portal.
Alternatively, you can activate the Server by:
license.file configuration property and set a different path to the license file, e.g. to a different volume.On Linux, if you bind a directory from the host OS, the data files will be owned by a user with UID 1000. You can override this by setting LOCAL_USER_ID environment variable by adding one more option to the command-line above:
-e LOCAL_USER_ID=`id -u $USER`
This will use the current user (the user starting the container in the host OS) as the user for the files created by the Server.
It is also possible to start CloverDX Server with minimal configuration that is suitable for example for evaluation use or a s smaller testing instance. This can be done by starting the Server without persistent storage and external database. To run a Server like this, you can use the following command:
docker run -d --name cloverdx --gpus 1 --memory=8g -p 8080:8080 cloverdx/cloverdx-server-ai:latest
When started like this, the Server will allow you to use the full functionality but will be started in a configuration that is suitable for evaluation or non-production use. It will use embedded Apache Derby database as system database and will not have any persistent storage attached to it.
Content type
Image
Digest
sha256:5964f6648…
Size
4.7 GB
Last updated
6 days ago
docker pull cloverdx/cloverdx-server-ai