This project is provides a deployable service for pulling SHACL records out of mobi with which to validate data from various types of semantic data stores. The resulting SHACL validation report data can be configured to be written to a number of data sinks, including mobi datasets.

The high-level CSDV architecture is to provide a containerized worker that can load SHACL shapes from a source or record, then validate sets of data against it. The resulting SHACL Validation Reports will then be written out to a reporting system (another triple store or other data sink).
The image itself provides a number of customization options with which to configure running containers (detailed below). The concept behind this image is a reusable compute engine for performing SHACL validation outside of your data repository to offload sometimes expensive operations.
This application exposes REST APIs for interacting with the CSDV system in order to validate data against SHACL shapes. It hosts a resource-efficient, reactive vertx.io services that wrap an RDF4j framework for evaluating SHACL against input data.
To trigger the systen, you can use Postman to POST at http://localhost:8888/ (port can be overridden) with JSON in the body of the request.
The structure of the JSON in the request to trigger CSDV is described in this section. Overall it can be thought of as a JSON document with 5 parts: SHACL, Data Sources, Reporters, Connections, and Engine Configuration.
A SHACL, Data Source, and Reporter configuration are all described as a Process Configuration (sharing the same base schema).
| Key | Description | Notes |
|---|---|---|
| type | The type of process configuration. Different types of process configurations have different available types | This is a required field to tell CSDV what type of thing to create. |
| source | Each configuration will link to a Source Configuration, which tells it how to connect to a remote system. | This is a required field to tell CSDV how to connect to a remote system to run the specified process. |
| config | This is a free-form map that will vary based upon the type of process being configured. | There are examples of values in this field in the sample JSON request, and in documentation further down below. |
This section of the request tells the CSDV system where to pull SHACL Shapes Records from to validate data from requested data sources. Each instance of this class in the shacl array are Process Configurations.
Note that you can specify multiple SHACL records to pull and instrument your validation against.
Currently there is only one supported type, mobi, which requires a shape_record in the config map in order to tell CSDV which SHACL record to pull and validate against. The source key should reference a specified Connection Configuration (described below) in order to describe how to connect and authenticate against the running Mobi system.
This section of the JSON request document describes where to pull data from to validate against configured SHACL shapes. Each data source entry specifies a set of data that should be validated against the SHACL configured in the previous section. You can specify multiple sources of data to validate.
The following tables detail the custom configuration each supported type of Data Source supports. Each Data Source entry in the JSON is of the Process Configuration type, so it should have a type and source key, then a config key with the configuration items specified below.
StarDog is an enterprise knowledge platform and RDF graph database that allows for sophisticated virtualization. CSDV can integrate with it to pull data to be validated against SHACL shapes.
| Key | Description | Notes |
|---|---|---|
| database | The name of the database within your StarDog database to connect to | Required; needs to exist on the StarDog side before execution. |
| types | An array of IRI strings (without the '<>' characters of a type of class to pull from the data source. | Not required, defaults to pulling raw statements. |
| limit | The number of instances to limit the pull against (if types isn't specified, it limits raw statements). | Not required, defaults to ALL. |
Mobi is a graph data platform, primarily focused on ontology and vocabulary reference data development.
| Key | Description | Notes |
|---|---|---|
| datasetRecord | The IRI of the dataset record to pull from Mobi. | Required, dataset record must exist in target Mobi system. |
| types | An array of IRI strings (without the '<>' characters of a type of class to pull from the data source. | Not required, defaults to pulling raw statements. |
| limit | The number of instances to limit the pull against (if types isn't specified, it limits raw statements). | Not required, defaults to ALL. |
Anzo is a comprehensive knowledge graph platform that can be leveraged as a source of data to validate against SHACL shapes.
| Key | Description | Notes |
|---|---|---|
| graphmart | The graphmart to query against within the Anzo system | Required, must exist on the Anzo system when run. |
| types | An array of IRI strings (without the '<>' characters of a type of class to pull from the data source. | Not required, defaults to pulling raw statements. |
| limit | The number of instances to limit the pull against (if types isn't specified, it limits raw statements). | Not required, defaults to ALL. |
GraphDB is a RDF graph database, built by OntoText, that can be leveraged as a source of data to validate against your SHACL shapes.
| Key | Description | Notes |
|---|---|---|
| repository | The repository to query against on the GraphDB system. | Required, must already exist on the GraphDB system. |
| types | An array of IRI strings (without the '<>' characters of a type of class to pull from the data source. | Not required, defaults to pulling raw statements. |
| limit | The number of instances to limit the pull against (if types isn't specified, it limits raw statements). | Not required, defaults to ALL. |
A reporter in this context is configuration that tells the CSDV process what to do with the SHACL validation report it generates. Reporters can insert the resulting validation report into a remote system based upon configuration here. Each reporter is a Process Configuration, as specified above, and should describe the type and source, as well as custom type configuration.
The following tables detail the custom type configuration available for each supported type of reporter.
SHACL validation reports can be inserted into a StarDog database.
| Key | Description | Notes |
|---|---|---|
| database | The database to insert the validation report into. | Required, must already exist on the StarDog system. |
| overwrite | Whether or not to clear out the location before writing the new report. | Not required, defaults to false. |
The CSDV system can write validation reports into existing layers of a GraphMart within an Anzo system. One thing to note is that GraphMarts are typically transient, meaning data will be lost when the GraphMart reloads unless otherwise configured.
| Key | Description | Notes |
|---|---|---|
| graphmart | The GraphMart IRI to connect and insert data into. | Required, the GraphMart must already exist on the remote Anzo. Do not include the '<>' characters on the IRI. |
| layer | The IRI of the layer to insert the validation report into. | Required, the layer must already exist on the Graphmart. Do not include the '<>' characters on the IRI. |
| overwrite | Whether or not to clear out the location before writing the new report. | Not required, defaults to false. |
The CSDV system can insert SHACL validation reports into a GraphDB database as well.
| Key | Description | Notes |
|---|---|---|
| repository | The name of the GraphDB repository to connect to. | Required, must exist on the GraphDB system. |
| graph | The named graph in which to insert the validation report. | Required, does not have to exist prior to use. |
| overwrite | Whether or not to clear out the location before writing the new report. | Not required, defaults to false. |
SHACL validation reports can also be inserted into a dataset record in Mobi.
| Key | Description | Notes |
|---|---|---|
| dataset | The IRI of the dataset record to insert the validation report into. | Required, must already exist on the Mobi system. Do not include the '<>' part of the IRI. |
| overwrite | Whether or not to clear out the location before writing the new report. | Not required, defaults to false. |
The Connections section of the request JSON specifies the various repositories that CSDV will need to know how to connect to and operate with. These are referenced by various Process Configurations using the source attribute that should overlap with the name attribute as described here.
A connection should follow a structure similar to:
{
"name": "mobienterprise",
"host": "https://mobienterprise.inovexcorp.com",
"auth": {
"type": "basic",
"user": "admin",
"password": "admin"
}
}
Where the:
The final portion of the configuration is slightly different in that is not an array of documents, but a single document that describes how to instrument the process performing the actual validation.
| Key | Description | Notes |
|---|---|---|
| parallelValidation | Whether or not to validate transactions in parallel. | Not required, default false. |
| validationResultsLimit | Limit the number of validation results in the output report. Can really speed up processing. | Not required, default 0 (meaning unlimited). |
| validationResultsLimitPerConstraint | Limit the number of validation results per constraint in the SHACL. Can really speed up processing. | Not required, default 0 (meaning unlimited). |
| storeType | Type type of RDF4j store. (native, memory, and lmdb are supported). | Required. Read about the RDF4j store types for more information. |
| index | The types of indexes to instrument the RDF4j store with. ("spoc,ospc,psoc" is suggested). | Required unless using memory store type. Index manipulation can result in performance vs memory trade-offs. |
| performanceLogging | Whether or not to produce additional logging around performance. | Not required, default false. Can have a logging/performance impact itself. |
| transactionalValidationLimit | The size at which a transaction in RDF4j needs to be validated in bulk. | Not required, default -1 (meaning always bulk). |
| rdf4jShaclExtensions | Whether or not to include RDF4j specific extensions to SHACL in the processing. | Not required, default false. |
| isolationLevel | The isolation level to process data in while validating. | Not required, default "READ_COMMITTED". Beware of this unless you know what you're doing in RDF4j. |
| isolateDataSources | Whether or not to produce a separate validation report per Data Source (true), or to cram them all into a single report (false. | Not required, default false (meaning cram violations into a single report). |
The following is a sample JSON payload that could be sent to CSDV to hook into several different systems to validate data against a SHACL shape against basic SKOS Concepts.
{
"shacl": [
{
"type": "mobi",
"source": "mobienterprise",
"config": {
"shape_record": "https://mobi.com/records#05e033a3-8ef9-49cb-b847-987f1c1d234f"
}
}
],
"dataSources": [
{
"type": "stardog",
"source": "local_stardog",
"config": {
"database": "local",
"limit": 1000,
"types": [
"http://www.w3.org/2004/02/skos/core#Concept",
"http://www.w3.org/2004/02/skos/core#ConceptScheme"
]
}
},
{
"type": "mobi",
"source": "mobienterprise",
"config": {
"datasetRecord": "https://mobi.com/records#95fa2559-2bf1-486d-a274-729211bff7b2",
"types": [
"http://www.w3.org/2004/02/skos/core#Concept",
"http://www.w3.org/2004/02/skos/core#ConceptScheme"
],
"limit": 1000
}
},
{
"type": "anzo",
"source": "anzodemo",
"config": {
"graphmart": "http://openanzo.org/Graphmart/89523810-6d55-45e6-881f-33ef44810e8d",
"limit": 1000,
"types": [
"http://www.w3.org/2004/02/skos/core#Concept",
"http://www.w3.org/2004/02/skos/core#ConceptScheme"
]
}
},
{
"type": "graphdb",
"source": "local_graphdb",
"config": {
"repository": "local",
"limit": 100,
"types": [
"http://www.w3.org/2004/02/skos/core#Concept",
"http://www.w3.org/2004/02/skos/core#ConceptScheme"
]
}
}
],
"reporters": [
{
"type": "mobi",
"source": "mobienterprise",
"config": {
"dataset": "https://mobi.com/records#404b513a-820f-4343-88cd-54027b7167a5",
"overwrite": true
}
},
{
"type": "anzo",
"source": "anzodemo",
"config": {
"graphmart": "http://openanzo.org/Graphmart/89523810-6d55-45e6-881f-33ef44810e8d",
"layer": "http://inovexcorp.com/Layer/c2374a8fff974c9e82d39d560d1e0aa4",
"overwrite": true
}
},
{
"type": "graphdb",
"source": "local_graphdb",
"config": {
"graph": "urn://report",
"repository": "local",
"overwrite": true
}
}
],
"connections": [
{
"name": "mobienterprise",
"host": "https://mobienterprise.inovexcorp.com",
"auth": {
"type": "basic",
"user": "username",
"password": "password"
}
},
{
"name": "anzodemo",
"host": "https://anzodemo.inovexcorp.com",
"auth": {
"type": "basic",
"user": "username",
"password": "password"
}
},
{
"name": "local_stardog",
"host": "http://10.5.1.2:5820",
"auth": {
"type": "basic",
"user": "username",
"password": "password"
}
},
{
"name": "local_graphdb",
"host": "http://127.0.0.1:7200",
"auth": {
"type": "basic",
"user": "username",
"password": "password"
}
}
],
"engine": {
"parallelValidation": true,
"validationResultsLimit": -1,
"validationResultsLimitPerConstraint": -1,
"storeType": "lmdb",
"index": "spoc,ospc,psoc",
"performanceLogging": false,
"transactionalValidationLimit": -1,
"rdf4jShaclExtensions": false,
"isolationLevel": "READ_COMMITTED",
"isolateDataSources": true
}
}
The csdv-vertx submodule uses a maven plugin to build a container. This container is built locally when you build the submodule.
The docker container spins up a non-clustered Vertx runtime, meaning all the actors in the setup will run on the same compute/container.
You can run the container in docker using a command like:
docker run --name csdv -p 8888:8888 -e CSDV_PORT=8888 inovexis/csdv:latest
The following environment variables are leveraged by the CSDV image to enable customization of the behavior of the running container:
CSDV_PORT
CSDV_HTTPS
CSDV_KEYSTORE
CSDV_KEYSTORE_PASS
CSDV_WORKER_INST
CSDV_SHACL_FETCHER_INST
CSDV_DATA_FETCHER_INST
CSDV_REPORTER_INST
CSDV_BASE_URL_OVERRIDE
CSDV_BASE_CONTEXT_OVERRIDE
You can set these variables using the -e flag (as seen in the example above) to allow customization of your
running docker container.
Longer term, this project would benefit from some of these actions:
Content type
Image
Digest
sha256:6df584b02…
Size
309.5 MB
Last updated
over 1 year ago
docker pull inovexis/csdv