Sign inSign up

mds4ul/pprl-broker-web

By mds4ul

•Updated almost 4 years ago

Image
0

978

mds4ul/pprl-broker-web repository overview

⁠PPRL Broker Service

The PPRL broker enables multi-party privacy-preserving record linkage in a trusted third party (TTP) setting. It handles match sessions, accepts bit vectors from different clients, performs matching in the background and provides results to these clients once matching is complete.

⁠Installation

The broker is split into two separate services. One service is the web API itself, the other is the background worker. You can choose to manually install the requirements, or you can choose to run everything via Docker. This document will explain both processes.

But first, you need to ensure that the required external services are running. All services need to be reachable by the web API and the worker.

  • Neo4j
  • Redis
  • AMQP service (e.g. RabbitMQ)
  • PPRL Match Service (version 0.6.0 or higher)

Once you have set up these services, you can continue setting up the PPRL broker.

⁠Running with Uvicorn

The PPRL broker web API can be run with any ASGI web server. This example will show the startup using Uvicorn⁠. It is highly recommended that you use Python 3.10 on the machine that you want to host the broker on.

First, clone this repository. Packaging is done using Poetry⁠, so install it if you haven't already. Navigate to the repository root, then run poetry install --without=dev. This will install all dependencies expect the ones necessary to run tests. It will also automatically install Uvicorn. Make sure to configure the service⁠ before running it. Run the following line to start the web server.

poetry run uvicorn pprl_broker.main:app --host 0.0.0.0 --port 8080

This will make the web server listen on port 8080 of all network interfaces and forward the requests to the PPRL broker. You will also need to run the worker that handles the heavy lifting in the background. Again, make sure to configure the worker⁠. Run the following line to start the worker.

poetry run celery -A pprl_broker.worker.tasks worker --loglevel=INFO
⁠Using a different ASGI server

If you choose to use a different server, do the same steps up to the installation of Poetry as described above. Run poetry install --without=dev,runtime. This will skip the installation of Uvicorn and you can choose to install another server instead.

⁠Running with Docker

We provide Docker images that you can use to run the services in Docker instead. Refer to the mds4ul/pprl-broker-web⁠ and mds4ul/pprl-broker-worker⁠ images. An example configuration using Docker Compose can be found in the Compose file at the root of this repository⁠. You can also simply run docker compose up -d --build instead to create containers based on the source code in this repository.

⁠Configuration

Both the web API and the worker can be configured in the same way. You can manually override the config file,⁠ or you can use environment variables to configure their behavior. If you're using Docker, it is advised to edit the .env file⁠ instead. The following table shows all configurable options.

SettingDescriptionDefault
match_endpoint_urlURL to the endpoint of the match servicehttp://localhost:8080/match⁠
neo4j_urlURL to the Neo4j graph databasebolt://localhost:7687
amqp_urlURL to the AMQP serviceamqp://guest:guest@localhost:5672//
redis_urlURL to the Redis key-value storeredis://localhost:6379/0
max_session_timeoutMaximum selectable duration after which a session expires1h
refresh_session_intervalDuration to extend session by on refresh1h
task_cleanup_intervalDuration after which to start (and repeat) the background cleanup task10s
vector_batch_sizeMaximum amount of client vectors to compare with another client's vectors at a time100
expose_docsExpose documentation on the web API servicetrue

The naming scheme for environment variables is converting the names of these settings into all uppercase and prepending SETTINGS__. So expose_docs would become SETTINGS__EXPOSE_DOCS.

⁠Using the API

This section provides an overview of the API and how you might use it in a PPRL workflow. Every time a PPRL execution is performed, it starts with the creation of a match session. Clients can then submit their bit vectors to a match session and wait for results.

⁠Creating and managing sessions

The /session endpoint is for managing the lifecycle of a match session. Create a new session by doing a POST request. You need to specify the session identifier, the match configuration (which consists of the similarity metric and the matching threshold), and optionally a duration after which the session should expire automatically. Keep in mind that the highest possible duration is determined by the server and sessions need to be refreshed to stay alive. By default, this limit is set to one hour. All example listings will use the Python requests library⁠.

import requests
from datetime import timedelta

r = requests.post("http://localhost:8080/session", json={
    "session": "my-session-id",
    "matchConfig": {
        "measure": "jaccard",
        "threshold": 0.8
    },
    "expiresIn": timedelta(minutes=30).seconds
})

print(r.status_code) # => 201

resp = r.json()

print(resp["session"]) # => "my-session-id"
print(resp["expiresAt"]) # => 1666768346
print(resp["token"]) # => "86aa57b9d1439b91831eb4d35e397194"

Treat the session identifier as a secret that you can only share with session participants you trust. It shouldn't be predictable. Anyone in possession of the session identifier can submit vectors to the session, affecting the quality of results. They can also choose to delete the session preemptively.

The session token is necessary to perform refresh and cancellation operations on a session. Only the client that issued the session creation request should be in possession of this token. It will only be sent once and cannot be requested again.

In the example above, the session is limited to 30 minutes. After that, there's no guarantee that clients will be able to submit any more bit vectors. To extend the lifetime of the session, it needs to be refreshed. This is done by performing a PATCH request to the same endpoint.

import requests

r = requests.patch("http://localhost:8080/session", json={
    "session": "my-session-id",
    "token": "86aa57b9d1439b91831eb4d35e397194"
})

print(r.status_code) # => 200

resp = r.json()

print(resp["session"]) # => "my-session-id"
print(resp["expiresAt"]) # => 1666771946

The session will be extended by a duration that is determined by the server. By default, this duration is one hour.

If you wish to stop the session before it expires, use a DELETE request to the session endpoint. This will immediately prevent clients from submitting any more bit vectors and cancel all match operations. Match results will be purged as soon as possible.

import requests

r = requests.delete("http://localhost:8080/session", json={
    "session": "my-session-id",
    "token": "86aa57b9d1439b91831eb4d35e397194"
})

print(r.status_code) # => 202
⁠Submitting vectors and receiving results

There are two endpoints that can be used by clients. One is for submitting bit vectors and one is for requesting results. To submit bit vectors, issue a POST request to the /session/submit endpoint.

import requests

r = requests.post("http://localhost:8080/session/submit", json={
    "session": "my-session-id",
    "client": "my-client-id",
    "vectors": [
        {
            "id": "001",
            "value": "CE9stxXqVmVQkHiZAZfE9w==",
            "metadata": [
                {
                    "name": "count",
                    "value": "10"
                }
            ]
        }
    ]
})

print(r.status_code) # => 202

Just like the session identifier, treat the client identifier like a secret. Anyone in the possession of the client identifier can submit vectors on the client's behalf, affecting results. They can also obtain unauthorized access to match results of that client.

This endpoint can be called as many times as necessary. Every time it is called, the vectors are stored at the broker and queued for matching with all other clients. This also enables a client to submit their vectors in fixed-size chunks instead of performing one massive request.

Once matching has concluded, clients can retrieve matches for the vectors they submitted by running a POST request against the /session/result endpoint.

import requests

r = requests.post("http://localhost:8080/session/result", json={
    "session": "my-session-id",
    "client": "my-client-id",
    "showUnfinishedResults": False
})

print(r.status_code) # => 200

resp = r.json()

print(resp["finished"]) # => true
print(len(resp["matches"])) # => 1

r_match = resp["matches"][0]

print(r_match["vector"]["id"]) # => "001"
print(r_match["similarity"]) # => 0.98
print(r_match["referenceMetadata"]) # => [BitVectorMetadata(name="count", value="20")]

The result contains a list of matched vectors which have been found to be sufficiently similar to vectors submitted by other clients. Every vector in the list contains the computed similarity, as well as the metadata of the other client's vector. This information can be used to perform automated decision-making.

If you're not sure whether matching has concluded or not, you can run the same request as above. The finished field in the response says whether matching has concluded or not. If showUnfinishedResults is set to false in the request and matching hasn't finished yet, then the result list is empty. However, you can set that field to true in order to receive preemptive results even when matching hasn't finished yet.

⁠Running tests

Run the linter in the root directory using poetry run flake8.

To run unit and integration tests, navigate to the tests directory⁠ and run docker compose up -d. This will bring up fresh instances of the required services. Then, go back to the root of this repository. To run the worker, execute the following line.

poetry run celery -A pprl_broker.worker.tasks worker -D --pidfile celery.pid --logfile celery.log --loglevel INFO

The worker might need a few seconds to connect to RabbitMQ. Finally, run poetry run pytest. After testing is complete, terminate the worker by running pkill -F celery.pid. Stop the services by running docker compose down -v in the tests directory⁠.

⁠License

MIT.

Tag summary

Content type

Image

Digest

sha256:be0a1f7a8…

Size

53.6 MB

Last updated

almost 4 years ago

docker pull mds4ul/pprl-broker-web