Federated learning within the bounds of patient consent: a multi-hospital use case on personal health data vaults
This repository holds the complete proof of concept: a federated learning (FL) system in which hospitals train a shared human-activity recognition model on accelerometer data that never leaves the participants' personal data pods without an explicit, machine-readable authorisation.
The accompanying paper describes the concepts, the actors and what the system does. This README covers what the paper does not: how the repository is laid out, how access control is actually wired, and what it takes to run.
git clone --recurse-submodules <this-repository>
# or, in an existing clone:
git submodule update --init --recursive| Submodule | Branch | Commit |
|---|---|---|
| aggregator/ | pacsoi-poc3-fl |
a7f922e7f790de919e62341a45ab543810c92010 |
| k8s-fl-services/ | main |
6eda81e0c0af1253e71993e907b75d4796a4d5ec |
| Path | What it is |
|---|---|
| aggregator/ | The Aggregator platform: an Aggregator Server that registers per-user aggregator instances and deploys UMA-protected services into Kubernetes |
| k8s-fl-services/ | The federated learning workloads: container images and the semantic service descriptions that let the Aggregator deploy them |
| pacsoi-cli/ | The pacsoi CLI that drives the whole setup: pods, slices, UMA policies, enrolment, FL services, outputs |
| dashboard/ | A Dash application that reads live training state from the deployed services |
Each directory keeps its own README. pacsoi-cli/README.md is a complete, ordered runbook from empty folders to trained model, and is the best single document for seeing how the pieces are wired together in practice.
Demo video: to be added — a short screen recording of the dashboard following a federated session from the first round to the trained global model.
Every protected read in this system resolves to an ODRL policy. Nothing is granted by configuration, by network position, or by the hospital holding a copy of the data; a request succeeds because some policy names its requester, its action and its exact target URL. This section is the short version of that model — the mechanism the rest of the architecture exists to serve.
Three layers are involved:
- Keycloak authenticates. Each actor's
subclaim is its identity, used in policies in IRI form,http://example.com/id/<sub>. - The UMA authorisation server decides. It stores the policies and issues, per request, a token scoped to one resource.
- Resource servers enforce: Kvasir pods for participant data, and each Aggregator instance's UMA ingress proxy for service endpoints and outputs.
- The requester calls the URL with no
Authorizationheader and receives401with aWWW-Authenticatechallenge carrying a ticket naming the resource. - It obtains an
id_tokenfrom Keycloak for its own identity. - It posts ticket and
id_tokento the UMA token endpoint undergrant_type=urn:ietf:params:oauth:grant-type:uma-ticket. - The authorisation server evaluates the policies whose target is that resource against the token's
sub, and returns an access token only if one permits the action. - The requester retries with
Authorization: Bearer ….
The deployed workloads never do it themselves — they send outbound requests through their Aggregator's egress proxy, which holds the tokens on their behalf.
One ODRL Agreement wrapping one Permission, submitted to the authorisation server as Turtle:
@prefix odrl: <http://www.w3.org/ns/odrl/2/> . # written out in full by the CLI
<urn:uuid:agreement> a odrl:Agreement ;
odrl:uid <urn:uuid:agreement> ;
odrl:permission <urn:uuid:permission> .
<urn:uuid:permission> a odrl:Permission ;
odrl:uid <urn:uuid:permission> ;
odrl:action odrl:read ;
odrl:target <https://kvasir.example/participant1/s3/7> ;
odrl:assignee <http://example.com/id/…> ; # the hospital's Keycloak sub
odrl:assigner <http://example.com/id/…> . # the participant, who owns the resourceOne permission, one action, one target, one assignee. There are no wildcards, and the only implicit reach is a partial grant: a policy on a slash-terminated URL also covers its children.
Enrolling a participant into a hospital's case creates a policy set built from datasetUmaPolicyTargets in dataset.js:
- one policy on
<pod>/slices/accellero/, whose partial grant covers the child/queryendpoint the hospital uses to read the dataset's metadata; and - one policy on each
<pod>/s3/<n>distribution because the distributions live outside the slice's URL tree and no partial grant reaches them.
The assignee is that one hospital's Keycloak subject, so a policy authorises exactly one hospital to read exactly one participant's data.
Aggregator instances enforce the same way, but a policy there binds a role that a service profile declares rather than a URL. That is what connects the federation: enrolling a hospital grants the hospital training-client on the researcher's weight-aggregation service, and the researcher coordinator on that hospital's training client. aggregator/docs/uma-policies.md documents the policy API, the default-policy rules and the full role model.
Three services sit underneath all of this and are deployed separately:
- a Kvasir Solid Server hosting the pods;
- a UMA authorisation server, which holds the ODRL policies and issues the tokens every protected request needs; and
- a Keycloak realm holding the identities of participants, hospitals and researchers, plus an admin client the CLI uses.
Participant data is likewise assumed to be in place: each participant pod must already hold that participant's accelerometer and ground-truth measurements as objects under <pod-url>/s3/, one per partition. The CLI publishes those objects as the dataset's distributions and attaches policies to them, but it never uploads measurement data itself, and the training clients read the data straight out of the pods.
Every URL, identifier and credential in this repository is a placeholder. A deployment supplies its own, so substitute the following before running anything:
| Placeholder | What it stands for |
|---|---|
https://kvasir.example |
Your Kvasir Solid Server |
https://kvasir.example/vocab#, https://kvasir.example/fine-grained-access# |
The Kvasir vocabulary namespaces the pod and slice documents are written against. Unlike the others these are not addresses but identifiers the server matches exactly, so they must be set to the ones your Kvasir Solid Server publishes. |
https://uma.example/uma |
Your UMA authorisation server |
https://keycloak.example/realms/kvasir |
Your Keycloak realm |
https://aggregator.example |
Your Aggregator Server |
hospitalx-aggregator-id, fl-aggregator-id, … |
The instance ID each Aggregator registration returns |
<…-uma-client-id>, <…-uma-client-secret> |
The UMA client credentials written when a pod is registered |
00000000-0000-4000-8000-0000000000NN |
Keycloak subjects, policy identifiers and pod revisions, all assigned at runtime |
The entity folders under pacsoi-cli/participants/, pacsoi-cli/hospitals/ and pacsoi-cli/research/ hold one worked example of each kind. Their contents are what pacsoi … register, … dataset setup, … enroll and … policies write, so they show the shape of the state a real run produces — pod configuration, UMA credentials, aggregator configuration, case plans and ODRL policy dumps.
The short version; each step is expanded in the component READMEs.
- Deploy an Aggregator Server and its definition catalogue (
aggregator/docs/local-setup.md), with a Kvasir Solid Server, a UMA authorisation server and Keycloak available. - Build and publish the three FL images and install the two service descriptions (
k8s-fl-services/README.md). - Install the CLI, create its configuration file, and follow
pacsoi-cli/README.mdfrom section 1 onward: register pods, create dataset and case slices over the data already in the participant pods, enrol participants, register one Aggregator instance per hospital and for the researcher, deploy the FL services, enrol hospitals, and start a session. - Point
dashboard/config.pyat the resulting service URLs to watch the session run.
PLACEHOLDERThe documentation in this repository was written with AI assistance and reviewed by a human.