Skip to content

Repository files navigation

Federated learning within the bounds of patient consent: a multi-hospital use case on personal health data vaults

This repository holds the complete proof of concept: a federated learning (FL) system in which hospitals train a shared human-activity recognition model on accelerometer data that never leaves the participants' personal data pods without an explicit, machine-readable authorisation.

The accompanying paper describes the concepts, the actors and what the system does. This README covers what the paper does not: how the repository is laid out, how access control is actually wired, and what it takes to run.

git clone --recurse-submodules <this-repository>
# or, in an existing clone:
git submodule update --init --recursive
Submodule Branch Commit
aggregator/ pacsoi-poc3-fl a7f922e7f790de919e62341a45ab543810c92010
k8s-fl-services/ main 6eda81e0c0af1253e71993e907b75d4796a4d5ec

Repository layout

Path What it is
aggregator/ The Aggregator platform: an Aggregator Server that registers per-user aggregator instances and deploys UMA-protected services into Kubernetes
k8s-fl-services/ The federated learning workloads: container images and the semantic service descriptions that let the Aggregator deploy them
pacsoi-cli/ The pacsoi CLI that drives the whole setup: pods, slices, UMA policies, enrolment, FL services, outputs
dashboard/ A Dash application that reads live training state from the deployed services

Each directory keeps its own README. pacsoi-cli/README.md is a complete, ordered runbook from empty folders to trained model, and is the best single document for seeing how the pieces are wired together in practice.

Demo video: to be added — a short screen recording of the dashboard following a federated session from the first round to the trained global model.

Access control and policies

Every protected read in this system resolves to an ODRL policy. Nothing is granted by configuration, by network position, or by the hospital holding a copy of the data; a request succeeds because some policy names its requester, its action and its exact target URL. This section is the short version of that model — the mechanism the rest of the architecture exists to serve.

Three layers are involved:

  • Keycloak authenticates. Each actor's sub claim is its identity, used in policies in IRI form, http://example.com/id/<sub>.
  • The UMA authorisation server decides. It stores the policies and issues, per request, a token scoped to one resource.
  • Resource servers enforce: Kvasir pods for participant data, and each Aggregator instance's UMA ingress proxy for service endpoints and outputs.

How one protected request works

  1. The requester calls the URL with no Authorization header and receives 401 with a WWW-Authenticate challenge carrying a ticket naming the resource.
  2. It obtains an id_token from Keycloak for its own identity.
  3. It posts ticket and id_token to the UMA token endpoint under grant_type=urn:ietf:params:oauth:grant-type:uma-ticket.
  4. The authorisation server evaluates the policies whose target is that resource against the token's sub, and returns an access token only if one permits the action.
  5. The requester retries with Authorization: Bearer ….

The deployed workloads never do it themselves — they send outbound requests through their Aggregator's egress proxy, which holds the tokens on their behalf.

What a policy is

One ODRL Agreement wrapping one Permission, submitted to the authorisation server as Turtle:

@prefix odrl: <http://www.w3.org/ns/odrl/2/> .   # written out in full by the CLI

<urn:uuid:agreement> a odrl:Agreement ;
  odrl:uid      <urn:uuid:agreement> ;
  odrl:permission <urn:uuid:permission> .
<urn:uuid:permission> a odrl:Permission ;
  odrl:uid      <urn:uuid:permission> ;
  odrl:action   odrl:read ;
  odrl:target   <https://kvasir.example/participant1/s3/7> ;
  odrl:assignee <http://example.com/id/…> ;   # the hospital's Keycloak sub
  odrl:assigner <http://example.com/id/…> .   # the participant, who owns the resource

One permission, one action, one target, one assignee. There are no wildcards, and the only implicit reach is a partial grant: a policy on a slash-terminated URL also covers its children.

Participant data: one policy per distribution

Enrolling a participant into a hospital's case creates a policy set built from datasetUmaPolicyTargets in dataset.js:

  • one policy on <pod>/slices/accellero/, whose partial grant covers the child /query endpoint the hospital uses to read the dataset's metadata; and
  • one policy on each <pod>/s3/<n> distribution because the distributions live outside the slice's URL tree and no partial grant reaches them.

The assignee is that one hospital's Keycloak subject, so a policy authorises exactly one hospital to read exactly one participant's data.

Service access: roles rather than URLs

Aggregator instances enforce the same way, but a policy there binds a role that a service profile declares rather than a URL. That is what connects the federation: enrolling a hospital grants the hospital training-client on the researcher's weight-aggregation service, and the researcher coordinator on that hospital's training client. aggregator/docs/uma-policies.md documents the policy API, the default-policy rules and the full role model.

Running it

What has to exist first

Three services sit underneath all of this and are deployed separately:

  • a Kvasir Solid Server hosting the pods;
  • a UMA authorisation server, which holds the ODRL policies and issues the tokens every protected request needs; and
  • a Keycloak realm holding the identities of participants, hospitals and researchers, plus an admin client the CLI uses.

Participant data is likewise assumed to be in place: each participant pod must already hold that participant's accelerometer and ground-truth measurements as objects under <pod-url>/s3/, one per partition. The CLI publishes those objects as the dataset's distributions and attaches policies to them, but it never uploads measurement data itself, and the training clients read the data straight out of the pods.

Endpoints, identifiers and credentials

Every URL, identifier and credential in this repository is a placeholder. A deployment supplies its own, so substitute the following before running anything:

Placeholder What it stands for
https://kvasir.example Your Kvasir Solid Server
https://kvasir.example/vocab#, https://kvasir.example/fine-grained-access# The Kvasir vocabulary namespaces the pod and slice documents are written against. Unlike the others these are not addresses but identifiers the server matches exactly, so they must be set to the ones your Kvasir Solid Server publishes.
https://uma.example/uma Your UMA authorisation server
https://keycloak.example/realms/kvasir Your Keycloak realm
https://aggregator.example Your Aggregator Server
hospitalx-aggregator-id, fl-aggregator-id, … The instance ID each Aggregator registration returns
<…-uma-client-id>, <…-uma-client-secret> The UMA client credentials written when a pod is registered
00000000-0000-4000-8000-0000000000NN Keycloak subjects, policy identifiers and pod revisions, all assigned at runtime

The entity folders under pacsoi-cli/participants/, pacsoi-cli/hospitals/ and pacsoi-cli/research/ hold one worked example of each kind. Their contents are what pacsoi … register, … dataset setup, … enroll and … policies write, so they show the shape of the state a real run produces — pod configuration, UMA credentials, aggregator configuration, case plans and ODRL policy dumps.

Steps

The short version; each step is expanded in the component READMEs.

  1. Deploy an Aggregator Server and its definition catalogue (aggregator/docs/local-setup.md), with a Kvasir Solid Server, a UMA authorisation server and Keycloak available.
  2. Build and publish the three FL images and install the two service descriptions (k8s-fl-services/README.md).
  3. Install the CLI, create its configuration file, and follow pacsoi-cli/README.md from section 1 onward: register pods, create dataset and case slices over the data already in the participant pods, enrol participants, register one Aggregator instance per hospital and for the researcher, deploy the FL services, enrol hospitals, and start a session.
  4. Point dashboard/config.py at the resulting service URLs to watch the session run.

Citing

PLACEHOLDER

Disclaimer

The documentation in this repository was written with AI assistance and reviewed by a human.

About

Repository for the paper 'Federated learning within the bounds of patient consent: a multi-hospital use case on personal health data vaults'

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages