Skip to content

SPI for formats and storage: load any format, store it any way, use it through the standard interfaces #110

Description

@jorabin

The problem

Only the Jackson model in kdbx-database can read and write proper KDBX files. Database, Group and Entry can't hold everything a KDBX file contains, so keeping KeePass data in another form, such as a different in-memory model or a SQL database, either loses data or means reimplementing the KeePass XML schema.

The basic module doesn't solve this. It uses the KDBX encryption layers but writes its own XML (<database>, not <KeePassFile>), so KeePass can't read its files, and it can't read files that KeePass wrote.

More generally, and a long-standing aim: databases in formats other than KDB and KDBX should be usable through the same Database, Group and Entry interfaces. So the goal is:

Load any format, store it any way, and make everything available to users through the standard interfaces.

Proposal: an SPI for formats and storage

Separate three things that are currently bound together in each implementation:

  1. Formats read and write a file format: KDBX 4.x, KDBX 3.1, KDB, and later others. They know about encryption, headers and the document schema, and nothing about storage.
  2. Storage holds the data: the current Jackson model, a plain in-memory model, a SQL database. It knows nothing about file formats.
  3. The API stays Database, Group and Entry, implemented by the storage, which is what applications use.

Formats and storage talk through an SPI: a contract the library calls, not something applications use. It doesn't have to involve ServiceLoader discovery, because the caller chooses the format and the storage. Discovery could be added later.

Content and events

  • Data objects carry the content: DatabaseData, GroupData and EntryData (with history), plus value types such as Times, CustomIcon and DeletedObject. 3.x requires Java 17 from 3.1.0, so these are records (with builders where there are many optional fields).
  • A sink receives the content in document order: database(DatabaseData), startGroup(GroupData), entry(EntryData), endGroup(), deletedObject(DeletedObject), end().
  • A source replays its content into a sink: void writeTo(Sink sink).

So:

  • a format reader is a source: it decrypts, parses and pushes into a sink;
  • a format writer is a sink: a source pushes into it, and it writes and encrypts;
  • storage is both: it's filled by a reader, and replays into a writer;
  • converting between formats or storages is source.writeTo(sink).

Common core and format-specific extensions

Formats don't share a schema, so the data objects have two parts:

  • A common core that every format and storage understands: UUID, name or title, properties (protected or not), binary properties, times, icon, child groups and entries in order. This is roughly what Database, Group and Entry expose now.
  • Format-specific extensions, each a typed value object owned by a format. For KDBX that's everything in the gap table below: AutoType, colours, tags, CustomData, custom icons, Meta settings, history and so on. KDB would have its own (e.g. meta-stream entries), and so would any later format.

Rules for extensions:

  • Storage must keep extensions it doesn't understand, unchanged, and hand them back on replay. A KDBX file then round-trips through any storage without loss, even one that knows nothing about KeePass. (An in-memory store just keeps the objects; a SQL store might keep them as a serialized blob, or map the ones it knows to columns.)
  • Writing to a different format uses the core and ignores extensions that belong to other formats, perhaps reporting what was dropped.
  • Users reach extensions through the standard interfaces, e.g. entry.getExtension(KdbxEntryExtension.class) returning an Optional. This adds no KeePass-specific methods to Entry, and fits the existing supports…() capability methods on Database.

What the formats and storage keep doing

Formats handle file details that aren't content: HeaderHash, the Binaries pool and Ref numbers, the Protected attribute and decrypting protected values in order, date encoding, and the KDBX 4 inner header. None of it appears in the SPI.

The sink has to say what happens with invalid input, such as endGroup without startGroup, or a missing custom icon. Either readers guarantee the schema's rules, or sinks check them.

Considered: derived interfaces

The first version of this issue proposed KeePassEntry extends Entry, KeePassGroup extends Group and KeePassDatabase extends Database, with getters and setters for every KDBX field, and the KDBX reader and writer working against them. It was set aside because:

  • every storage would have to implement a full live object graph (parents, searching, the recycle bin) plus about fifty KeePass setters, just so the reader and writer could reach the fields;
  • it ties storage to one format, and doesn't extend to KDB or other formats;
  • it mixes the application API with the storage contract.

Derived interfaces for applications that want KeePass fields directly could still come later, built on the extensions.

What KDBX needs that the core doesn't hold

This is the specification for the KDBX extensions. It was checked against Reichl's KDBX 4.1 schema (XSD/KDBX.4.1.reichl.xsd), not sample files, which may not contain every element. "Get" and "set" refer to the current Entry, Group and Database interfaces.

Entry (TEntry)

Schema Now
UUID get only
IconID get and set (Icon)
String (protected or not) get and set (properties)
Binary get and set (binary properties)
ExpiryTime, Expires get and set
CreationTime, LastModificationTime, LastAccessTime get only
UsageCount, LocationChanged none
CustomIconUUID, ForegroundColor, BackgroundColor, OverrideURL, QualityCheck, Tags, PreviousParentGroup none
AutoType (Enabled, DataTransferObfuscation, DefaultSequence, Associations) none
CustomData none
History (a list of whole entries) none

Group (TGroup)

Schema Now
UUID get only
Name, IconID, child entries and groups in order get and set
Times same as entries
Notes, CustomIconUUID, IsExpanded, DefaultAutoTypeSequence, EnableAutoType, EnableSearching (true, false or null), LastTopVisibleEntry, PreviousParentGroup, Tags, CustomData none

Database (TMeta, TRoot)

Schema Now
DatabaseName, DatabaseDescription get and set
RecycleBinEnabled get and set
RecycleBinUUID only indirectly (isRecycleBin)
MemoryProtection roughly, via shouldProtect
Generator, SettingsChanged, the …Changed time for each Meta field, DefaultUserName, MaintenanceHistoryDays, Color, MasterKeyChanged, MasterKeyChangeRec, MasterKeyChangeForce, MasterKeyChangeForceOnce, CustomIcons (UUID, Data, Name, LastModificationTime), EntryTemplatesGroup, HistoryMaxItems, HistoryMaxSize, LastSelectedGroup, LastTopVisibleGroup, CustomData with times none
DeletedObjects (UUID, DeletionTime) none

Setting UUIDs and times also comes up in #96.

Rules from the schema

  • UUIDs are unique, apart from history entries, which have their parent entry's UUID.
  • Every CustomIconUUID refers to an icon in Meta/CustomIcons.
  • The order of String, Binary and History items, and of child entries and groups, matters.

Questions

  • The core: exactly which fields are common to all formats, and which are extensions? E.g. is history core or a KDBX extension? Are times core? KDB has them too.
  • History: presumably read-only snapshots with the parent's UUID; the SPI has to say how they're delivered (inside EntryData seems simplest).
  • Dates: decided: Instant. From 3.1.0 Entry and Group times are java.time.Instant rather than Date (KDBX times are UTC and KDB times are treated as UTC, so no zone is needed), and the SPI uses Instant too.
  • KDBX 3.1: its schema differs in places (Meta/Binaries, HeaderHash), but those are format details, so one KDBX extension should cover both.
  • Existing implementations: does the Jackson model become one storage, filled through the SPI, or keep its direct path for speed and also implement the SPI?

Testing

  • A test XML file that uses every element and attribute in the KDBX 4.1 schema at least once, validated against the XSD.
  • A recording sink that captures every event, to check the reader field by field.
  • Round trips of that file through each storage, and through KDBX 4 (and 3.1), compared with the original. A missing field then fails the test, which no sample file can guarantee.

Examples

  • Read a KDBX 4.1 file into storage that doesn't use Jackson, use it through Database, Group and Entry, and write it back with nothing lost.
  • Convert between formats, e.g. KDB to KDBX.

Order

After #109 (write/read) and the move to Java 17, both in 3.1.0, so the new code uses the new stream handling and records from the start. The SPI is planned for release 3.2.0, as experimental at first, possibly with prototype releases for comment. The design is in FormatsAndStorageSPI.md on branch issue-110-spi.

Activity

  1. changed the title [-]Interfaces can't hold everything in a KDBX file, so storage other than the Jackson model can't read and write KDBX[/-] [+]SPI for formats and storage: load any format, store it any way, use it through the standard interfaces[/+] on Oct 3, 2026
  2. astrapi69 commented on Oct 4, 2026

    @astrapi69

    Thanks - yes, this meets what #96 was after, and better than setters would have.

    In this SPI's terms: mystic-crypt-ui exports its own vault to KDBX and imports KDBX into it. Today the export builds Jackson objects and reaches uuid, times and history by reflection, because nothing else can set them. With #110 the export becomes a source pushing DatabaseData / GroupData / EntryData into the KDBX 4 writer, and the import a sink building our own objects from what the reader pushes. No Jackson objects in between, and the reflection goes in both directions.

    For that to work from outside the library:

    1. Public construction of the data records (and their builders), so a source outside the library can supply UUID, all four times with the expiry flag, and history.
    2. History inside EntryData, as a list of EntryData carrying the parent's UUID. For us history is content: our round-trip test asserts it field by field.
    3. Times in the core - KDB has them too.
    4. The protected flag per property in the core, as you have it; we round-trip which properties are protected.

    One question on the extensions rule: for a store that isn't Java objects - ours persists through XStream, you mention SQL - keeping an extension means serializing it. Will an extension have a defined serialized form (the XML fragment it came from would do), or is that left to each store? That decides whether a store like ours can carry AutoType, colours, tags and custom data at all, or has to drop them as it does now.

    Java 17 is no problem here; the application is on 25.

    If it helps, I can check the SPI against our export and import once there is a branch: our round-trip test reads both files back with keepassxc-cli and compares field by field, which might be a useful second check next to the XSD-validated file.

  3. jorabin-sense commented on Oct 4, 2026

    @jorabin-sense

    Will an extension have a defined serialized form (the XML fragment it came from would do), or is that left to each store?

    The intention is that you can do what you want, SQL merely given as an illustration. The mantra being:

    Load any format, store it any way, and make everything available to users through the standard interfaces.

    Pleased to hear that you'll be interested in checking the work as we proceed. It's good to know hat there is one customer out there!

    Bear with me as I go along here, it's going to take a while to do and the work will benefit from your feedback please. Also though I'm really pleased to do it, this is rather a lot of work which doesn't appear in yesterday's to do list! I'd be grateful if you would comment on the design document.

    I'm going to release 3.1.0 shortly. It will feature the Java 17 uplift. And will sort out load/save stream handling which has been wrong from the dawn of time (don't close streams that you did not open) and changing of Date to Instant which should have happened in the interfaces at 3.0.0. This may or may not be relevant to you, will appreciate your view on it as it changes the internal formal as well as at the interface.

  4. astrapi69 commented on Oct 7, 2026

    @astrapi69

    Thanks. On 3.1.0 first, since you asked: I compiled mystic-crypt-ui against it. 18 errors, all in our two KDBX converters, all Date to Instant - and our own model is java.time already, so the change removes conversions on our side rather than adding them. The stream change doesn't touch us: we open and close both streams ourselves around load and save, and will move to read/write with the upgrade. So no objection, and Instant is welcome.

    On the design document, from the point of view of a source and a sink outside the library:

    1. The four points from my last comment are covered: public records with builders, history inside EntryData with the parent's UUID, times in the core, protection per property. One question on the last: a source can make a protected value today with new PropertyValue.SealedStore(...). Will a sink keep isProtected() as it receives it, or re-derive protection from its own Strategy by property name? We round-trip which properties are protected, including ad hoc ones, so we'd need the first.
    2. Dropped data (decision 4): an application wants to show a person what an import or export left out, per entry, and to test it. A Consumer<String> gives a log line; a small record - the element's UUID, what was dropped, and why - would let us list it per entry and assert it without parsing text. The string could still be its toString().
    3. Extensions in other storage (decision 7): understood that the codec is deferred. Until it exists we keep dropping AutoType, colours and CustomData on import, as we do now. When it comes, an XML fragment per element is exactly what we would store - one vote for it, no hurry.
    4. A small one: the core table says "the five times", while Times has four Instants and the expires flag. If LocationChanged is meant to be core, it's missing from the record; if the fifth is the flag, fine.

    The offer stands: once there is a prototype build, I'll run our KDBX round trip against it - export through the SPI, read the file back with keepassxc-cli, compare field by field - and report what differs.

  5. jorabin commented on Oct 7, 2026

    @jorabin
    OwnerAuthor

    thanks for feedback, I'm afraid (though I suppose better to fix bugs, really? 😄) yet another bug-fix release 3.1.1 is to follow please see #114

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions