Connection Pooling Is a Distributed Systems Problem

One of the first things we usually do when building a Java application that talks to a relational database is configure a connection pool. The reason is straightforward: opening a database connection is expensive, and an application runs a huge number of database operations over its lifetime. Paying that price every single time is not something we want to do. With PostgreSQL the cost is even more tangible than with other databases, because every connection is served by its own backend process on the server. Each one carries its own memory, its own share of the scheduler, and its own contribution to lock and snapshot bookkeeping, which is why max_connections starts to hurt long before the machine looks busy. So, instead of creating a new connection for every operation, we keep a collection of connections around and reuse them. In the Java ecosystem, HikariCP has become the default implementation of this pattern, and frameworks such as Spring Boot will configure it for us without asking us to think about the details.

The architecture is quite simple. The application borrows a connection from HikariCP, uses it to execute one or more statements, and returns it to the pool. HikariCP takes care of creating connections when necessary, keeping the pool within its configured limits, retiring old connections, and coordinating the threads that are waiting for one. From the application’s perspective, it is a very convenient abstraction: the database connection becomes just another managed resource.

The problem appears when we stop running one application instance and start running many. Suppose our HikariCP pool is configured with a maximum of 20 connections. With a single instance, the application can hold at most 20 database connections. If we deploy ten instances, there are now ten independent pools, each allowed to open up to 20 connections, so we are potentially at 200. If we are in an environment with auto-scaling, and the number of instances grows to, let’s say, 100, the same configuration can end up asking PostgreSQL for 2,000 connections, and nobody changed a single line of configuration to get there.

Nothing strange has happened inside HikariCP. Every instance is behaving exactly as we asked it to behave. The problem is that HikariCP only has local knowledge. Instance A knows how many connections it is using, instance B knows how many it is using, and so on, but there is no shared resource manager coordinating those decisions. The database, on the other hand, is a shared resource. It has a finite amount of CPU, memory, and I/O, and a finite practical limit on how much concurrent work it can do.

This creates a mismatch between two scaling models. Application infrastructure is usually designed to scale horizontally: if traffic increases, we add more instances. Database infrastructure does not scale that way, or at least not as cheaply. Adding another application instance can indirectly increase the pressure on the database for no other reason than that the new instance brings another connection pool with it.

The obvious response is to make the HikariCP pool smaller. Instead of allowing every instance to hold 20 connections, perhaps we configure five. Many teams go one step further and do the arithmetic explicitly: take the connection budget PostgreSQL can afford, divide it by the maximum number of replicas the autoscaler is allowed to create, and use the result as the pool size. It works, but it is fragile. The number has to be recalculated every time someone raises the autoscaler limit, adds a new service against the same database, or runs a deployment that briefly doubles the number of pods during a rolling update. More importantly, it doesn’t change the fundamental relationship between application scaling and database connections. With 100 instances and a pool of five, we are still potentially at 500 connections, and most of them will be sitting idle on instances that happen to be quiet while the busy instances queue for their five. The database connection capacity is still coupled to the number of application instances.

This is where an external connection pooler such as PgBouncer changes the architecture. Instead of every application instance opening its physical connections directly against PostgreSQL, we introduce another layer between the applications and the database.

There are now two different kinds of connections. The connections between the applications and PgBouncer are client connections, while the connections between PgBouncer and PostgreSQL are server connections, the real backend processes. PgBouncer keeps a pool of the latter and multiplexes activity from a much larger number of clients onto a smaller number of actual PostgreSQL connections.

This distinction matters because the number of application-side connections no longer has to match the number of backend processes PostgreSQL has to run. We might have hundreds of application instances and thousands of client connections while deliberately capping the connections PgBouncer opens to PostgreSQL at, let’s say, 50. The pooler becomes the point at which we impose a global limit on the database, rather than relying only on independent limits inside each application process.

It also explains why, contrary to what many people think, HikariCP and PgBouncer are not really competitors. They operate at different layers. HikariCP is a library running inside each JVM, and manages the connections available to that application instance. PgBouncer is an external process, and manages the connections between itself and the database. It is perfectly valid to have both, and it is a very natural setup when we want local connection management and a central limit at the same time: HikariCP stops an individual instance from opening an uncontrolled number of connections, while PgBouncer stops the fleet as a whole from consuming an uncontrolled number of backend processes.

There is a detail I skipped over in the previous paragraphs, and it is the one that decides whether any of this works: PgBouncer’s pool mode. PgBouncer can operate in three modes, and they give very different results.

In session mode, a client connection is assigned a server connection when it connects and keeps it until it disconnects. That is the most compatible mode, but if we put HikariCP in front of it, which opens its connections at startup and keeps them open for as long as it can, we have gained almost nothing. Every Hikari connection pins a PostgreSQL backend, and we are back to 100 instances times 20 connections, now with an extra hop in the middle.

In transaction mode, a server connection is assigned to a client only for the duration of a transaction and goes back to PgBouncer’s pool when the transaction commits or rolls back. This is the mode where the multiplexing really happens, and the reason it works so well is that most application connections are idle most of the time. A Hikari connection spends the bulk of its life waiting between transactions, or waiting on the application to do something with the results. In transaction mode, that idle time no longer costs a backend process.

There is also a statement mode, which releases the server connection after every single statement and therefore forbids multi-statement transactions. It has its uses, but for the typical Java service it is not the mode you want.

Transaction mode comes with a catch, though, and it is one that Java applications run into quite frequently. Because consecutive transactions from the same client connection can land on different server connections, anything that relies on session state stops being reliable. A SET executed outside a transaction, session-level advisory locks, LISTEN/NOTIFY, temporary tables that outlive a transaction, all of them can silently end up on a different backend than the one we expect. The classic trap is prepared statements. The PostgreSQL JDBC driver switches to named server-side prepared statements after a statement has been executed a few times (controlled by prepareThreshold, five by default), and historically that produced errors such as prepared statement "S_1" does not exist when the next execution landed on a different backend. For years the answer was to add prepareThreshold=0 to the JDBC URL and give up server-side prepared statements. Since PgBouncer 1.21, setting max_prepared_statements makes PgBouncer track protocol-level prepared statements itself and re-prepare them transparently on whatever server connection the client lands on, which removes most of the pain.

As a reference, a minimal setup could look something like this:

; pgbouncer.ini
[databases]
orders = host=postgres.internal port=5432 dbname=orders
[pgbouncer]
listen_port = 6432
pool_mode = transaction
max_client_conn = 5000 ; client connections PgBouncer will accept
default_pool_size = 40 ; server connections per database/user pair
reserve_pool_size = 5
query_wait_timeout = 30 ; seconds a client may wait for a server connection
server_lifetime = 3600
server_idle_timeout = 600
max_prepared_statements = 200 ; PgBouncer 1.21+

And on the application side:

spring:
datasource:
# With PgBouncer < 1.21 add ?prepareThreshold=0 to the URL
url: jdbc:postgresql://pgbouncer.internal:6432/orders
hikari:
maximum-pool-size: 20
connection-timeout: 5000
max-lifetime: 1800000
idle-timeout: 600000

Notice that default_pool_size is per database and user pair, not per PgBouncer instance as a whole, so every service that connects with its own credentials gets its own pool of server connections. The real ceiling on PostgreSQL is the sum of those pools.

There is a price for introducing the second layer. We now have two systems managing database connectivity, and their configuration and behaviour have to be considered together. Connection limits, timeouts, connection lifetimes, and failure behaviour exist at both layers, and they interact in ways that are not obvious at first. Take waiting as an example. A request thread in our service first waits for HikariCP to hand it a connection, bounded by connection-timeout. Once it has one, the first statement of the transaction may wait again inside PgBouncer for a server connection, bounded by query_wait_timeout. That is two queues, with two timeouts, surfacing two different errors, and whoever is on call at three in the morning needs to know which one fired. The same happens with lifetimes: Hikari’s max-lifetime now governs the cheap connection to PgBouncer, while server_lifetime governs the expensive one to PostgreSQL, and tuning one does nothing for the other.

That queue inside PgBouncer is also worth pausing on, because it is already a form of backpressure. When all server connections are busy, clients do not get new backends, they wait, and after query_wait_timeout they are rejected. That is admission control at the database boundary, and it is one of the main reasons a central pooler protects the database better than a hundred independent pools.

And there is one more detail that is easy to miss: the limit is only global if PgBouncer itself is central. PgBouncer is single-threaded, so at high throughput teams run several instances, and each one maintains its own pools. Five PgBouncer instances with default_pool_size = 40 means up to 200 server connections, not 40. Some teams also deploy PgBouncer as a sidecar in every application pod, which is convenient but brings us right back to the original multiplication problem, just one layer further down. Where the pooler runs is as much an architectural decision as whether we run one at all.

If you are on a managed PostgreSQL offering, a lot of this may already be available as a service. Amazon RDS Proxy, Supabase’s Supavisor, and the built-in PgBouncer in Azure Database for PostgreSQL are all variations of the same idea, and the same questions about pool modes, session state, and timeouts apply to them.

PgBouncer sits at the PostgreSQL protocol boundary. It does not need to know whether the application using it is written in Java, Go, or Python. As long as the client speaks the PostgreSQL protocol, PgBouncer can sit in the middle. That is one of the reasons this architecture works so well in organisations where many services and languages share the same PostgreSQL infrastructure.

Other PostgreSQL proxies take this idea further. PgCat, for example, combines connection pooling with health checking, load balancing, failover, and read/write splitting. PgDog, written by one of PgCat’s original authors, is a multi-threaded proxy that provides pooling and routing, including support for sharding. These are not alternative implementations of HikariCP. They are PostgreSQL-aware infrastructure components that sit at the database boundary and can make decisions based on the topology and behaviour of the PostgreSQL cluster.

Here, the proxy is no longer only a mechanism for reducing the number of database connections. It becomes part of the database topology, deciding where different types of traffic should go and reacting when one of the nodes becomes unavailable.

For a long time, these were usually the options people compared and combined when making this decision. Recently, though, Open J Proxy (OJP) reached version 1.0.0, its first production-ready release (JavaPro article).

OJP approaches the problem from a different direction. Rather than being a proxy for one database protocol, it is built around the Java/JDBC boundary. The application uses the OJP JDBC driver, which talks to an OJP server over gRPC, and the OJP server becomes the component responsible for the actual database connections. Because it sits at the JDBC level instead of the wire protocol, it is not tied to PostgreSQL; the same server can front PostgreSQL, MySQL, MariaDB, Oracle, SQL Server, DB2, and others with a JDBC driver, which is a real difference from everything else in this article. The flip side is that it only helps Java clients. A Python batch job or a Go service hitting the same database will go around it.

The OJP server uses HikariCP by default (DBCP is also available), so the underlying pooling mechanism has not disappeared; what has changed is where the pool lives. With a normal deployment, every application instance has its own HikariCP pool; with OJP, the physical pool is centralised in the OJP server. The OJP driver’s connections are virtual, and putting HikariCP on top of them would add a queue that knows nothing about the real limit, so it makes sense to remove the application-side pool and let the OJP server be the single place where connections are managed.

So OJP is not replacing HikariCP with another implementation of the same abstraction. It moves the physical connection pool out of the application instances and into a shared service. The application still talks to the database through JDBC, but the physical connection becomes an infrastructure concern rather than something each JVM manages on its own. In that sense, OJP and HikariCP are not mutually exclusive either; it is closer to OJP wrapping HikariCP than to OJP versus HikariCP.

This lets us look at database connectivity as an admission-control problem. Suppose the database can safely sustain a certain amount of concurrent work. It doesn’t really matter whether that work comes from ten application instances or one hundred; what matters is the total amount reaching the database, and a central component can enforce that limit across the whole fleet.

It is fair to say, though, that PgBouncer already gives us that global point, as we saw with its wait queue. So the interesting question is not whether OJP can limit concurrency (both can) but what it can do because it sits at the JDBC level rather than at the wire protocol. Because the server sees JDBC operations with the context of the datasource they come from, it can make decisions a protocol-level pooler has a harder time making. One example is its slow query segregation, which separates slow and fast operations into different lanes so that a handful of heavy reporting queries cannot occupy every connection and starve the quick transactional ones. Features like backpressure, concurrency limits, circuit breaking, and query monitoring are not different ways of pooling connections; they are mechanisms for controlling the relationship between an elastic application layer and a comparatively constrained database layer, and the layer we put them in determines how much they know.

OJP deserves the same critical look we gave PgBouncer, though. Every database call now makes an additional network hop, from the application to the OJP server over gRPC, before it even reaches the database, and that latency is paid on every statement, not only when connections are opened. The OJP server is also now on the critical path for every Java service using it, so it needs to be deployed with redundancy, scaled, monitored, and upgraded like any other piece of shared infrastructure. None of this is unique to OJP; a central PgBouncer has exactly the same concerns, but it is important not to forget that centralising the pool also means centralising a failure point.

The important change is not that OJP has a pool and PgBouncer has a pool. Both do. The question is what the component surrounding that pool knows about and what responsibilities it has. HikariCP manages a pool inside one Java process. PgBouncer manages PostgreSQL connections at the protocol boundary. PgCat and PgDog add topology and routing concerns. OJP puts the pool behind a Java-aware service and can provide application-level controls around database access.

Do we actually need several layers? There is no universal answer. HikariCP plus PgBouncer is a good fit when we want local pooling in each service and a central limit that works for any language. OJP provides a similar centralisation model for Java applications while adding controls at the JDBC boundary. A PostgreSQL-specific proxy such as PgCat or PgDog may be preferable when the central problem is topology, such as routing reads to replicas or distributing work across shards.

What I would avoid is stacking all of them blindly. It is technically possible, but probably not a good idea. For example, we could build something like Java -> OJP -> PgBouncer -> PostgreSQL, but now there are two components managing connection limits, queueing, timeouts, and connection lifetimes. There may be valid reasons for it, for example, if OJP provides application-level controls while a PostgreSQL proxy handles routing to replicas, but adding another pool does not automatically improve the system. Each additional layer is another place where requests can queue, another set of timeouts to understand, and another failure mode to reason about. The same applies to PgBouncer, PgCat, and PgDog among themselves. They occupy roughly the same boundary, so in most architectures we would choose the one whose capabilities match our problem rather than chaining them.

The deeper problem behind all of these technologies is not connection pooling. Connection pooling is the mechanism we started with because database connections are an expensive and finite resource. Once the application becomes a distributed system, however, we discover that the resource is shared by many independent processes, while the original pool was designed to manage only one.

That leads to a more general distributed-systems problem: how do we govern a shared, finite resource when the clients consuming it are elastic?

  • For a small Java application, the answer can remain entirely local: Application -> HikariCP -> Database.
  • As the fleet grows, we may introduce a PostgreSQL-level pooler: Application Fleet -> HikariCP pools -> PgBouncer -> Database.
  • If database topology becomes part of the problem, a PostgreSQL-aware proxy can take responsibility for routing and failover: Application Fleet -> PgCat / PgDog -> [Primary | Replicas].
  • And if the environment is predominantly Java, and we want database access itself to be a centrally managed capability, we can move the physical pool behind OJP: Application Fleet -> OJP (HikariCP) -> Database.

The architectural decision is not which connection pool to use. It is where we want the responsibility for database resource management to live. HikariCP places it inside each application instance. PgBouncer places it at the PostgreSQL protocol boundary. PgCat and PgDog extend that boundary into routing and topology management. OJP moves it into a shared, Java-aware infrastructure layer while still using HikariCP underneath.

Once we look at it this way, these technologies stop looking like competing connection pools and start looking like different answers to the same scaling problem: the application layer wants to scale independently, while the database remains a shared and finite resource.

The connection pool is simply the first place where that tension becomes visible.

Connection Pooling Is a Distributed Systems Problem

How Object-Storage-Native LSM-Trees Work Under the Hood

In the last few years, data stores have changed, and “a lot” feels like an understatement. Around 2013 RocksDB was shipped, and it assumed a POSIX filesystem underneath it, because back then that was just what a fast key-value engine ran on. A decade or so later, key-value engines like SlateDB, and analytical table formats like Iceberg or Delta Lake, are running LSM-style structures straight against S3, where that assumption doesn’t hold anymore. This post is about what actually has changed to make that work.

An object-storage-native LSM tree isn’t a fundamentally new database architecture, it’s the same Log-Structured Merge-tree RocksDB has shipped for over a decade, but with the traditional POSIX filesystem replaced by immutable objects and a manifest updated via compare-and-swap (CAS). By running directly on cloud storage like AWS S3 or Google Cloud Storage, it adapts the classical design around three core characteristics:

  • Separation of compute & storage: State is persisted in scalable, low-cost object stores rather than local NVMe drives.
  • Immutable file alignment: Because LSM trees naturally write data sequentially into immutable files (Static Sorted Tables – SSTs), they natively match object storage’s write-once, read-many design.
  • Cloud-optimised I/O: Compaction and read paths are optimised to handle object store latency, high GET/PUT bandwidth, and explicit API call costs.

RocksDB’s write path depends on three things that don’t really have anything to do with LSM trees, they’re just what a POSIX filesystem gives us for free: we can append a few bytes to an existing file, we can fsync those bytes and know they survived a crash, and we can atomically rename a temp file over a real one to make a change visible in one step. The WAL leans on the first two. The MANIFEST, RocksDB’s own record of “which SSTs currently exist”, leans on the third.

If any one of those is taken away, the system does not degrade gracefully, it just stop working. S3 takes away all three. There’s no append, a PUT replaces the whole object. There’s no fsync, durability is whatever the object store’s replication does behind the scenes, and it happens on a timescale of tens to hundreds of milliseconds instead of the low single digits a local NVMe fsync costs. And there’s no rename, only, since August 2024, a conditional PUT: “create this key, but only if it doesn’t already exist“.

S3 is limited by its design, we cannot make it faster, so what is the smallest change to an LSM tree’s write path that survives losing append, fsync, and rename, while leaving everything else, memtables, sorted runs, compaction, exactly as it was? To figure out the answer we need to inspect what each missing primitive was actually protecting.

Append and fsync existed to make a partial write durable, a handful of bytes, safely, before the file that holds them is complete. If we can’t do that cheaply anymore, the fix isn’t to find a workaround, it’s to stop needing partial durability at all: buffer writes in memory until we have a whole, complete, self-contained object worth writing, and pay one network round trip for the whole thing instead of one round trip per record. This is exactly what a memtable already is. Object storage doesn’t force a new component into the design here; it just makes the memtable’s flush threshold matter for latency in a way it never did locally.

Rename existed to make a set of files change atomically, so a reader never observes “half the new SSTs, half the old ones”. Once we can’t rename, the only way to keep that guarantee is to never let the set of files change in place at all: every SST, once written, is permanently immutable, and the only thing that ever changes is a single small pointer, a manifest, listing which immutable SSTs are currently live. Updating that pointer is now the one operation in the entire system that needs an atomic primitive, and it’s small and infrequent enough that a conditional PUT and a retry loop can carry it.

RocksDB’s SSTs were already immutable once flushed, that part isn’t new. What’s new is that immutability stops being an implementation detail and becomes a first-level citizen of the entire system. Locally, “immutable” mostly meant “compaction rewrites files instead of editing them“, a convenience for concurrency control inside one process. On object storage, immutability is the only reason concurrent readers and writers can share a table at all without coordinating. A reader holding an old manifest can keep reading old SSTs indefinitely, safely, even while a compaction job somewhere else is busy writing brand-new ones, because nothing the reader is looking at will ever be touched again. Nobody has to lock anything. Nobody has to tell the reader to wait. The old SSTs just sit there, unreferenced eventually, garbage, but never wrong.

This is the part worth a deep consideration, because it’s a genuine inversion of where durability lives. In RocksDB, the WAL is the thing standing between us and data loss, and the MANIFEST is comparatively an afterthought, a bookkeeping file rebuilt from the WAL if it ever gets confused. In an object-storage-native LSM, that hierarchy flips. The manifest, one small JSON or Avro object, updated by Compare-And-Swap (or Compare-And-Set, CAS), is the database. It’s the single point that defines “what does this table currently contain“, and every SST it doesn’t list, however durably it sits in S3, might as well not exist.

That’s why the commit path collapses to one operation: read the current manifest, compute the new one, try to write it at the next version number with put_if_absent. If someone else got there first, the write fails, not with data loss, with a clean, detectable rejection, and we retry against their version instead of ours. This is optimistic concurrency control, the same pattern MVCC databases have used internally for decades, except here the granularity is “the whole table’s file list” instead of “one row“, and the retry cost is a network round trip instead of a spinlock.

SlateDB applies this architecture directly to low-latency key-value workloads by flushing memtables and conditional-PUTing manifest updates directly to S3. Analytical table formats like Iceberg and Delta Lake aren’t point-lookup KV engines, but they apply this exact same paradigm to columnar datasets: raw data is stored in immutable Parquet objects, while state changes (like Merge-on-Read or Copy-on-Write updates) are committed by racing to swap a manifest pointer, either native in S3 or via an external catalog (e.g., Hive metastore, Glue, REST catalog). The underlying engine goals differ, but the storage mechanics are identical: never mutate in place, always write new immutable objects, and guard the active manifest with compare-and-swap.

The design reads cleanly on paper, but as always the best way to learn is hands-on. Below is a minimal engine, in memory buffering, immutable SSTs flushed as whole objects, a manifest committed by CAS-and-retry, and a compaction pass that merges and swaps atomically, against a fake object store that only exposes what S3 actually gives us: put, put_if_absent, get, list.

The example is going to be in Python for convenience using only built-in standard library modules. A simple python script.py should suffice to run it.

import bisect
import json
import time
from dataclasses import dataclass, field
class ConditionalWriteFailed(Exception):
"""The S3 analogue of a failed compare-and-swap on If-None-Match."""
class FakeObjectStore:
"""No append, no in-place edits. The only concurrency primitive is
'create this key, but only if it doesn't already exist.'"""
def __init__(self, put_latency_ms=80):
self._objects: dict[str, bytes] = {}
self.put_latency_ms = put_latency_ms # simulated network cost
def put(self, key: str, data: bytes) -> None:
time.sleep(self.put_latency_ms / 1000)
self._objects[key] = data
def put_if_absent(self, key: str, data: bytes) -> None:
time.sleep(self.put_latency_ms / 1000)
if key in self._objects:
raise ConditionalWriteFailed(key)
self._objects[key] = data
def get(self, key: str) -> bytes:
return self._objects[key]
def list(self, prefix: str) -> list[str]:
return sorted(k for k in self._objects if k.startswith(prefix))
@dataclass
class SSTable:
"""Written once, never touched again. A real SST carries a sparse
block index and a Bloom filter so a miss doesn't cost a full fetch.
This one is small enough that the whole thing is the index."""
sst_id: str
entries: list[tuple[str, str | None]] # (key, value); None = tombstone
def get(self, key: str) -> str | None:
i = bisect.bisect_left([k for k, _ in self.entries], key)
if i < len(self.entries) and self.entries[i][0] == key:
return self.entries[i][1]
return None
def to_bytes(self) -> bytes:
return json.dumps(self.entries).encode()
@classmethod
def from_bytes(cls, sst_id: str, data: bytes) -> "SSTable":
return cls(sst_id, [tuple(e) for e in json.loads(data)])
@dataclass
class Manifest:
"""The one thing in this whole system that ever changes. Everything
it doesn't list might as well not exist."""
version: int
sst_ids: list[str] = field(default_factory=list)
def to_bytes(self) -> bytes:
return json.dumps({"version": self.version, "sst_ids": self.sst_ids}).encode()
@classmethod
def from_bytes(cls, data: bytes) -> "Manifest":
d = json.loads(data)
return cls(d["version"], d["sst_ids"])
class ObjectStoreLSM:
def __init__(self, store: FakeObjectStore, table: str = "t1"):
self.store = store
self.table = table
self.memtable: dict[str, str | None] = {}
self.local_cache: dict[str, SSTable] = {}
self._cached_manifest: Manifest | None = None
self._ensure_manifest_exists()
def _manifest_key(self, version: int) -> str:
return f"{self.table}/manifest/{version:06d}.json"
def _ensure_manifest_exists(self):
if not self.store.list(f"{self.table}/manifest/"):
self.store.put_if_absent(self._manifest_key(0), Manifest(0, []).to_bytes())
def _current_manifest(self) -> Manifest:
"""Real systems cache this pointer and only re-fetch on a CAS
conflict, rather than paying LIST+GET on every read; that's what
the cache below is for. (They still need some way to notice a
*different* writer moved the pointer without telling this
process: a poll, a watch, or a version check on some other
operation. This toy has exactly one writer, so it never has to
solve that half of the problem.)"""
if self._cached_manifest is None:
latest_key = self.store.list(f"{self.table}/manifest/")[-1]
self._cached_manifest = Manifest.from_bytes(self.store.get(latest_key))
return self._cached_manifest
def _commit_append(self, new_sst_ids: list[str], retries: int = 5) -> Manifest:
"""For flush(): a new SST doesn't depend on anything else that
might land first, so on conflict it's always safe to replay it
on top of whatever the latest version turns out to be."""
for _ in range(retries):
current = self._current_manifest()
candidate = Manifest(current.version + 1, current.sst_ids + new_sst_ids)
try:
self.store.put_if_absent(
self._manifest_key(candidate.version), candidate.to_bytes()
)
self._cached_manifest = candidate
return candidate
except ConditionalWriteFailed:
self._cached_manifest = None # someone else landed that version -> re-read
raise RuntimeError("manifest commit did not converge. contention too high")
def _commit_replace(self, sst_ids: list[str], based_on: Manifest) -> Manifest | None:
"""For compact(): the merged output was computed from a specific
snapshot of SSTs (`based_on`). If someone else committed in the
meantime, that output may already be missing data. It can't be
patched by appending, only discarded. Returns None on conflict
so the caller redoes the merge from scratch, rather than risking
a manifest that silently drops or resurrects files."""
candidate = Manifest(based_on.version + 1, sst_ids)
try:
self.store.put_if_absent(
self._manifest_key(candidate.version), candidate.to_bytes()
)
self._cached_manifest = candidate
return candidate
except ConditionalWriteFailed:
self._cached_manifest = None
return None
def put(self, key: str, value: str | None) -> None:
self.memtable[key] = value # None = delete
def flush(self) -> None:
"""The entire durability cost of a batch: one PUT for the SST,
one CAS for the manifest. Compare that to a local WAL paying a
network-grade fsync on every single write (this is the whole
reason batching stopped being optional)."""
if not self.memtable:
return
sst_id = f"sst-{int(time.time() * 1_000_000)}"
sst = SSTable(sst_id, sorted(self.memtable.items()))
self.store.put(f"{self.table}/data/{sst_id}.json", sst.to_bytes())
self.local_cache[sst_id] = sst
self._commit_append([sst_id])
self.memtable.clear()
def get(self, key: str) -> str | None:
if key in self.memtable:
return self.memtable[key]
manifest = self._current_manifest()
for sst_id in reversed(manifest.sst_ids): # newest first
sst = self.local_cache.get(sst_id)
if sst is None:
data = self.store.get(f"{self.table}/data/{sst_id}.json")
sst = SSTable.from_bytes(sst_id, data)
self.local_cache[sst_id] = sst
if key in dict(sst.entries):
return sst.get(key)
return None
def compact(self, retries: int = 5) -> None:
"""Merge, write once, swap the manifest. A reader never sees a
half-merged table, only the manifest before this call, or after.
Note this can't reuse flush's retry strategy. Flush's new SST is
independent of whatever else lands first, so replaying it on top
of the latest version is safe. Compact's merged SST is a snapshot
of a specific set of inputs. If a conflicting write landed in
between, that merge might already be missing an SST's worth of
data, or about to make an already-superseded one look live again.
Patching the manifest instead of redoing the merge is how a
compaction can silently resurrect files it just made obsolete."""
for _ in range(retries):
manifest = self._current_manifest()
if len(manifest.sst_ids) < 2:
return
merged: dict[str, str | None] = {}
for sst_id in manifest.sst_ids: # oldest to newest, newer wins
sst = self.local_cache.get(sst_id) or SSTable.from_bytes(
sst_id, self.store.get(f"{self.table}/data/{sst_id}.json")
)
merged.update(dict(sst.entries))
new_id = f"sst-compacted-{int(time.time() * 1_000_000)}"
new_sst = SSTable(new_id, sorted(merged.items()))
self.store.put(f"{self.table}/data/{new_id}.json", new_sst.to_bytes())
self.local_cache[new_id] = new_sst
if self._commit_replace([new_id], based_on=manifest) is not None:
return
# someone else committed first. this merge is stale, redo it
del self.local_cache[new_id]
raise RuntimeError("compaction did not converge. contention too high")
# old SSTs are now garbage, unreferenced, but still sitting in
# S3 until something is confident no reader still needs them

Now let’s execute a simple example to see how it works. If everything goes as expected we will see the number ’43’ listed twice.

store = FakeObjectStore(put_latency_ms=20)
lsm = ObjectStoreLSM(store)
lsm.put("alice", "42")
lsm.put("bob", "17")
lsm.flush() # one PUT, one CAS (that's the whole commit)
lsm.put("alice", "43") # overwrite, still just sitting in memory
lsm.put("carol", "9")
lsm.flush() # manifest now references two SSTs
print(lsm.get("alice")) # "43" -> the newer SST wins
lsm.compact() # merge both, swap the manifest atomically
print(lsm.get("alice")) # still "43", now from one merged SST

Two methods carry the entire idea:

  • flush, which turns “durable” from a per-write cost into a per-batch one
  • the pair of commit strategies that replace fsync-then-rename with compare-and-swap-then-retry

“Retry on conflict” isn’t a single reusable pattern, it depends on whether the operation retrying is additive (safe to replay on top of whatever won), or a function of a specific snapshot (unsafe to replay, has to be redone). Sorted runs, tombstones, newest-wins reads, merge-based compaction, all of it was present in the RocksDB approach. The object-storage-native part is entirely contained in how visibility gets established, not in the data structure.

Once the write path is solved, what’s left is making the read path fast, and that’s where Arrow, Flight SQL, and multi-tier caching actually earn their place as answers to problems the manifest-and-immutable-objects design creates on the read side.

Immutable SSTs mean a reader fetches whole objects or byte ranges from S3 constantly, so whatever format those objects are stored in had better not cost us a deserialisation pass on every fetch. That’s what Arrow buys: a columnar layout specified exactly enough, down to the byte, that a process can operate on a block pulled straight off the wire without constructing row objects first. Flight extends that further, its wire format is the in-memory format, so shipping a batch of results to a client skips the usual serialise-deserialise-reserialise round trip entirely.

And because every SST is now a network fetch away instead of a disk seek away, the cache in front of it has to be shaped for the access pattern that actually dominates, range scans, not point lookups. A three-tier hierarchy, RAM block cache, local NVMe as a cache of raw S3 bytes, and S3 itself as the source of truth is the same shape a buffer pool always had. What changes is the eviction policy: a point-lookup cache scores blocks mostly by recency, because a miss costs roughly the same either way. A range-scan cache has to score by how expensive a given fetch was relative to its size, and prefetch ahead of the scan cursor, because S3 rewards large sequential GETs far more than it rewards many small ones. Get that scoring wrong and every cache miss becomes a synchronous network stall sitting directly on our p99, no amount of clean manifest design upstream saves us from that.

None of this is a new architectural idea, though. It’s engineering effort spent making the consequences of “the filesystem is gone” fast, once we have already accepted the one substitution that made the whole thing possible in the first place.

Note on the toy engine: It skips a real local WAL for the sub-flush window (something still has to survive a crash between “written to the memtable” and “SST flushed”), garbage collection of orphaned SSTs after compaction, and Bloom filters, all things a production SlateDB or Iceberg deployment can’t skip. None of them change the substitution this whole post has been trying to explain for: an object-storage-native LSM tree is a normal LSM tree with the filesystem replaced by immutable objects and a manifest CAS. Everything harder than that is just making that one idea fast enough to matter.

How Object-Storage-Native LSM-Trees Work Under the Hood

Implementing Durable Execution

Reading and writing about topics we are learning is great, but there is nothing better than some hands-on approach, as such, let’s build a couple implementations of the pizza example described in yesterday’s article.

The first implementation is using the Temporal SDK to allow us to get more familiar with the details of Durable Execution, and consolidate a bit better what we are reviewing. In the second example, we will try to implement our incredible tiny very reduced version of the whole thing.

First example: Using the Temporal SDK

The full implementation of this example can be found in GitHub in the repository pizza-durable-execution. The code and the repository have been heavily documented, which what I think should be enough information to understand the example. But some quick overview is:

  • PizzaActivities: Activities are the side-effecting operations in Durable Execution.
  • PizzaActivitiesImpl: Concrete implementation of the activities.
  • PizzaOrderStarter: Starts a new pizza order workflow instance.
  • PizzaOrderWorkflow: The workflow interface defines the durable process contract.
  • PizzaOrderWorkflowImpl: The durable workflow implementation with pure orchestration, zero side effects.
  • PizzaWorker: The Worker process.

Once we run it, we should see something like:

The pizza worker

=================================================
Pizza Worker started. Polling: pizza-order-queue
Now run PizzaOrderStarter to place an order.
=================================================
10:50:09.222 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Starting pizza order for: margherita
10:50:09.247 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Taking order for pizza: margherita
10:50:09.247 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Order created → ORDER-MARGHERITA
10:50:09.256 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Order accepted → ORDER-MARGHERITA
10:50:09.259 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Kitchen preparing pizza for order: ORDER-MARGHERITA
10:50:09.260 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Pizza ready → PIZZA-ORDER-MARGHERITA
10:50:09.262 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Pizza prepared → PIZZA-ORDER-MARGHERITA
10:50:09.262 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Waiting 5 seconds for delivery window (durable timer)...
10:50:14.291 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Dispatching delivery for: PIZZA-ORDER-MARGHERITA
10:50:14.292 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Delivery confirmed → DELIVERED-PIZZA-ORDER-MARGHERITA
10:50:14.297 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Pizza delivered → DELIVERED-PIZZA-ORDER-MARGHERITA
10:50:14.301 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Sending receipt for delivery: DELIVERED-PIZZA-ORDER-MARGHERITA
10:50:14.301 [Activity Executor taskQueue="pizza-order-queue", namespace="default": 1] INFO d.b.pizza.PizzaActivitiesImpl - [ACTIVITY] Receipt sent. Workflow complete.
10:50:14.305 [workflow-method-pizza-order-margherita-001-019e3031-78f6-7222-9627-e9f11a3367de] INFO d.b.pizza.PizzaOrderWorkflowImpl - [WORKFLOW] Workflow complete for order: ORDER-MARGHERITA

The pizza order starter

=================================================
Starting pizza order workflow...
=================================================
=================================================
Workflow finished. Pizza delivered!
=================================================

We can check the Temporal UI, and see our execution:

All necessary instructions for running it are present in the README of the project.

Second example: Implementing our own

The full implementation of this example can be found in GitHub in the repository mini-durable-execution-platform. The code and the repository have been heavily documented, which what I think should be enough information to understand the example. But some quick overview is:

  • MiniTemporal: Single class containing the whole project.
  • PizzaWorkflow: Durable orchestration of a pizza order.
  • WorkflowContext: The replay engine: the heart of durable execution.

Once we run it, we should see something like:

The pizza worker

worker started
[EXECUTING] prepare-dough
[EXECUTING] add-toppings
[EXECUTING] bake-pizza
[EXECUTING] prepare-dough
[EXECUTING] add-toppings
[EXECUTING] bake-pizza
[EXECUTING] deliver-pizza

The pizza order starter

[WAITING] prepare-dough
[WAITING] add-toppings
[WAITING] bake-pizza
=== JVM CRASH SIMULATED ===
workflowId=d2a8dc13-e428-4bc2-996e-c721d71592f7
...
[WAITING] prepare-dough
[WAITING] add-toppings
[WAITING] bake-pizza
[WAITING] deliver-pizza
===== ORDER COMPLETED =====
dough=dough-ready
toppings=toppings-added
baked=pizza-baked
delivery=pizza-delivered
=== WORKFLOW COMPLETED ===

All necessary instructions for running it are present in the README of the project.

Implementing Durable Execution

The Zero Knowledge Era

This article is going to be a bit controversial. So let me start by saying that I have nothing against AI. I think it is an amazing tool with plenty of use cases where it is useful and helpful, but like many tools before it, it is simply a tool. It is up to us, as professionals, regardless of the field, to decide when to use it and how to use it. Making this decision should be a conscious action based on knowledge and experience.

With that said, I must add that we are taking the wrong approach. Let’s see a few scenarios:

Scenario 1

An engineer is reviewing a pull request created by another engineer to fix a performance problem. While reviewing it, one of the changes looks ‘weird’. At this point, the reviewer decides to ask the author what the logic behind the change is. The reviewer wants to know why they think the change will offer better performance and solve the problem. Surprisingly, the author responds with ‘I don’t know, the AI suggested that’.

Scenario 2

An engineer is digging into a project and finds some code belonging to a not-very-popular framework. While the engineer is not very experienced with this particular framework, they know that their teammates have been battling with it for some time now. They turn around and ask the rest of the team for help. Most of them look at the code and come back with ‘I have no idea’, but one of them provides what looks like a very solid answer. As a follow-up, and trying to learn a bit more, the initial engineer asks some further questions and wonders whether references to documentation can be provided. At some point in that conversation, the second engineer ends up admitting that they have no idea either; the first response was simply what the AI told them.

Scenario 3

Two engineers have been working together for a very long time. After all this time, they know each other and they know their code styles: how each one structures code, what constructions they favour, and so on. One of them, while reviewing a PR, realises that the style in which the code is written deviates from what their colleague usually writes. Additionally, they see some constructions that they have never seen in the codebase, such as some ‘clever’ bitwise logic. After thinking hard to understand the logic, the reviewer realises that some edge cases are not covered. With that information, they go to the author and ask about it. The author replies with something similar to: ‘I have no idea; the AI wrote it, and the code looks elegant and efficient.’

Individually, these seem harmless. Collectively, they point to something more concerning.

I am sure that if you work regularly with other engineers, you will recognise some, if not all, of those scenarios. And that is the problem that I am finding lately: people applying modifications or creating new code in complex and critical codebases without having an understanding of what they are doing, just trusting AI responses without double-checking why the response was suggested, or why something was implemented in this or that way.

As I said at the beginning of the article, AI is a fantastic tool. If you want to implement scripts, one-off tools, or anything you do not care about how it was built, just about the final result, you can do in hours what used to take days. But I think we need to have more discipline when we are modifying codebases that run in production, that are complex and critical, that need troubleshooting at 3 a.m. when we are on call. Making changes without understanding does not seem like the right way to go, or to survive in the long run.

Maybe, one day, AI will be reliable and trustworthy enough to write code without human supervision, AI reviews, for example, for mission-critical projects, but until we get there, we need to keep humans in the loop. And not only to push the ‘Approved’ button to comply with SOC2 requirements, but making the effort to understand what we submit for review and what we review.

It seems that the pressure to be more productive, and especially the FOMO (fear of missing out) is pushing us to be less effective, less disciplined, less knowledgeable. How long will it take for a project to turn into a beast that can only be modified by using AI because no one knows anymore what is under the hood? How long will it take to troubleshoot it when the AI cannot do it, which is not uncommon nowadays?

Let me be clear: the problem is not AI; it is engineers outsourcing understanding. It is engineers pushing changes without understanding what they are pushing. This is why a ‘zero-knowledge era’ is emerging, an environment where code is written, modified, and deployed without anyone fully understanding it, and where systems continue to function until they suddenly don’t, unless we stop it.

What do you think?

The Zero Knowledge Era

Think in Tradeoffs, Not Best Practices

“Best practices” is one of the most popular phrases in software engineering, and also one of the most misleading. It carries an air of safety and responsibility, suggesting that difficult decisions have already been settled elsewhere by wiser people through experience, or communities of people by consensus, or technical maturity over time, and that a careful team only needs to identify the correct practice and apply it consistently.

Sometimes that assumption holds. There are areas of software where reinvention is wasteful, where certain defaults are demonstrably safer, and where repeated failure has already taught the lessons worth preserving. Some habits are justified often enough that ignoring them is simply inefficient. But the phrase becomes dangerous when it obscures the real nature of engineering work. Most meaningful decisions in software are not about selecting a “best” practice in isolation; they are about choosing a trade-off within a specific context. And context changes everything.

The right testing strategy depends on risk, team size, system shape, and release pressure. Architecture is shaped as much by domain complexity and operational maturity as by any abstract principle. Delivery processes reflect failure cost, regulatory expectations, rollback capability, and trust in automation. Even practices that seem straightforward, such as code review, abstraction, documentation, or service decomposition, shift in value depending on environment and consequence. This is why experienced engineers grow cautious around advice that sounds universal. Not because experience is unhelpful, but because it reveals where general guidance stops being general.

The idea of best practices exists for a reason. Engineering teams cannot rediscover every lesson from first principles; they need shared heuristics, conventions, and defaults that work well often enough to reduce unnecessary debate. In that sense, many so-called best practices are simply compressed experience. Advice such as validating inputs, using version control properly, automating builds and tests, avoiding hardcoded secrets, keeping dependencies updated, reviewing production changes carefully, monitoring systems, and limiting privilege is broadly sound. Much of it is essential.

The problem begins when this compressed experience is mistaken for complete reasoning. A principle can be widely useful and still be applied poorly if the team stops asking what it is trying to achieve, what assumptions the practice depends on, and what costs it introduces in a particular situation. Guidance is most valuable when it supports thought; it becomes harmful when it replaces it.

At the heart of this is a simple reality: every non-trivial engineering decision buys something and costs something. More abstraction may improve reuse but reduce clarity. More process may reduce accidental risk but slow change. More services may increase team autonomy while introducing operational complexity. More tests can improve confidence while adding maintenance overhead. Stronger security controls reduce exposure but often introduce friction and recovery costs. Flexibility can reduce lock-in but increase design burden.

This is not a flaw in engineering; it is the work itself. Good engineers learn to evaluate decisions in terms of consequences: what is gained, what becomes more difficult, who benefits, who pays later, and what must remain true for the decision to continue working well. These questions tend to be more valuable than asking whether a practice is modern, popular, or widely recommended. A best practice usually captures a remembered benefit; a trade-off analysis accounts for the cost as well.

One of the more subtle mistakes teams make is confusing “good in general” with “right now”. Strong testing discipline is valuable, but a team may still need to decide whether its next hour is better spent increasing unit coverage, fixing a failing deployment pipeline, or addressing a visibility gap that repeatedly causes production uncertainty. Documentation is important, yet not all documentation carries equal value, as probably every engineer has seen; some supports operational continuity, while some quickly becomes stale and adds maintenance noise. Loose coupling is desirable, but pursuing it too early can result in abstractions that serve hypothetical futures that may never arrive, better than present understanding.

The same applies at larger scales. Microservices may eventually be appropriate, but a modular monolith is often the better choice while a team is still clarifying the product and stabilising its delivery practices. Even code review, one of the most widely defended practices, does not deliver equal value in all contexts; its effectiveness depends on risk, team trust, system criticality, and release cadence. A shallow, ritualised review can be less useful than fewer, more deliberate reviews on meaningful changes. The relevant question is not whether a practice is respectable, but whether it represents the best use of time, complexity, and attention in the current situation.

Disagreements in engineering often reveal another limitation of the “best practice” framing: it tends to hide assumptions. Teams can argue passionately while invoking the same language. One group may describe microservices as best practice for scalability, while another argues that simpler monoliths are best practice for maintainability. One engineer may advocate strict test pyramids; another may favour end-to-end verification. One architect may emphasise standardisation; another, team autonomy.

These conflicts are rarely about the practices themselves. They are about the conditions those practices assume: expected scale, number of teams, failure tolerance, regulatory burden, tooling maturity, cost of change, team skill distribution, operational support quality, and the stability of the domain. Once those assumptions are made explicit, disagreements become easier to understand and often easier to resolve. The conversation shifts from competing claims of correctness to differing views of the environment being optimised for. Precise teams therefore spend less time appealing to abstract best practices and more time discussing constraints, risks, and desired outcomes.

What distinguishes a mature engineer is not a longer list of approved practices, but a stronger ability to trace consequences. Questions such as “What operational load will this introduce?”, “What delivery friction will this create?”, “What failures become easier if we relax this control?”, or “What debt becomes more expensive if we take the faster path now?” lead to better decisions than appeals to convention. This way of thinking is slower than slogan-driven decision-making, but far more reliable, and it produces healthier forms of disagreement. Instead of arguing at the level of identity, teams can argue at the level of impact: which risks are reduced, which costs are increased, and whether that exchange is worthwhile.

This clarity also explains why trade-off-aware teams are not necessarily more cautious. In some cases, they move faster than others precisely because they understand which risks are acceptable and which costs are not worth paying. Their speed comes from deliberate choice rather than adherence to fashion.

Another practical test of any practice is whether it can be sustained under ordinary conditions. Much engineering advice sounds compelling in ideal circumstances, but real systems operate under pressure: deadlines, fatigue, incomplete information, and evolving requirements. A review process that collapses under time pressure, a testing strategy that becomes unmanageable as the system grows, or a documentation model that cannot survive team turnover may not be best practice at all. It may simply be aspirational. Good teams therefore optimise not only for technical correctness, but for durability by choosing approaches that remain functional when systems are messy and time is limited. Sustainability is part of technical quality.

There is also a quieter benefit to thinking in trade-offs: it encourages honesty. When teams rely on the language of best practices, they can present decisions as if they were externally validated, borrowing certainty from the industry instead of owning the consequences themselves. Trade-off thinking removes that cover. It leads to more explicit reasoning: accepting certain risks because delivery speed matters more in a given context, introducing complexity because coordination costs have already become too high, deferring improvements because current failure modes are tolerable, or deliberately avoiding flexibility because the domain is not yet well understood.

This kind of clarity makes decisions easier to revisit and easier for future engineers to understand. It captures not just what was chosen, but why it made sense at the time, which is a far more durable form of knowledge than a claim that something was “best practice”.

Over the course of a technical career, many engineers move from a desire for certainty to a greater appreciation of nuance. Early on, best practices are reassuring; they provide direction and reduce ambiguity. With experience, working through projects, outages, migrations, failed abstractions, and conflicting constraints, confidence in universal answers tends to soften. Ideally, this does not lead to cynicism, but to precision. Experience should widen judgement, not harden it into dogma.

This does not mean that everything is relative or that no principles are worth defending. Some practices are strongly justified, and some trade-offs consistently favour one side. The difference is that experienced engineers tend to understand the boundary conditions more clearly: when a principle holds, when an exception is dangerous, and when competing concerns deserve more weight than usual. Trade-off thinking is not an excuse for vagueness; it is a discipline that requires attention to consequences, constraints, and priorities.

In practice, best practices remain useful as starting points. They help teams prevent avoidable mistakes, preserve lessons that should not need to be relearned, and provide shared defaults that reduce chaos. But they are not substitutes for engineering judgement. Good engineers do not ignore them; they interrogate them. They ask what a practice is protecting, what it costs, what assumptions it carries, and whether those assumptions hold in the system in front of them.

Software engineering is not a search for approved answers. It is a discipline of constrained choices, where every meaningful improvement competes with costs in complexity, speed, flexibility, or operational burden. The teams that understand this tend to build better systems, not because they know more slogans, but because they know how to think in trade-offs.

Think in Tradeoffs, Not Best Practices

Resilience. Keep Distributed Systems Alive

Talk to enough backend engineers and you will eventually hear some version of this story:

“Nothing actually broke. Everything just got slower… until it stopped working.”

Distributed systems rarely fail with a bang. A service times out, clients retry, queues fill, latency spreads, and suddenly the entire platform behaves like a crowded motorway where every driver keeps tapping the brakes.

What’s striking is not that this happens, it’s that many engineers building production systems have never been formally introduced to the ideas designed to prevent it. Terms like exponential backoff, circuit breaker, bulkhead, token bucket, or load shedding sound esoteric, even though they describe mechanisms as fundamental as memory management or indexing. These are not implementation details. They are the control theory of modern software.

And as AI makes it trivial to generate functioning services, this kind of systems thinking is becoming the real differentiator between software that works and software that survives. And, probably, the difference between engineers who design durable systems and those who unknowingly ship fragile ones.

In traditional software, failure was often discrete. A process crashed, a machine went offline, a database corrupted. You debugged, fixed, restarted (the good old times).

Cloud-native systems introduce an entirely new class of failure modes. They are alive with partial availability:

  • A dependency slows but does not fail
  • A region degrades but still responds
  • Requests succeed… just too slowly
  • Retries amplify load
  • Healthy components become collateral damage

This phenomenon is explored deeply in works like Release It! and Designing Data-Intensive Applications, but many engineers encounter it only during their first major incident. The core danger is not failure itself. It is uncontrolled reaction to failure. The following ideas didn’t emerge from theory. They emerged from postmortems on systems that failed in exactly these ways.

Exponential Backoff

Let’s imagine we are on-call, and a service call times out. In this scenario, the most instinctive answer is to try again. While it is not wrong, it can be incomplete. If thousands of clients retry immediately, the struggling service receives a sudden surge of new requests precisely when it is least capable of handling them. The system is not recovering; it is being hammered.

This is where exponential backoff enters the picture. The idea is simple: the more failures you observe, the longer you wait before trying again. Crucially, different callers wait for different lengths of time, so they don’t all stampede back at once. Conceptually, this mirrors real-world congestion control. When traffic jams form, metered ramps and staggered entry prevent waves of cars from worsening the blockage.

While the pattern doesn’t fix the underlying issue, it prevents panic from making it worse.

Circuit Breaker

But retries alone cannot solve everything. If a dependency is failing consistently, continuing to call it at all may be wasteful or dangerous.

Borrowed from electrical systems, the idea is almost philosophical: after enough failures, stop trying. Fail fast. Give the system space to recover. Instead of waiting on timeouts that tie up resources, the application immediately returns an error or fallback response. After a cooling-off period, it cautiously tests whether the dependency has recovered. While this behaviour feels counterintuitive because engineers are trained to maximise success rates, in distributed systems, refusing work can be the act that preserves the ability to do any work at all.

Bulkhead Pattern

Even with smart retries and fast failure, trouble in one part of a system can spread through shared resources.

Consider a service that talks to multiple downstream systems. If one of them becomes slow, threads accumulate waiting for responses. Eventually, there are no threads left for anything else, including healthy dependencies.

The bulkhead pattern addresses this by isolating resources. Just as ships are divided into watertight compartments, systems allocate separate pools for different activities. One flooding compartment does not sink the vessel. This principle appears everywhere once you start looking for it: separate queues, isolated worker groups, per-tenant limits, even independent microservices.

Rate-limiting

So far we’ve discussed reactions to failure. But many outages are caused not by faults, but by sheer volume. Every system has a finite processing capacity. When incoming requests exceed that capacity, queues grow, latency spikes, and eventually the system collapses under its own backlog.

Rate-limiting mechanisms enforce a simple rule: requests are allowed at a sustainable pace, with limited tolerance for bursts. Excess traffic is delayed or rejected. This is not just about protecting infrastructure. It’s about fairness and predictability. Without limits, a single noisy client can degrade service for everyone.

Large platforms use these mechanisms not as emergency tools but as everyday traffic shaping, the software equivalent of speed limits and traffic lights.

Load Shedding

Load shedding may be the most counterintuitive pattern of all. When a system is overwhelmed, the instinct is to try harder: spin up more workers, process faster, squeeze every ounce of throughput from the hardware. But beyond a certain point, this effort becomes self-destructive. The system spends more time managing overload than serving useful work.

Load shedding flips the perspective. Instead of attempting to serve everyone poorly, the system deliberately refuses some requests so it can serve others well. Nonessential features may be disabled. Expensive operations deferred. Low-priority traffic rejected.

Airlines do this. Power grids do this. Even the human body does this under stress. Graceful degradation is not failure. It is survival.

Individually, each technique addresses a specific problem. Together, they express a deeper principle: distributed systems must regulate themselves under stress. One pattern slows demand. Another isolates damage. Another prevents futile work. Another enforces fairness. Another sacrifices noncritical functionality to preserve core operations. Seen this way, resilience engineering begins to resemble ecology or economics more than programming. You are designing feedback loops, not just writing code.

AI tools can now generate working services in seconds (let’s not go deeper in this assessment). They can scaffold APIs, configure deployments, and even suggest architecture diagrams. What they do not yet do reliably is reason about emergent behavior under failure. As software creation accelerates, two trends emerge:

  • Systems become more interconnected
  • Failure modes become more complex

The bottleneck shifts from writing code to designing systems that remain stable under unpredictable conditions. In that environment, understanding resilience patterns is less like knowing a framework and more like understanding physics. It shapes every design decision, even when invisible. Engineers who internalise these ideas will build platforms that feel calm and dependable. Those who don’t will unknowingly construct systems that work beautifully, right up until they don’t.

Users rarely notice resilience when it works. They only experience its absence. Behind every highly available service is not just redundancy or scaling, but a network of small, deliberate decisions about how the system behaves when things go wrong.

  • Retry: but not too quickly
  • Call dependencies: but not blindly
  • Share resources: but not indiscriminately
  • Accept traffic: but not endlessly
  • Serve features: but not at the cost of survival

These decisions are not implementation details. They are the difference between a platform that collapses under pressure and one that bends without breaking. In the end, resilience is not a component you install. It is a mindset you design into the system from the start, a quiet architecture of restraint, isolation, and controlled imperfection that keeps everything running when the world inevitably catches fire.

Resilience. Keep Distributed Systems Alive

Technical career progression: Code less, think more

As software engineers progress in their careers, the scope of their responsibilities and expected impact changes. From a simple graduate, apprentice or self-taught position, all the way to senior positions, the more engineers grow, the bigger is expected to be the blast radius of their decisions and area of influence. However, as individual contributors (ICs), their mindset does not need to shift much. Engineers can focus on creating, designing and building software. In general, while not always the case, the more experienced are the engineers, the more complexity they can deal with, which usually translates into the impact they are having.

All of this makes sense, the more you learn and experience, the more your capacities grow. This happens, in general, in any other aspect of life and any other field. Of course, this is a very simplistic statement that does not factor in a lot of other important considerations. But, that is a whole other conversation and not pertinent to the present article.

In an ideal world, engineers, if they wish, should be able to stay as ICs all their careers. Unfortunately, in the corporate world, there is this non-written expectation of an ever-growing career progression. Additionally, in the socioeconomic societies we live in, where the costs of life are always growing and we have constant changes in the workforce, deciding to stay as ICs is not always feasible.

At this point, when engineers have reached the limits as ICs, they are suggested or pushed to take a step up on the corporate ladder, usually as engineer managers or staff engineers. There are more unorthodox paths such as starting a startup, breaking on your own, or taking a brave jump into ‘head of‘ or C-level roles, but, in general, at this point, any type of growth means to stop being ICs. And, that usually involves a big change of mentality. Some people have a very smooth transition, perhaps they have had great mentors or role models, or they have been surrounded by a great growth environment, but other people have a tough time adapting. The rest of the article is going to focus on some suggestions for the latter, and more specifically on the staff engineer role. However, it is probably applicable to other roles and, even, other fields.

ICs’ goals are tangible, and their impact is easy to see and measure. Their main role is to build systems, design them, extend them with new features, maintain and evolve existing systems, and fix potential problems. All of this generates some sort of deliverable artefacts. The ICs get the satisfaction of delivering one thing, and moving to the next one, the feeling of progression, achieving a goal, and keep moving to achieve the next one. When stepping into a non-IC role, you start working with more abstract deliverables such as deciding the roadmap for a team, creating your own tasks, deciding what is important and their priorities, adding perspective to what a team is building to see the bigger picture and how it affects the company, the business and its clients, and you will be evaluated, and evaluate yourself, based on the impact of your decisions and your influence. There are no concrete deliverables any more, and the feedback loop on your work is completely different. It seems so different that, initially, you will probably feel that you are doing nothing, and you are just losing your time. But, if you are in that position or you feel that way, don’t worry, it’s not your problem, it’s just that you have not shifted your perspective yet. This transition can be challenging, but here are a few strategies to help you adjust:

Shift Your Mindset from Executor to Strategist

  • Focus on outcomes, not just output. You’re used to delivering code and features, but now you need to think about “why” the team is building something and the value it provides to the business. Outcomes like improving customer satisfaction, reducing technical debt, or enabling future scalability are key.
  • Be comfortable with ambiguity. While senior roles often involve making decisions with incomplete information, as a non-IC you will see this only increased. Instead of always looking for the “right” answer, work with your team to define acceptable risks and build mechanisms for course correction.

Start with a High-Level Vision

  • Define long-term goals. Think about where you want the team to be in six months or a year. What systems, processes, or architectures will help achieve those goals? Align your work with the company’s overall objectives and communicate this vision clearly to your team.
  • Develop roadmaps. A roadmap gives direction. Focus on key milestones that align with business priorities, but remain flexible in how to get there. You’ll find it’s not as much about specific tasks, but about defining priorities and ensuring that the team is working on the highest-impact problems.

Learn to Delegate and Empower

  • Trust your team. As a leader, your job now is less about directly writing the code and more about enabling your team to be effective. Delegate tasks and empower others to take ownership, while you focus on removing blockers and setting strategic direction.
  • Coach instead of direct. Help your team members grow by guiding them through decision-making processes instead of giving them answers. Provide context and goals, and let them fill in the technical details.

Master Prioritisation

  • Focus on impact. Not all tasks are equally important. Evaluate them based on impact on the business, technical risk, and long-term benefits. Prioritisation is about balancing short-term needs (e.g., addressing technical debt) with long-term goals (e.g., new feature development).
  • Use frameworks. Tools like OKRs (Objectives and Key Results) or the Eisenhower Matrix (Urgent vs. Important) can help you make decisions that align with the bigger picture.

Communicate Strategically

  • Stakeholder management. Your role may now involve more communication with upper management, product teams, or other non-technical stakeholders. Translate technical concepts into business value. What’s important is how the team’s work contributes to company goals, not just how the system works.
  • Create visibility. Make sure your team’s progress and challenges are visible to leadership in a clear, concise manner. Reports, roadmaps, and updates should focus on outcomes rather than low-level details.

Develop Abstract Thinking Skills

  • See the forest, not just the trees. Abstract thinking involves stepping back from the day-to-day coding details and viewing the system holistically. Consider scalability, maintainability, and the business problems the system is solving.
  • Stay curious. Think about “why” a particular feature or task is important, and how it fits into the larger ecosystem. Are there other stakeholders or systems impacted by the team’s work? Are you solving the right problems?

Get Comfortable with Measurement Through Influence

  • Measure team effectiveness. Your success is no longer measured by the lines of code or features you deliver. Instead, focus on team metrics: velocity, team satisfaction, alignment with business goals, etc.
  • Focus on alignment and value. Even though the work feels abstract, the results are measurable through customer impact, system stability, and business success. These are now your “deliverables”.

Embracing the leadership mindset will help you move forward confidently., and be fulfilled in to your new role. Transitioning to a more strategic role takes time, but with practice, you’ll find new ways to make an impact.

Technical career progression: Code less, think more

Learning properly

If you are a regular reader of this blog, you know that some, if not most, of my articles are inspired by conversations with less experienced developers or situations in my day-to-day work life. Sometimes, they are about questions that need answering, topics that need discussion, or similar subjects. Today, while the topic is somehow related and fits that criteria, it is more of a broad observation and a compilation of multiple conversations.

All of us, everyone who works with technology, especially in the field of software development, regardless of our background (i.e., university degree, professional degree, boot camp, self-taught), have learned things on our own. This learning capacity, in my opinion, is what differentiates good engineers from great ones.

For a long time, I set my learning goals following two strategies:

  • If I had a problem to solve, I would ask myself how I would solve it and start digging, leading me to learn different things.
  • If there is a new technology I want to understand better, I would think or search for a series of problems, from simpler to more complex, that I can solve with the technology I want to understand.

This is probably what is called “Learning by doing”. I did not discover it, I did not label it; that is just how I learned to do things during my studies. As it was naturally taught to me through conversations with my teachers, attending classes, or even just discussing with my peers, friends, and colleagues, I never thought about going in a different direction or approaching it differently. However, over the last few years, I have seen more and more people utilizing a different learning approach. While on multiple occasions, a lot can be learned from less experienced people – sometimes from their naivety, passion, boldness, or digital nativity – they can bring a lot to the table if you listen to them. But, I think this time they are doing it wrong. Let me explain.

I am talking about their learning methodology. I have observed the tendency to learn exclusively by following tutorials and assume that after finishing one or more, regardless of the length, they know the technology. I agree that some tutorials are not bad; they can give you a brief understanding of the basics of a technology and can guide you in designing your learning path. However, you are not going to truly learn a technology by just following tutorials. Yes, they may provide a vague understanding of the technology, and if that is all you want, that is perfect, but do not lie to yourself – you have not truly learned a technology.

Everything you usually do when you follow a tutorial falls under the umbrella of copy and paste. The tutorial mentors present you with a problem, and they solve it for you. If the teaching skills of these mentors are above average, they will show you errors and problems you may encounter, but still, it is basically copy and paste. I can copy and paste a physics manual, but that does not mean I have learned physics. While following a tutorial, you are not thinking on your own, you are not investigating, and you are not putting in the effort required for that knowledge to solidify in your mind.

Learning by doing should be the goal of everyone who wants to truly learn something beyond a simple understanding of what it is. Yes, it requires time, effort, and persistence, like almost everything we do or learn in life. This is why, with exceptions, people grow through their careers, and this is why there are differences in seniority. More senior people are not in that position because they were born knowing something you do not; it is because they have dealt with countless hours of trying to fix real-life problems.

If you are a less experienced developer, just starting, or even if you have not started but are thinking about it, do not fall into the tutorial trap. If you want to learn something, find a problem you can solve, invent a problem where that technology can be applied, use Google to find one, and try to solve it. Read the documentation, explore blogs where there are discussions similar to what you are trying to do, and think. One or, even, a few tutorials to get an overview are fine, but you should quickly move past that stage. Sometimes you will get stuck for days, and it will be frustrating, but believe me when I say this – all of us have been there, and most of us are still there. Sometimes, I waste eight hours thinking about a problem without finding a solution. I go to bed, and the next morning, it takes me five minutes to find the solution. Do you think I was smarter this morning than yesterday? It was just that my mind had time to deal with the problem, and the solution was the addition of the “wasted” time plus the five minutes of “geniality.” All of it is part of the process.

Your ability to learn new things properly, your capacity for critical thinking, and your skill in breaking down a problem into smaller components are, without any doubt, some of the best skills you can nurture in your career.

Happy learning!

Learning properly

Maintaining compatibility

Most of the companies nowadays are implementing or want to implement architectures based on micro-services. While this can help companies to overcome multiple challenges, it can bring its own new challenges to the table.

In this article, we are going to discuss a very concrete one been maintaining compatibility when we change objects that help us to communicate the different micro-services on our systems. Sometimes, they are called API objects, Data Transfer Objects (DTO) or similar. And, more concretely, we are going to be using Jackson and JSON as a serialisation (marshalling and unmarshalling) mechanism.

There are some other methods, other technologies and other ways to achieve this but, this is just one tool to keep on our belt and be aware of to make informed decisions when faced with this challenge in the future.

Maintaining compatibility is a very broad term, to establish what we are talking about and ensure we are on the same page, let us see a few examples of real situations we are trying to mitigate:

  • To deploy breaking changes when releasing new features or services due to changes on the objects used to communicate the different services. Especially, if at deployment time, there is a small period of time where the old and the new versions are still running (almost impossible to avoid unless you stop both services, deploy them and restart them again).
  • To be forced to have a strict order of deployment for our different services. We should be able to deploy in any order and whenever it best suits the business and the different teams involved.
  • The need of, due to a change in one object in one concrete service, being forced to deploy multiple services not directly involved or affected by the change.
  • Related to the previous point, to be forced to change other services because of a small data or structural change. An example of this would be some objects that travel through different systems been, for example, enriched with extra information and finally shown to the user on the last one.

To exemplify the kind of situation we can find ourselves in, let us take a look at the image below. In this scenario, we have four different services:

  • Service A: It stores some basic user information such as the first name and last name of a user.
  • Service B: It enriches the user information with extra information about the job position of the user.
  • Service C: It adds some extra administrative information such as the number of complaints open against the user.
  • Service D: It finally uses all the information about the user to, for example, calculate some advice based on performance and area of work.

All of this is deployed and working on our production environment using the first version of our User object.

At some point, product managers decided the age field should be considered on the calculations to be able to offer users extra advice based on proximity of retirement. This added requirement is going to create a second version of our User object where the field age is present.

Just a last comment, for simplicity purposes, let us say the communication between services is asynchronous based on queues.

As we can see on the image, in this situation only services A and D should be modified and deployed. This is what we are trying to achieve and what I mean by maintaining compatibility. But, first, let us explore what are the options we have at this point:

  1. Upgrade all services to the second version of the object User before we start sending messages.
  2. Avoid sending the User from service A to service D, send just an id, and perform a call from service D to recover the User information based on the id.
  3. Keep the unknown fields on an object even, if the service processing the message at this point does not know anything about them.
  4. Fail the message, and store it for re-processing until we perform the upgrade to all services involved. This option is not valid on synchronous communications.

Option 1

As we have described, it implies the update of the dependency service-a-user in all the projects. This is possible but it brings quickly some problems to the table:

  • We not only need to update direct dependencies but indirect dependencies too what it can be hard to track, and easy to miss. In addition, a decision needs to be done about what to do when a dependency is missed, should an error be thrown? Should we fail silently?
  • We have a problem with scenarios where we need to roll back a deployment due to something going wrong. Should we roll back everything? Good luck! Should we try to fix the problem while our system is not behaving properly?
  • Heavy refactoring operations or modifications can make upgrades very hard to perform.

Option 2

Instead of sending the object information on the message, we just send an id to be able posteriorly to recover the object information using a REST call. This option while very useful in multiple cases is not exempt from problems:

  • What if, instead of just a couple of enrichers, we have a dozen of them and they need the user information? Should we consolidate all services and send ids for the enriched information crating stores on the enrichers?
  • If, instead of a queue, other mechanisms of communications are used such as RPC, do now all the services need to call service A to recover the User information and do their job? This just creates a cascade of calls.
  • And, under this scenario, we can have inconsistent data if there is any update while the different services are recovering a User.

Option 3

This is going to be the desired option and the one we are going to do a deep dive on this article using Jackson and JSON how to keep the fields even if the processing service does not know everything about them.

To add in advance that, as always, there are no silver bullets, there are problems that not even this solution can solve but it will mitigate most of the ones we have named on previous lines.

One problem we are not able to solve with this approach – especially if your company perform “all at once” releases instead of independent ones – is, if service B, once deployed tries to persist some information on service A before the new version has been deployed, or tries to perform a search using one criterion, in this case, the field age, on the service A. In this scenario, the only thing we can do is to throw an error.

Option 4

This option, especially in asynchronous situations where messages can be stored to be retried later, can be a possible solution to propagate the upgrade. It will slow down our processing capabilities temporarily, and retrying mechanism needs to be in place but, it is doable.

Using Jackson to solve versioning

Renaming a field

Plain and simple, do not do it. Especially, if it is a client-facing API and not an internal one. It will save you a lot of trouble and headaches. Unfortunately, if we are persisting JSON on our databases, this will require some migrations.

If it needs to be done, think about it again. Really, rethink it. If after rethinking it, it needs to be done a few steps need to be taken:

  1. Update the API object with the new field name using @JsonAlias.
  2. Release and update everything using the renamed field, and @JsonAlias for the old field name.
  3. Remove @JsonAlias for the old field name. This is a cleanup step, everything should work after step two.

Removing a field

Like in the previous case, do not do it, or think very hard about it before you do it. Again, if you finally must, a few steps need to be followed.

First, consider to deprecate the field:

If it must be removed:

  1. Explicitly ignore the old property with @JsonIgnoreProperties.
  2. Remove @JsonIgnoreProperties for the old field name.

Unknown fields (adding a field)

Ignoring them

The first option is the simplest one, we do not care for new fields, a rare situation but it can happen. We should just ignore them:

A note of caution in this scenario is that we need tone completely sure we want to ignore all properties. As an example, we can miss on APIs that return errors as HTTP 200 OK, and map the errors on the response if we are not aware of that, while in other circumstances it will just crash making us aware.

Ignoring enums

In a similar way, we can ignore fields, we can ignore enums, or more appropriately, we can map them to an UNKNOWN value.

Keeping them

The most common situation is that we want to keep the fields even if they do not mean anything for the service it is currently processing the object because they will be needed up or downs the stream.

Jackson offers us two interesting annotations:

  • @JsonAnySetter
  • @JsonAnyGetter

These two annotations help us to read and write fields even if I do not know what they are.

class User {
    @JsonAnySetter
    private final Map<String, Object> unknownFields = new LinkedHashMap<>();
    
    private Long id;
    private String firstname;
    private String lastname;

    @JsonAnyGetter
    public Map<String, Object> getUnknownFields() {
        return unknownFields;
    }
}

Keeping enums

In a similar way, we are keeping the fields, we can keep the enums. The best way to achieve that is to map them as strings but leave the getters and setters as the enums.

@JsonAutoDetect(
    fieldVisibility = Visibility.ANY,
    getterVisibility = Visibility.NONE,
    setterVisibility = Visibility.NONE)
class Process {
    private Long id;
    private String state;

    public void setState(State state) {
        this.state = nameOrNull(state);
    }

    public State getState() {
        return nameOrDefault(State.class, state, State.UNKNOWN);
    }

    public String getStateRaw() {
        return state;
    }
}

enum State {
    READY,
    IN_PROGRESS,
    COMPLETED,
    UNKNOWN
}

Worth pointing that the annotation @JsonAutoDetect tells Jackson to ignore the getters and setter and perform the serialisation based on the properties defined.

Unknown types

One of the things Jackson can manage is polymorphism but this implies we need to deal sometimes with unknown types. We have a few options for this:

Error when unknown type

We prepare Jackson to read an deal with known types but it will throw an error when an unknown type is given, been this the default behaviour:

@JsonTypeInfo(
    use = JsonTypeInfo.Id.NAME,
    include = JsonTypeInfo.As.PROPERTY)
@JsonSubTypes({
    @JsonSubTypes.Type(value = SelectionProcess.class, name = "SELECTION_PROCESS"),
})
interface Process {
}

Keeping the new type

In a very similar to what we have done for fields, Jackson allow as to define a default or fallback type when the given type is not found, what put together with out unknown fields previous implementation can solve our problem.

@JsonTypeInfo(
    use = JsonTypeInfo.Id.NAME,
    include = JsonTypeInfo.As.PROPERTY,
    property = "@type",
    defaultImpl = AnyProcess.class)
@JsonSubTypes({
    @JsonSubTypes.Type(value = SelectionProcess.class, name = "SELECTION_PROCESS"),
    @JsonSubTypes.Type(value = SelectionProcess.class, name = "VALIDATION_PROCESS"),
})
interface Process {
    String getType();
}

class AnyProcess implements Process {
    @JsonAnysetter
    private final Map<String, Object> unknownFields = new LinkedHashMap<>();

    @JsonProperty("@type")
    private String type;

    @Override
    public String getType() {
        return type;
    }

    @JsonAnyGetter
    public Map<String, Object> getUnknownFields() {
        return unknownFields
    }
}

And, with all of this, we have decent compatibility implemented, all provided by the Jackson serialisation.

We can go one step further and implement some basic classes with the common code e.g., unknownFields, and make our API objects extend for simplicity, to avoid boilerplate code and use some good practices. Something similar to:

class Compatibility {
    ....
}

class MyApiObject extends Compatibility {
    ...
}

With this, we have a new tool under our belt we can consider and use whenever is necessary.

Maintaining compatibility

Diagrams as Code

I imagine that everyone reading this blog should be, by now, familiar with the term Infrastructure as Code (IaC). If not because we have written a few articles on this blog, probably because it is a widely extended-term nowadays.

At the same time I assume familiarity with the IaC term, I have not heard a lot of people talking about a similar concept called Diagram as Code (DaC) but focus on diagrams. I am someone that, when I arrive at a new environment, finds diagrams very useful to have a general view of how a new system works or, when given explanations to a new joiner. But, at the same time, I must recognise, sometimes, I am not diligent enough to update them or, even worst, I arrived at projects where, or they are too old to make sense or there are none.

The reasons for that can be various, it can be hard for developers to maintain diagrams, lack of time, lack of knowledge on the system, not obvious location of the editable file, only obsolete diagrams available and so on.

I have been playing lately with the DaC concept and, with a little bit of effort, all new practices require some level of it till they are part of the workflow, it can fill this gap and help developers and, other profiles in general, to keep diagrams up to date.

In addition, there are some extra benefits of using these tools such as easy version control, the need for only a text editor to modify the diagram and, the ability to generate the diagrams everywhere, even, as part of our pipelines.

This article is going to cover some of the tools I have found and play with it lately, the goal is to have a brief introduction to the tools, be familiar with their capabilities and restrictions and, try to figure out if they can be introduced on our production environments as a long term practise. And, why not, compare which tool offers the most eye-catching ones, in the end, people are going to pay more attention to these ones.

Just a quick note before we start, all the tools we are going to see are open source tools and freely available. I am not going to explore any payment tool and, I am not involved in any way in the tools we are going to be using.

Graphviz

The first tool, actually a library, we are going to see is Graphviz. In advance, I am going to say, we are not exploring this option in depth because it is too low level for my taste or the purpose we are trying to achieve. It is a great solution if you want to generate diagrams as part of your applications. In fact, some of the higher-level solution we are going to explore use this library to be able to generate the diagrams. With that said, they define themselves as:

Graphviz is open source graph visualization software. Graph visualization is a way of representing structural information as diagrams of abstract graphs and networks. It has important applications in networking, bioinformatics, software engineering, database and web design, machine learning, and in visual interfaces for other technical domains.

The Graphviz layout programs take descriptions of graphs in a simple text language, and make diagrams in useful formats, such as images and SVG for web pages; PDF or Postscript for inclusion in other documents; or display in an interactive graph browser. Graphviz has many useful features for concrete diagrams, such as options for colors, fonts, tabular node layouts, line styles, hyperlinks, and custom shapes.

https://graphviz.org

The installation process in an Ubuntu machine is quite simple using the package manager:

sudo apt install graphviz

As I have said before, we are not going to explore deeper this library, we can check on the documentation for some examples built in C and, the use of the command line tools it offers but, for my taste, it is a too low-level solution to directly use it when implementing DaC practises.

PlantUML

The next tool is PlantUML. Regardless of what the name seems to imply, the tool is able to generate multiple types of diagrams, all of them listed on their web page. Some examples are:

The diagrams are defined using a simple and intuitive language and, images can be generated in PNG, SVG or, LaTeX format.

They provide an online tool it can be used for evaluation and testing purposes but, in this case, we are going to be installing it locally, especially to test how hard it is and how integrable is in our CI/CD pipelines as part of our evaluation. By the way, this is one of the tools it uses Graphviz behind the scenes.

There is no installation process, the tool just needs the download of the corresponding JAR file from their download page.

Now, let’s use it. We are going to generate a State Diagram. It is just one of the examples we can find on the tool’s page. The code used to generate the example diagram is the one that follows and is going to be stored in a TXT file:

@startuml
scale 600 width

[*] -> State1
State1 --> State2 : Succeeded
State1 --> [*] : Aborted
State2 --> State3 : Succeeded
State2 --> [*] : Aborted
state State3 {
  state "Accumulate Enough Data\nLong State Name" as long1
  long1 : Just a test
  [*] --> long1
  long1 --> long1 : New Data
  long1 --> ProcessData : Enough Data
}
State3 --> State3 : Failed
State3 --> [*] : Succeeded / Save Result
State3 --> [*] : Aborted

@enduml

Now, let’s generate the diagram:

java -jar ~/tools/plantuml.jar plantuml-state.txt

The result is going to be something like:

State Diagram generated with PlantUML

As we can see, the result is pretty good, especially if we consider that it has taken us around five minutes to write the code without knowing the syntax and we have not had to deal with positioning the element on a drag and drop screen.

Taking a look at the examples, the diagrams are not particularly beautiful but, the easy use of the tool and the variety of diagrams supported makes this tool a good candidate for further exploration.

WebSequenceDiagrams

WebSequenceDiagrams is just a web page that allows us to create in a quick and simple way Sequence Diagrams. It has some advantages such as offering multiple colours, there is no need to install anything and, having only one purpose, it covers it quite well in a simple way.

We are not going to explore this option further because it does not cover our needs, we want more variety of diagrams and, it does not seem integrable on our daily routines and CI/CD pipelines.

Asciidoctor Diagram

I assume everyone is more or less aware of the existence of the Asciidoctor project. The project is a fast, open-source text processor and publishing toolchain for converting AsciiDoc content to HTML5, DocBook, PDF, and other formats.

Asccidoctor Diagram is a set of Asciidoctor extensions that enable you to add diagrams, which you describe using plain text, to your AsciiDoc document.

The installation of the extension is quite simple, just a basic RubyGem that can be installed following the standard way.

gem install asciidoctor-diagram

There are other options of usage but, we are going to do an example using the terminal and, using the PlantUML syntax we have already seen.

[plantuml, Asciidoctor-classes, png]     
....
class BlockProcessor
class DiagramBlock
class DitaaBlock
class PlantUmlBlock

BlockProcessor <|-- DiagramBlock
DiagramBlock <|-- DitaaBlock
DiagramBlock <|-- PlantUmlBlock
....

The result been something like:

Generated with Asciidoctor Diagrams extension

One of the advantages of this tool is that it supports multiple diagram types. As we can see, we have used the PlantUML syntax but, there are many more available. Check the documentation.

Another of the advantages is it is based on Asciidoctor that is a very well known tool and, in addition to the image it generates an HTML page with extra content if desired. Seems worth it for further exploration.

Structurizr

I was going to skip this one because, despite having a free option, requires some subscription for determinate features and, besides, it does not seem as easy to integrate and use as other tools we are seeing.

Despite all of this, I thought it was worth it to mention it due to the demo page they offer where, with just some clicking, you can see the diagram expressed on different syntaxes such as PlantUML or WebSequenceDiagrams.

Diagrams

Diagrams is a tool that seems to have been implemented explicitly to follow the Diagram as Code practice focus on infrastructure. It allows you to write diagrams using Python and, in addition to support and having nice images for the main cloud providers, it allows you to fetch non-available images to use them in your diagrams.

Installation can be done using any of the available common mechanism in Python, in our case, pip3.

pip3 install diagrams

This is another one of the tools that, behind the scenes, uses Graphviz to do its job.

Let’s create our diagram now:

from diagrams import Cluster, Diagram
from diagrams.onprem.analytics import Spark
from diagrams.onprem.compute import Server
from diagrams.onprem.database import PostgreSQL
from diagrams.onprem.inmemory import Redis
from diagrams.onprem.aggregator import Fluentd
from diagrams.onprem.monitoring import Grafana, Prometheus
from diagrams.onprem.network import Nginx
from diagrams.onprem.queue import Kafka

with Diagram("Advanced Web Service with On-Premise", show=False):
    ingress = Nginx("ingress")

    metrics = Prometheus("metric")
    metrics << Grafana("monitoring")

    with Cluster("Service Cluster"):
        grpcsvc = [
            Server("grpc1"),
            Server("grpc2"),
            Server("grpc3")]

    with Cluster("Sessions HA"):
        master = Redis("session")
        master - Redis("replica") << metrics
        grpcsvc >> master

    with Cluster("Database HA"):
        master = PostgreSQL("users")
        master - PostgreSQL("slave") << metrics
        grpcsvc >> master

    aggregator = Fluentd("logging")
    aggregator >> Kafka("stream") >> Spark("analytics")

    ingress >> grpcsvc >> aggregator

And, let’s generate the diagram:

python3 diagrams-web-service.py

With that, the result is something like:

Diagram generated with Diagrams

As we can see, it is easy to understand and, the best part, it is quite eye-catching. And, everything looks in place without the need to mess with a drag and drop tool to position our elements.

Conclusion

As always, we need to evaluate which tool is the one that best fit our use case but, after seeing a few of them, my conclusions are:

  • If I need to generate infrastructure diagrams I will go with the Diagrams tools. Seems very easy to use been based on Python and, the results are very visually appealing.
  • For any other type of diagram, I will be inclined to use PlantUML. It seems to support a big deal of diagram types and, despite not being the most beautiful ones, it seems the results can be clear and useful enough.

Asciidoctor Diagrams seems a good option if your team or organisation is already using Asciidoctor and, it seems a good option if we want something else than just a diagram generated.

Diagrams as Code