All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
- Metrics integrations —
rmcache-micrometer(MicrometerCacheMeterBinder) andrmcache-opentelemetry(OpenTelemetry observable instruments) expose cache statistics with zero hot-path cost (pull-based; readgetStats()only on the collection/export interval).rmcache-metricsaddsMeteredOffHeapCache, an opt-in latency-sampling decorator. See docs/metrics.md. - JCache (JSR-107) provider —
rmcache-jcacheis a standardjavax.cacheprovider (drop-in for Spring Cache / Hibernate L2): store-by-value, atomicinvokevia per-key striped locks,ExpiryPolicy→TTL, and JMX statistics. See docs/jcache.md. - Multi-module Maven Central publishing — all modules (
rmcache-metrics,-micrometer,-opentelemetry,-jcache) are now signed and published alongside the core in a single Central Portal deployment bundle (gradle/maven-publish-conventions.gradle). - Peer benchmark suite —
FairComparisonScaleBenchmarknow compares RMCache (plain + OFF_HEAP GhostCache) against Caffeine (on-heap reference), Chronicle Map, OHC, MapDB, and EhCache in one internally-consistent run; newOHCComparisonBenchmarkisolates OHC'sUnsafe-based allocator. Results published in the README. - Tail-latency benchmark —
TailLatencyBenchmark(JMHSampleTime) reports p50/p90/p99/p99.9 for GET and PUT at 1M entries; the README shows RMCache's GET tail beating even on-heap Caffeine (off-heap means no GC jitter). - Open-source governance —
NOTICE,CLA.md(Contributor License Agreement enabling the open-core model), GitHub issue/PR templates,CODEOWNERS, and Dependabot configuration. examples/subproject — 5 runnable examples:BasicCacheExample,TTLExample,ZeroCopyExample,EvictionExample,CustomSerializerExample. Run via./gradlew :examples:run<Name>.- JaCoCo coverage reporting —
jacocoTestReporttask (HTML + XML) with 80% instruction coverage minimum (jacocoTestCoverageVerification). Current baseline: ~88.6% instruction, ~87.0% line. - MapDB and ChronicleMap as benchmark competitors — replace NMA in
FairComparisonScaleBenchmark,ComparativeWorkloadBenchmark,ThroughputBenchmark, andMemoryScalabilitySuite. asyncExecutor(Executor)builder option — custom bounded executor forputAsync/getAsync(replaces unboundedForkJoinPool.commonPool())- Per-cause eviction counters —
evictionsBySize,evictionsByTtl,evictionsByExplicitinCacheStatsrecord, backed byLongAdder module-info.java— explicit Java module declaration (com.codeabbot.rmcache) with documented exports, unexported internals, and test reflection contractzeroHeapProfile()Javadoc — precise documentation of what the preset eliminates and what it does not eliminate from the heapSECURITY.md— vulnerability reporting process and security modelARCHITECTURE.mdrewrite — merged TECHNICAL.md content, corrected entry header layout (ISSUE-017), updated file index, added design decisions tableARCHITECTURE-DEEP-DIVE.md— complete configuration reference and developer guide (replacesDEVELOPER.md)docs/getting-started.md— quickstart, patterns, sizing guidedocs/zero-copy-access.md—getZeroCopy/getViewsafety constraints and usage patternsdocs/eviction-policies.md— LRU, TTL, composite, eviction listener/filter referencedocs/custom-serialization.md— custom serializer implementations, segment serializer, framework integrationsdocs/heap-profile.md— heap breakdown, zero-heap profiles, 1B-entry scale projection
- Build dependency management — every dependency now declared through the Gradle version catalog (
gradle/libs.versions.toml); JUnit (Jupiter + Platform) versions aligned via the JUnit BOM to prevent drift. - JDK requirement raised to 25+ (LTS) — build config, CI matrix, all documentation updated
entryCountrename —_sizefield renamed toentryCountinLRUPolicyandTTLPolicyfor readabilityBackgroundEvictionTest— replaced busy-sleep (20×10ms polling) with deadline-based polling (2s window, 5ms sleep)
- Close lifecycle race —
OffHeapCacheImpl.close()now closes the eviction policy before freeing native resources, preventing the LRU maintenance thread from touching unmapped off-heap segments during shutdown. - Close idempotency — repeated
close()calls now return immediately after the first shutdown, preventing double-free of native segments. CacheValueView.isValid()— now detects views whose entry slot has already been removed or evicted before the call; Javadoc clarifies the remaining slot-reuse and concurrent-read limits.- Broken
byte[]examples — public quickstarts and docs now configure.forByteArrayValues()instead of relying on the default string value serializer. - GhostCache AUTO default —
GhostCacheMode.AUTOnow resolves toOFF_HEAP, so default caches keep the L1 shortcut off-heap unlessHEAPis explicitly requested. maxEntriesresidency cap. With background eviction enabled, the async drain could lag behind a write burst and let the cache grow toward the memory limit instead ofmaxEntries(~6× overshoot observed at small caps). The new-key insert path now also evicts synchronously while over the cap, somaxEntriesbounds steady-state residency — a convergent cap (transient overshoot under concurrent bursts is expected and documented, as in Caffeine), not a hard per-instant limit. Eviction quality (hit rate) verified on par with Caffeine's W-TinyLFU.- Put failure is observable, not silent. A
put,putIfAbsent, orcomputeIfAbsentthat cannot allocate under memory pressure is no longer counted as a successful put; it increments the newCacheStats.rejectedPuts()counter. Micrometer and OpenTelemetry also exportcache.puts.rejected. A bounded cache may decline an entry — this is not an error and does not throw, preserving lossy-cache semantics while keeping theputsstat honest. - HEAP ghost staleness on update — an in-place value update now invalidates the on-heap (
GhostCacheMode.HEAP) L1 entry, so a subsequentget()can no longer return the stale pre-update value. The defaultOFF_HEAPghost was already correct. LRUPolicy.close()shutdown race — close now force-stops (shutdownNow) and re-awaits the maintenance thread if it does not terminate within the grace window, before freeing native shards. Exportedevictionpolicies (LRUPolicy/TTLPolicyand their off-heap structures) are also idempotent on directclose().- Custom segment serializer documented as a trusted extension —
SegmentValueSerializerwrites directly to native memory with no per-write bounds check by default (zero-copy fast path; implementations must honormaxLen). NewCacheBuilder.strictSegmentSerializerBounds(true)opts custom serializers into amaxLen-bounded slice (over-write throws instead of corrupting memory) for development/untrusted use; built-in serializers and the commonbyte[]PUT path are unchanged. - Gradle 10 readiness — replaced the deprecated
required { … }signing assignment withrequired = { … }, clearing the space-assignment deprecation. - Docs accuracy — documented the
close()/maxEntries/ resize / ABA concurrency contracts (ARCHITECTURE.md §11); corrected the TTL docs (per-entryput(…, Duration)is the supported path); fixed the packed allocation-handle layout (40-bit offset, not 48-bit); JCache mentions now flagged Phase 1. - ISSUE-017: Entry header alignment — swapped
slotIdandexpiresAtpositions soexpiresAtis at offset 8 (8-byte aligned). ReplacedUNALIGNED_LONGwithValueLayout.JAVA_LONG. 17–37% latency improvement across all benchmark scales. - ISSUE-014:
OffHeapFrequencySketchthread safety — replaced plain read/write withVarHandle.getVolatile+ single-attemptcompareAndExchange. Eliminates lost increments under concurrent access. - ISSUE-018: CAS double-free guard — replaced volatile-read + free in
SubPool.free()with a compare-and-swap. Prevents concurrent double-free corrupting the slab allocator. - ISSUE-015: Duplicate LRU insert guard — added
getSegment(slot) != NONEguard in all threeaddTo*()methods ofOffHeapCompactLRU. Prevents list corruption from duplicate inserts during concurrent access. - ISSUE-019:
RetiredSegmenttyped record — replacedObject[]retired segment entries with aprivate record RetiredSegment(long retiredAtNanos, MemorySegment segment). Eliminates unsafe casts and clarifies intent. - ISSUE-020:
RETIRED_GRACE_NANOSconstant — extracted hardcoded500_000_000Lto named constantRETIRED_GRACE_NANOS = 2_000_000_000L(2 seconds). Safer grace period for concurrent optimistic readers during resize. - ISSUE-011: Murmur-style hash spread — applied
h ^= h>>>16; h *= 0x85ebca6b; h ^= h>>>13at all 5keyHashcomputation sites. Prevents clustering on low-entropy keys (e.g., sequential integers). - Deleted dead code — removed
CacheContext.java(outer public class, never used — shadowed by private inner class),CacheStatistics.java(public record, never used — superseded byOffHeapCache.CacheStats), andTimingWheel.java(package-private, never instantiated — superseded byOffHeapTimingWheel) - Removed NMA dependency — removed
com.target:native-memory-allocatorfrom all benchmark code; replaced with MapDB and ChronicleMap - Fixed Javadoc errors —
<=,<<HTML escaping and heading hierarchy inAllocationHandle,OffHeapGhostCache,SegmentValueSerializer,ValueWriter,OffHeapTimingWheel,CacheBuilder
Hot-path read/write optimizations (work-removing — no path does more than before):
- No-TTL read fast path —
get/getView/getZeroCopyskip the per-entry expiry check entirely until a TTL (eviction policy or per-entry) is first used, so caches that never use TTL pay nothing for expiry on the read path. - Skip redundant ghost write on hit — an off-heap-ghost GET hit no longer re-writes the
(hash, slot)mapping it just read. putIfAbsentshort-circuit —putIfAbsent/computeIfAbsenton an existing key return before serializing the value, avoiding wasted serialization on CAS-miss workloads.- byte[]-key word-wise lookup + direct value path — caches using
BYTE_ARRAY_KEYnow compare keys 8 bytes at a time (getWithLen/ off-heap-ghostgetSlotWithLen) instead of byte-by-byte on every lookup; built-inbyte[]values bypass the generic segment serializer; and same-size updates take a dedicatedupdateValueWithLenSameSizeFastpath. Measured on byte[] keys + 256 B values at 1M entries: GET throughput +22–37%, PUT update latency −21% (non-byte[]key types are unaffected).
4 threads, JDK 25, macOS, 256 B values (FairComparisonScaleBenchmark + OHCComparisonBenchmark, JMH AverageTime, all caches measured in one run). Lower is better.
GET (ns/op):
| Cache | 10K | 100K | 1M |
|---|---|---|---|
| RMCache | 107 | 257 | 424 |
| RMCache+Ghost | 126 | 236 | 413 |
| Chronicle Map | 251 | 305 | 445 |
| OHC | 284 | 430 | 629 |
| MapDB | 1,098 | 1,731 | 2,214 |
| EhCache | 1,639 | 1,765 | 2,003 |
| Caffeine (on-heap ref.) | 65 | 105 | 254 |
PUT (ns/op):
| Cache | 10K | 100K | 1M |
|---|---|---|---|
| RMCache+Ghost | 130 | 278 | 474 |
| RMCache | 169 | 320 | 488 |
| Chronicle Map | 603 | 622 | 715 |
| OHC | 426 | 595 | 1,044 |
| EhCache | 2,460 | 2,740 | 3,095 |
| MapDB | 2,621 | 4,205 | 4,660 |
| Caffeine (on-heap ref.) | 152 | 248 | 511 |
RMCache is the fastest off-heap cache measured — faster than Chronicle Map, OHC, MapDB, and EhCache at every scale on both GET and PUT. On PUT it stays within range of on-heap Caffeine despite living entirely off-heap. Caffeine is listed only as an on-heap reference point, not a direct competitor.
- Off-heap cache engine built on Java Foreign Function & Memory (FFM) API — zero
Unsafedependency - Slab allocator with lock-free bitmap allocation (CAS-based
AtomicLongArray), 11 size classes (64 B – 64 KB) - Buddy allocator for large values exceeding slab size (>64 KB)
- Robin Hood hash table with
StampedLockoptimistic reads and graceful resizing - EntryPool with packed slot metadata and partitioned locking (128 partitions default)
- W-TinyLFU eviction (3-segment SLRU + Count-Min Sketch frequency filter, fully off-heap)
- Off-heap timing wheel for TTL-based expiration with per-entry TTL support
- Ghost Cache L1 — direct-mapped off-heap shortcut bypassing hash table probes
- Zero-copy access via
CacheValueView/getZeroCopy()API - Background eviction with configurable high/low watermarks (95%/90% default)
- Memory estimator for capacity planning (
CacheBuilder.estimateMemory()) - Index memory budgeting to constrain hash table memory usage
- Built-in serializers for String (UTF-8/Latin-1) and byte arrays
- Segment value serializer for direct native memory writes without intermediate heap buffers
- Latin-1 fast path for ASCII string keys (zero-allocation key encoding via
ThreadLocalKeyBuffer) - Eviction listeners and filters for custom eviction control
- Composite eviction policy combining LRU + TTL
- JMH benchmarks comparing against NMA and Ehcache
CacheStatsrecord with hit rate, miss rate, memory usage, and eviction countscomputeIfAbsentwith non-atomicity contract documentedputIfAbsent,putAll,getAll,putAsync,getAsyncbulk and async APIscleanupThreadLocals()static method for thread-pool thread retirement
4 threads, JDK 25, macOS:
| Operation | Scale | Latency |
|---|---|---|
| GET | 10 K | 179 ns |
| GET | 100 K | 360 ns |
| GET | 1 M | 590 ns |
| PUT | 10 K | 281 ns |
| PUT | 100 K | 403 ns |
| PUT | 1 M | 695 ns |
- Memory leaks: Native slab blocks freed on failed allocation
- Overflow guard: Fail-fast when slot capacity exceeds
Integer.MAX_VALUE - LRU 30-bit guard: Prevent slot corruption for >1.07B entries
- CoarseClock lifecycle: Reference-counted start/stop with synchronized release
- Frequency sketch thread safety: Atomic sample counter + CAS-guarded reset
- Hash table resize safety: Grace-period retired segment cleanup
- Ghost cache performance: Eliminated redundant hash table re-validation on ghost hits
- Zero-heap profile: Conditional key deserialization in eviction path
- Allocation-free value reads: Replaced
getValuePosition()with offset/length primitives