Exact and semantic response caching for Java LLM tasks — skip a repeat call to
the model entirely instead of paying for it twice. The Java port of
cacheflow (Rust): same idea, adapted
to Java: no framework dependency, just a Handler (String -> String)
decorated the way a Servlet Filter decorates a request handler.
Not published to Maven Central — pull it straight from GitHub via JitPack:
<repositories>
<repository>
<id>jitpack.io</id>
<url>https://jitpack.io</url>
</repository>
</repositories>
<dependency>
<groupId>com.github.thaicn1712</groupId>
<artifactId>cacheflow-java</artifactId>
<version>main-SNAPSHOT</version> <!-- or a tagged release -->
</dependency>Wrap a handler with an exact-match cache and a function that extracts the cache key from the input:
Cache cache = new InMemoryCache();
Handler cached = CacheFlow.withCache(myLlmHandler, cache, Function.identity(), null);
String first = cached.run("what is the meaning of life?"); // runs myLlmHandler, caches the response
String second = cached.run("what is the meaning of life?"); // identical prompt: cache hit, myLlmHandler never runsmyLlmHandler is any Handler — String run(String input) throws Exception, the shape most LLM call wrappers in Java already have. No interface to implement beyond that, no framework to adopt.
Give entries a TTL so stale answers expire:
Handler cached = CacheFlow.withCache(myLlmHandler, cache, keyFn, Duration.ofHours(1));Exact-match misses on paraphrases. SemanticCache hits on similar prompts by comparing embeddings, not strings — implement Embedder against whatever embedding API or local model you already use:
class MyEmbedder implements Embedder {
@Override
public float[] embed(String text) {
// call an embeddings API, or a local model
throw new UnsupportedOperationException("todo");
}
}
Cache cache = new SemanticCache(new MyEmbedder(), 0.92f); // cosine similarity threshold
Handler cached = CacheFlow.withCache(myLlmHandler, cache, keyFn, null);Cache is a plain interface — InMemoryCache and SemanticCache ship built in, implement it yourself against Redis/Postgres/pgvector for a cache that survives a restart.
The same prompt often repeats — a retried request, a templated prompt with few variables, a user asking the same question a different way. Re-running it through the model burns latency and money for an answer you already computed. cacheflow-java sits between the caller and the model call, exact-matching or semantically-matching the input against what's already been answered, and only calling through on a genuine miss.
mvn -q compile exec:java -Dexec.mainClass=io.github.thaicn1712.cacheflow.examples.SkipARepeatPromptExampleMIT