Warning
Iris is in beta. It is young, it has one user, and anything may change between minor versions.
One small interface for sending prompts and conversations to a large language model, with adapters for Anthropic, OpenAI and Google Gemini, so that switching provider is a matter of configuration. It is built on Cats Effect, sttp and Circe.
Named for Iris, messenger of the gods, who carried their words to mortals along the rainbow.
Add the following to your build.sbt:
libraryDependencies += "com.alecdorrington" %% "iris" % "0.1.0"Compiled with Scala 3.8.4, with no intention to explicitly support older versions. JVM only.
The public interface is LlmClient.
Build one for an explicit LlmConfig with LlmClient.resource,
over an sttp backend of your own with LlmClient(config, backend),
or from the environment (see Configuration) with LlmClient.fromEnv.
import com.alecdorrington.iris.{LlmClient, Prompt}
LlmClient.fromEnv[IO].use {
case Some(client) => client.complete(Prompt("Hello!"))
case None => // No provider configured.
}Conversations are stateless: nothing is remembered between calls, and the full history
travels with every request as a Chat.
To continue a conversation, append the model's reply and the next user message, then send the chat again:
import com.alecdorrington.iris.Chat
val chat = Chat().withSystem("You are terse.").user("Name a colour.")
for
first <- client.send(chat)
second <- client.send(chat.assistant(first.text).user("And another."))
yield second.textEach request accepts CompletionOptions
(model override, token limit, temperature, top-p, stop sequences), and each
Completion carries the reply text along with a
normalised StopReason and token usage counts.
client.send(chat, CompletionOptions(maxTokens = Some(1024), stopSequences = List("\n\n")))Note
Anthropic's newer models, the default claude-sonnet-5 among them, no longer accept sampling
parameters and reject a request which carries them. Iris omits temperature and top-p for those
models rather than let the request fail; both still apply to every other provider, and to
Anthropic's older models.
A message is a list of Parts, so text may travel beside
the media it refers to. A message of text alone is still sent as plain text, so nothing
changes for a conversation which carries none.
import com.alecdorrington.iris.Part
chat.user(Part.Text("What is in this picture?"), Part.Media("image/png", base64))Anthropic and Gemini take both pictures and documents. OpenAI's chat completions take
pictures only, and a document there fails with LlmError.Unsupported rather than being
dropped or sent as something it is not.
Where a reader is waiting, LlmStream delivers the reply
as it is written. It is a capability apart from LlmClient, because it needs a backend which
can stream.
import com.alecdorrington.iris.{Delta, LlmStream}
LlmStream.fromEnv[IO].use {
case Some(llm) => llm.stream(Prompt("Tell me a story.")).evalMap {
case Delta.Text(text) => IO.print(text)
case Delta.End(reason, usage) => IO.println(s"
($reason)")
}.compile.drain
case None => IO.unit
}Nothing is sent until the stream is run, and a failure reaches it as an LlmError as it
would anywhere else. Delta.End carries the token counts where the provider reports them
as it ends; Anthropic counts the input at the start instead, so send remains the way to
have both halves together.
Offer the model tools it may ask to have run, and it will say so in Completion.toolCalls.
Iris never runs a tool itself; what one does, and whether it is allowed to, is yours to decide.
import com.alecdorrington.iris.{Part, Tool}
val weather = Tool("weather", "Looks up the weather.", schema)
for
asked <- client.send(chat, CompletionOptions(tools = List(weather)))
result = Part.ToolResult(asked.toolCalls.head.id, "weather", lookUp(asked.toolCalls.head))
answer <- client.send(chat.reply(asked).results(result))
yield answer.textA reply which asked for a tool has StopReason.ToolUse, on every provider, and so does a
streamed reply's Delta.End β Gemini reports no such reason of its own, so the asking is what
says so.
count asks the provider what a chat would cost to send, so a conversation can be
checked against a budget or a context window before a completion is spent finding out.
client.count(Prompt("How long is a piece of string?"))OpenAI offers no such endpoint, and fails with LlmError.Unsupported rather than
guessing with a tokeniser of its own.
When many requests begin alike, as when one document is asked several questions, put what
they share first and end it with a Part.CacheBreakpoint, so that the provider can keep what it
made of the prefix rather than read it afresh every time. Usage.cachedTokens says how much of a
prompt was read from the cache.
val asking = (question: String) =>
Chat().withSystem("Answer from the document.")
.user(Part.Text(document), Part.CacheBreakpoint, Part.Text(question))
for
_ <- client.warm(asking(""))
answers <- questions.parTraverse(question => client.send(asking(question)))
yield answersRequests sent at once cannot read what none of them has written yet, so warm writes the prefix
first: everything up to the last breakpoint, asking for as little reply as the provider allows.
A provider answers only a chat which ends with the user saying something, so a prefix which does
not, as when the last breakpoint opens a message to cache only what came before it, is followed
by the least a user can say.
Anthropic caches only what it is told to, and a breakpoint marks the block before it, at most four blocks in one chat (breakpoints in a row mark one). OpenAI and Gemini cache long prefixes of their own accord and are sent nothing. Each provider keeps a cache per model, and caches nothing shorter than its own minimum.
A provider's refusal fails the effect with an LlmError:
Http for an unsuccessful response, carrying its status and body, and Malformed for a response
that could not be understood. Response bodies may contain provider detail you would rather not show
to your own users, so consider logging them rather than passing them on.
LlmConfig.fromEnv (and so LlmClient.fromEnv) reads these environment variables:
| Variable | Meaning | Default |
|---|---|---|
LLM_PROVIDER |
anthropic, openai or gemini |
Inferred from which key exists |
ANTHROPIC_API_KEY |
API key for Anthropic | - |
OPENAI_API_KEY |
API key for OpenAI | - |
GEMINI_API_KEY |
API key for Gemini (or GOOGLE_API_KEY) |
- |
LLM_MODEL |
Model name to use | Provider-specific default |
LLM_MAX_TOKENS |
Maximum number of tokens in each completion | 8192 |
LLM_BASE_URL |
Overrides the provider's API origin | The provider's own origin |
LLM_TIMEOUT |
Seconds to wait for a completion | 300 |
With no key set, fromEnv yields None. So does a variable which is set but cannot be used:
an unrecognised LLM_PROVIDER, rather than falling back to whichever key exists, and an
LLM_MAX_TOKENS which is not a positive whole number, rather than quietly reverting to the
default and hiding the mistake.
Iris is developed as part of a larger private project, of which this repository is an automatically synchronised mirror (by GitHub Graph), so changes made here directly would be overwritten. Issues are very welcome; for anything more, please open an issue first.
- Hecate, a sibling, for user accounts, sessions, groups and permissions.
- Eunomia, a sibling, for filtering, ordering and paging lists.
- Dike, a sibling, for ranking by pairwise comparison, with a model as the judge.
- qr4s, a sibling, for generating QR codes, on the JVM and in the browser.
- This library was made using Scala Library Template.