Here's how">
Skip to navigation
Getting Started

Bug Fixes & Improvements

  • Large threads no longer fail to load — threads/retrieve had been intermittently returning a 500 since a ClickHouse join between traces and their aggregated spans could pick either side to build the join from; on a large thread it occasionally picked the multi-gigabyte side and ran out of memory. The join now always builds from the smaller, pre-aggregated side, so large threads load reliably.

  • Diagnostics checks your credit balance before starting a run — Starting a diagnostic run when a workspace was out of credits used to fail only after the run had already started. The run button now checks affordability upfront, and the page shows a clear out-of-credits state instead of a run that fails partway through.

  • Image and other media in experiment output now renders in the comparison views — The experiment comparison sidebar and the test-suite sidebar only extracted media (e.g. images) from a dataset item’s input; a trace’s output went straight into the raw JSON view with no media extraction, so an image an SDK logged as output never showed. Output media now renders the same way input media does, including images served from URLs with no file extension.

  • Online evaluation rule duration filters now use the right unit — A duration filter typed as ”> 5” on a rule was compared against the underlying duration column in milliseconds, so on an LLM-as-judge rule it matched effectively every trace instead of only the slow ones. The rule dialog now converts the value the same way the traces table already does, and the duration column is labeled “Duration (s)” so the unit is clear.

  • Bedrock and Mistral streaming integrations no longer lose parts of the response — OpenAI models called through Bedrock’s invoke_model (gpt-oss, GPT-5.x, GPT-6) stream text and stop-reason in a shape the aggregator didn’t recognize, so those spans ended up with empty output and zero token usage. Bedrock’s converse_stream sent a tool call’s arguments as JSON fragments that got overwritten instead of merged, so a streamed tool call kept only its last fragment, and parallel tool calls collapsed into one. Mistral’s reasoning models (Magistral) stream content as a list of thinking/text chunks rather than a string, which crashed the aggregator and left the whole span without output. All three now aggregate correctly.

  • Judge calls to Bedrock-hosted GPT-5.x and GPT-6 models no longer fail — Opik already dropped the temperature parameter for GPT-5-family judge models, since these models reject it, but only recognized them by their plain gpt-5* names. The same models routed through Bedrock (e.g. bedrock/converse/us.openai.gpt-6-...) keep a bedrock/ prefix, weren’t recognized, and had their judge calls rejected. They’re now matched and handled the same way.

  • Hallucination and SycEval judge reasons render as text, not a Python list — Both metrics ask the judge for a list of reason strings, but the parser rendered that list with Python’s str(), so the explanation shown to you (and uploaded with the score) was a literal ['reason 1', 'reason 2'] instead of readable prose. It’s now joined into text the same way other list-based judge reasons already are.

  • Context-precision and context-recall judges get the same prompt-injection hardening as last week’s hallucination and G-Eval fix — Content being evaluated by these metrics is now clearly namespaced from the rest of the judge prompt, so an input, output, or context that happens to contain something resembling an instruction or a closing delimiter can no longer influence the verdict.

  • Several built-in metrics are more robust on edge-case input — BLEU now rejects a non-positive or non-integer n_grams instead of silently scoring 0 or raising a raw KeyError; METEOR tokenizes its input before handing it to NLTK instead of raising TypeError on every call; ChrF now scores a candidate against each reference separately and keeps the best match instead of passing the whole reference list into NLTK’s single-reference scorer (which scored a perfect match as low as 0.26); and VADER sentiment now downloads its lexicon on first use instead of failing with a bare LookupError on a fresh install.

  • @track now logs slots=True dataclasses correctly — Encoding a dataclass defined with slots=True read from __dict__, which such classes don’t have, so the logged value was either the object’s repr string or an empty object instead of its actual fields.

  • Config files with % in a value no longer break opik configure — ~/.opik.config is read and written with Python’s ConfigParser, which by default treats % as the start of an interpolation sequence; a value containing % (for example, in a URL) could fail to load or save. Interpolation is now disabled for this file.

  • Custom scorers now get a real name in results instead of crashing — Passing a functools.partial-wrapped function or a callable class instance as a scorer raised AttributeError because neither has a __name__. Both are now named sensibly (the wrapped function’s name, or the class name as a fallback).

  • Gemini usage without candidates_token_count is accepted — A Gemini response that omits this field (for example, at content-filter length) is no longer rejected while building token usage; the count is treated as unknown rather than causing an error.

  • 19 more providers are priced correctly — Hyperbolic, Baseten, Lambda AI, nscale, OCI, Replicate, watsonx, Cohere, Novita, Cloudflare, Anyscale, Scaleway, OVHcloud, GMI, Gradient AI, Libertai, Azure AI, Vercel AI Gateway, and OpenRouter are now registered as canonical providers, so calls through them resolve a model price instead of going unpriced.

  • Large dataset batch uploads no longer exceed the configured size cap — StreamingBatchWriter decided when to flush a batch based only on the serialized items’ size, without counting the surrounding envelope (dataset name, project name, batch id, and JSON wrapper). Every flushed request body was slightly larger than the configured cap; the cap now accounts for the full request body.

  • video_url placeholders in chat prompts are validated — Placeholder validation (validate_placeholders=True) checked {{...}} templates inside text and image_url prompt parts but skipped video_url, even though it’s rendered the same way, so a templated video URL could silently fail to substitute.

  • Overlapping evaluations no longer leave your app’s HTTP connections short-lived — evaluate() temporarily patches httpcore’s keep-alive behavior for the duration of a run and restores it afterwards. When two evaluations ran at the same time, the first one to finish could restore the patch early (disabling it for the run still in progress) or the last one to finish could restore the wrong thing, leaving the host application’s own connections short-lived after both runs completed. The patch is now reference-counted so it’s only removed once every overlapping run is done.

Performance Improvements

  • Reading large experiments page by page is up to ~2.6x faster — The query behind experiment item paging was re-evaluated up to six times per page because it’s referenced from multiple places in the query plan; it’s now evaluated once per page and the result reused. Measured on a 100k-item experiment at a page size of 2000: a 50-page read dropped from 14.48s to 5.56s.
  • The experiment compare view no longer re-scans the whole experiment on every page — Each page of the compare view re-resolved the same target-project lookup by reading the experiment’s entire trace set, even though the answer is identical for every page of that read. The lookup is now cached per read, so its cost no longer scales with the size of the experiment.
  • Trace and thread reads use less CPU and read less data — find_trace_stream and find_thread_by_id deduplicated spans with a ClickHouse FINAL read, which is expensive. Both now use a pre-aggregated form instead, cutting CPU by 22–50% and roughly halving bytes read in production measurements, with identical results.
  • Span-by-ID reads scan less data — Reads that look up a span by ID now bound themselves to the week(s) that ID’s timestamp resolves to, instead of scanning across all partitions.

Bug Fixes & Improvements

  • opik import can import traces of any age again — The trace importer derived each imported trace’s ID from the exported trace’s start_time, which bakes that original timestamp into the ID’s UUIDv7 prefix. Opik validates that an ingested ID’s embedded timestamp falls inside an ingestion window around now — 24 hours by default on Opik Cloud, which is what keeps its storage layer partitioned correctly — so importing anything older failed with Invalid UUID for id ... reason 'too_old'. Imported traces now get IDs minted at import time, which also keeps IDs unique across projects and workspaces. start_time and end_time still carry the original values, spans and feedback scores are remapped onto the new IDs, and each trace records the ID it came from in its metadata as _import_id. Traces are imported in the order the source project listed them, so the destination trace list and thread view keep that order; note that the time-bucketed project metrics key off the ID, so imported traces count at import time. If you hit the error above, upgrade the SDK. See Imported trace and span IDs.

Map Trace and Span Fields onto Dataset Item Columns

Adding traces or spans to a dataset always used a fixed shape — input from the trace’s input, expected_output from its output, nothing else. Getting any other field into the dataset meant exporting and reshaping the data by hand.

The “Add to dataset” dialog now has an Advanced mapping toggle. With it on, you can re-point input and expected_output at a different field, add extra columns from a path explorer over the sampled traces/spans, and use Quick add chips to pull in common enrichment fields (feedback scores, metadata, tags) or columns the target dataset already has. A coverage badge on each mapped field and a preview table show what will actually land in the dataset before you commit.

Name Experiments from the Playground

Playground experiments were always given a random adjective_noun_number name, with no way to set one — making it harder to find a specific run later in the Experiments list.

The Playground’s output toolbar now has an optional experiment name field. Each output column derives its own name from it by appending a letter (concise → concise_a, concise_b, …), and a preview shows what a run will produce before you start it. The field is optional: leave it empty and Opik generates a name exactly as before, so nothing changes if you don’t use it. A completion toast reports the experiments created and links straight to the comparison view.

Bug Fixes & Improvements

  • Exporting experiment results from the browser now covers the full result set — The comparison table’s export button only ever covered the rows selected on screen, and was disabled with nothing selected. With no selection, it now exports the whole result set behind the current filters (capped at 2000 rows, since a browser export holds everything in memory); past that cap, it points you to the SDK export script, which has no such limit.
  • Claude models no longer 400 when routed through Bedrock, OpenRouter, or a custom provider — Claude rejects a request that sets both temperature and top_p. Opik already enforced that, but only for the Anthropic provider directly; the same Claude model called through Bedrock, OpenRouter, or an OpenAI-compatible custom provider sent both and was rejected. The either/or choice now applies to every provider on the backend, and the Playground’s model panels all show the same either/or control for any Claude model instead of only some of them.
  • Copy actions are now one click away on trace, thread, and span panels — Copy-ID and copy-link used to live inside an overflow menu; they now render as icon buttons right next to the panel title, with the overflow menu reduced to Export and Delete.
  • A setup hint now appears on expanded trace errors — Expanding an error in a trace now offers to connect an AI coding assistant (Claude Code, Cursor, VS Code, and others) via MCP, with install instructions and a ready-to-use prompt for investigating that specific error.
  • opik configure and opik mcp configure now ask the same questions the same way — The two commands used to ask about registering an MCP server differently, and a bare Enter could silently decline or silently accept depending on which one you ran. Both now show the same explanation and client picker, which gained an “All” option and a manual-setup fallback (with a docs link) for AI clients Opik doesn’t auto-detect.
  • Threshold alerts no longer silently stop firing — An alert whose threshold trigger was missing a window (including some created before this fix) failed internally on every evaluation and was skipped without notifying anyone. Alerts in this state now evaluate using their configured or default window instead of failing silently, and creating or editing one with an invalid threshold or window is now rejected up front with a clear error.
  • Bulk-deleting alerts now requires the same permission as editing them — deleteAlertBatch was missing the permission check its create/update siblings already had.
  • Webhook destinations are now checked before Opik connects to them — Webhook deliveries and the “test connection” action now validate the destination URL first; on Comet-managed workspaces this blocks destinations that only Opik’s own network can reach, and a failed test no longer echoes the destination’s response body back to the caller.
  • LLM-as-judge evaluations (hallucination, G-Eval) can no longer have their verdict overridden by the content being judged — A judge prompt whose evaluated input, context, or output happened to contain something that looked like a closing delimiter could break out of its section and forge a verdict. The judged content is now clearly namespaced from the rest of the prompt.
  • G-Eval scores are now accurate for two-digit scores and more tokenizers — The logprob-based scorer assumed a score is always a single digit at a fixed token position. A two-digit score (e.g. 10) or a tokenizer that folds whitespace into the score token produced a badly wrong score instead of the intended one; both are now located and parsed correctly.
  • LangChain streaming integrations no longer drop token usage — A streaming run with an empty generations list (seen on the Anthropic-Vertex AI path) raised internally and discarded token usage for that call instead of reporting None.
  • Playground dataset runs are now scored by the rules you actually selected — An online evaluation rule scoped to “Production traces” was scoring (or failing to score) Playground dataset runs regardless of the rules picked in the Playground’s own metric selector. Playground dataset runs are now scored only by the selected rules plus any rule explicitly scoped to experiments, and a Playground run with no dataset is no longer treated as production traffic.

Performance Improvements

  • Dataset and experiment transfers now default to 8 threads — The Python SDK’s bulk paths shipped with inconsistent worker counts: dataset reads (get_items(), stream_items()) and Dataset.insert() ran on 4 threads, while Experiment.batch_upload_items() uploaded batches sequentially unless you passed num_threads yourself. All three now default to 8, the setting we benchmark against — the experiment upload path benefits most, since it went from sequential to parallel. Every call still takes num_threads if you want to push harder or ease off; see Tuning SDK throughput. Note that raising it much further does not help: on a 119,903-item upload, 16 threads measured slightly slower than 8, because the SDK saturates a CPU core serializing and compressing payloads before thread count becomes the limit. Parallel dataset upload still requires an Opik backend of 2.2.8 or newer, and falls back to a sequential upload against anything older.
  • Dataset and experiment uploads encode payloads up to 5x faster — The Python SDK now uses orjson to encode request bodies and compute content-deduplication hashes where it’s available, instead of the standard library’s json module. On a 30,000-item upload over 8 threads, end-to-end time dropped from 148.7s to 78.2s. This is automatic and requires no changes to your code; installs without an available orjson wheel fall back to the standard library unchanged.

Custom AI Providers Can Use OAuth2 Token Auth

Custom AI provider integrations only accepted a single static API key, so any provider that rotates or expires credentials — an enterprise gateway sitting behind OAuth2, for example — had to be re-configured by hand every time a key expired.

The provider configuration dialog now has an Authentication mode switch: choose Static API key, or OAuth2 client credentials and enter the token URL plus the client_id and client_secret rows (add more rows, such as scope or audience, if your auth service needs them). This applies to custom providers and to Bedrock.

Fine-Grained Thinking Controls for Gemini Models

Gemini models increasingly default to an internal “thinking” step before answering, which adds cost and latency that isn’t always worth paying — Opik previously had no way to influence that.

The Playground, Online Evaluation rules, and experiments can now set a thinking level per Gemini call, for both Vertex AI and Google AI Studio: Auto, None, Minimal, Low, Medium, or High (availability depends on the model — Gemini 2.5 Pro, for example, can’t disable thinking). Vertex AI translates the level into the thinking_budget value it expects; Google AI Studio sends the level directly. A follow-up fix added the dedicated None level after Flash Lite models — which don’t think by default — were found to have thinking silently re-enabled by a preselected Minimal level, adding several seconds of latency to every call.

Bug Fixes & Improvements

  • Judge scoring no longer 500s or stalls on edge-case content — A judge message whose content legitimately started with [ (for example, a prompt beginning “[Source Text]…”) could return a 500 for an entire project’s list of evaluation rules and silently stop sampling for that project. A trace whose mapped input, output, or metadata was a bare JSON value instead of an object failed its whole evaluation instead of just skipping that field, and a Vertex AI evaluation using the tool-calling path failed outright because Vertex rejects a forced tool choice. All three now evaluate instead of failing.

  • Online-scoring queues no longer get stuck behind one bad message — A single oversized or undecodable message could wedge an entire scoring stream, blocking every other trace behind it (one recorded case reached tens of gigabytes of stuck messages), and a permanently failing provider error (like an invalid-credentials response) was retried repeatedly instead of being retired immediately. Both cases are now dropped or retired instead of stalling the queue.

  • Online evaluation sampling now applies only to production traces — The sampling rate is meant to thin a continuous production stream, but it was also being applied to traces from experiments, playground runs, and optimizations — one-off runs a user starts intentionally. Those are now always scored in full, regardless of the rule’s sampling rate.

  • Prompt version history now loads and labels every version correctly — The Prompt tab’s version history only ever loaded the first 25 versions, so older versions were unreachable and a deep link to one silently fell back to showing the latest version under a stale label. Version history, the Compare and Deploy menus, and version labels throughout the app (Playground, trace details, Optimizations, Agent Runner) now paginate through the full history and use each version’s real, persistent number.

  • Experiment items export no longer drops rows or loses sort order — Exporting experiment items truncated long field values and could ignore the sorting and search filters applied in the table. Exports now include full field values and respect the same sort and search as the on-screen view.

  • Dataset item and version data is more accurate — A dataset version’s reported item count came from the raw number of items submitted, not the number actually stored, so a batch with duplicate item ids showed an inflated count. Separately, filters on a dataset item listing weren’t always applied consistently between the page and count queries, and page reads weren’t bounded to the requested page size. All three are now correct.

  • Annotation queue names can no longer cause lost form data — A whitespace-only queue name passed client-side validation, was rejected by the API, and closed the create/edit dialog anyway — discarding everything entered in the form. Names are now trimmed and validated before submission, and the dialog stays open until the save actually succeeds.

  • Dashboards can now be filtered by description — Filtering the dashboards list by its description field returned “No matching results” for every operator, because the field wasn’t registered as filterable on the backend. It’s now supported like any other field, and a failed list request now shows an error instead of silently rendering as empty.

  • Optimization runs list now shows the actual best trial — The runs list reported the baseline trial’s latency and cost as if it were the run’s best result, with the cost/latency delta always showing 0%, while the run’s own detail page correctly showed the genuine best trial. The list now matches the detail page.

  • Cost tracking recognizes more providers and model naming schemes — Calls through Cerebras, Snowflake Cortex, and DeepInfra weren’t recognized as canonical providers, models identified by OpenTelemetry semantic-convention provider names weren’t mapped to Opik’s provider list, and models with compact YYYYMMDD date suffixes in their id weren’t matched to a price. All of these now price and attribute correctly.

  • Streaming SDK integrations no longer swallow errors — The Anthropic, Bedrock, and Mistral stream wrappers’ cleanup step could silently swallow an exception raised while finalizing a stream instead of surfacing it, and the LangChain integration stopped extracting token usage whenever a call’s model metadata was absent. Both are fixed, and the Bedrock integration now correctly passes through Claude’s cache-read/write token counts instead of undercounting them.

  • Cursor extension reports accurate usage and structure — The Opik extension for Cursor was logging zero-token spans after Cursor stopped populating the field it read from, one flat span per conversation turn instead of a span per model or tool call, and undercounted prompt and total tokens (by up to ~6x) whenever cache reads dominated a call. All three are fixed.

  • MCP OAuth consent screen preselects your default workspace — Approving an AI tool’s access via MCP OAuth used to preselect whichever workspace the backend returned first; it now preselects your actual default workspace.

  • opik configure can set up MCP and skill packs in one step — Running opik configure --install-mcp --install-skills registers the Opik MCP server with detected AI coding tools (Claude Code, Cursor, VS Code, Codex, opencode) and installs the Opik skill pack that teaches the coding agent how to instrument your code, instead of setting each one up by hand.

  • Dataset inserts can skip deduplication — Dataset.insert() (and the batch, pandas, and JSONL variants) now accepts a deduplication argument; setting it to False skips downloading and hash-comparing existing items before insert, trading duplicate-safety for speed on large datasets.

  • Python SDK reports anonymous feature usage — The SDK now reports which integrations and features are used (for example, which LLM provider you’re calling through) to help prioritize development; no trace, span, prompt, or dataset content is included. This is on by default and can be disabled by setting OPIK_ANALYTICS_ENABLE=false. Running opik configure or opik mcp configure additionally attaches an account identifier so a setup run can be tied to your workspace.

  • Self-hosted online evaluation rules get the same tool-calling judge as Comet-managed workspaces — LLM-as-judge rules that reference {{trace}}, {{span}}, or {{spans}} can let the judge model call tools to fetch large trace or span data on demand instead of inlining all of it into the prompt, avoiding context-window overflow. This previously only ran for Comet-managed workspaces; the feature toggle gating it off for self-hosted instances has been removed.

Performance Improvements

  • Dataset writes and reads are significantly faster — Inserting into a dataset serialized every batch behind a lock only needed for creating a new version, item-count updates cost three database round-trips instead of one, and the enrichment queries used when loading dataset details ran one after another instead of concurrently. A redundant item-count scan also ran even when a version already tracked its own total, and a per-dataset experiment summary query scanned the whole workspace’s experiment items instead of just the requested dataset’s. All of these are fixed, substantially reducing dataset save and read latency.

  • Reading large datasets from the Python SDK is up to ~3x faster — A new parallel, chunked Dataset.stream_items() reader (which get_items() now uses internally) measured 3.3x faster than the previous single-threaded read path at 8 threads, with no change to its output.

  • Traces, trace logs, and experiment results tables render faster on large projects — These tables now virtualize their rows and columns, so only what’s visible on screen is rendered, instead of paying render cost for every column and row up front.

MCP + Skill Pack Setup Can Now Run Without a Terminal

Setting up the Opik MCP server previously meant sitting through an interactive wizard, and it only reached three of the AI assistants people actually use. opik configure --install-mcp and opik mcp configure --ai-client <host> now run non-interactively — pass the client explicitly (or --ai-client all) and the command completes on its own, so it can run from a coding agent, a Dockerfile, or CI. Codex and opencode join Claude Code, Cursor, and VS Code Copilot as supported hosts, and every install now ends with a real verification call that reports the workspace and project count instead of assuming the config it just wrote actually works.

Bug Fixes & Improvements

  • Self-hosted: agentic tool-calling scoring is on by default — LLM-as-judge online scoring’s agentic tool loop (previously introduced behind a toggle) was already enabled on Comet-hosted workspaces; self-hosted installs inherited the toggle’s off default. The toggle has been removed and the behavior is now unconditional everywhere.

  • Exports now match what’s on screen — Exporting experiment items ignored the active sort and search, and truncated long field values regardless of the on-screen truncation setting. Exporting traces or threads with a search term containing leading or trailing whitespace could also return different rows than the table showed. All three now export exactly what’s displayed.

  • Optimization runs list reports the same best trial as the run page — The runs list computed its “best” latency and cost from a column that dataset-based runs (Studio runs and every SDK optimizer) never populate, so it silently fell back to the baseline trial with a flat 0% delta on every row. It also didn’t discount a candidate that had only evaluated part of the dataset. Both now match the run page’s logic, and the previously mislabeled “Opt. cost” column is named consistently between the two screens.

  • Dashboards can be filtered by description, and a failed list load shows an error instead of “no results” — Filtering dashboards by description previously returned a 400 that several list views quietly rendered as an empty state rather than surfacing. Both are fixed: the filter now works, and a failed request shows an error message on affected pages.

  • Annotation queue names can no longer be blank, and a rejected save no longer discards your edits — A name of only whitespace passed client-side validation and was rejected by the API, closing the dialog and losing everything typed. Whitespace-only names are now blocked before submission, and the dialog stays open on any save failure so nothing is lost.

  • Dataset item pages with large payloads no longer fail, and filters return correct results — Reading a page of a large dataset version could exceed memory limits and return a server error, since fetching a single page still had to sort the entire version in memory. Pages are now resolved in two steps that avoid the full sort, and a follow-up fix ensures filters and sorting apply to the right columns so filtered pages return the correct items instead of a short page.

  • More accurate cost and usage across several providers and integrations — DeepInfra model prices (including its Claude, DeepSeek, Qwen, and Llama catalogs) now load instead of costing 0,sinceDeepInfrawasmissingfromtheinternalproviderlist.CustomOpenTelemetryinstrumentationreportingaprovidernameintheOTelsemantic−conventionvocabulary(forexample‘vertexai‘or‘aws.bedrock‘)nowmapstoOpik′scanonicalprovidernamesinsteadofsilentlycosting0, since DeepInfra was missing from the internal provider list. Custom OpenTelemetry instrumentation reporting a provider name in the OTel semantic-convention vocabulary (for example `vertex_ai` or `aws.bedrock`) now maps to Opik's canonical provider names instead of silently costing 0. Claude’s cache-read and cache-write token counts now flow through correctly when using Claude on Bedrock with streaming. LangChain usage and cost tracking no longer disappears entirely when model metadata can’t be resolved.

  • Streaming SDK integrations no longer silently swallow errors — The Anthropic, Bedrock, and Mistral integrations patch shared streaming classes process-wide to add tracing. A cleanup step meant to run only for tracked calls used a return inside a finally block, which discarded any in-flight exception — so once any traced call ran, an untracked stream elsewhere in the same process that failed would complete silently instead of raising. For Bedrock this could also return None in place of real response data from any botocore call, not just Opik’s. Tracked and untracked streams now both behave as expected.

  • Cursor extension logs each tool and model call as its own span — A Cursor agent turn previously logged as a single trace and a single span, so a multi-step turn showed only the initial question and the final answer. Each model call now logs as its own LLM span and each tool call as its own tool span, nested under the turn, with tool names interleaved with assistant messages in the trace output.

Performance Improvements

  • Faster dataset experiment summaries — The query behind a dataset’s experiment summary scanned every experiment item in the whole workspace before filtering down to the requested dataset. It now prunes to the relevant experiments upfront, cutting rows read by roughly 13x on a workspace with 3M experiment items across 200 experiments, and skipping the scan entirely for a dataset with no experiments.

  • Faster dataset item inserts — Adding items to a dataset updated the version’s item count with a read-modify-write cycle inside the per-dataset lock. It’s now a single atomic increment, cutting the database round trips on that path from three to one.

evaluate_threads Can Now Judge RAG Agents on Retrieved Context

evaluate_threads only accepted trace_input_transform and trace_output_transform, so a thread was flattened to plain role/content messages with no way to pass along the documents each answer was grounded on — judging relevance or faithfulness for multi-turn RAG agents meant falling back to scoring individual traces instead.

A new trace_context_transform argument receives the whole trace object and attaches the extracted context to that trace’s message:

Bug Fixes & Improvements

  • Daily briefings report “Not enough credits” instead of failing silently — A daily report that ran out of LLM credits mid-run used to just say “Briefing failed,” with the pod already allocated and the spend already burned. The trigger now checks credits up front, and a briefing that can’t run for that reason shows “Not enough credits” with a direct link to add more, alongside Run again.

  • Provider errors keep their real HTTP status instead of turning into a 500 — Traces and online-scoring rules calling out to an LLM provider reported a generic server error whenever the provider’s response couldn’t be parsed into its usual error shape — a plain-text upstream outage message, or a rate-limit response wrapped by a retry layer. The provider’s actual status is now recovered from these cases too, so a caller-side or provider-side failure is no longer reported as an Opik fault.

  • LLM-as-judge scoring retries on empty provider responses and no longer hangs indefinitely — An empty structured-output response from the judge model previously fell outside the retry policy and cost the item its whole score. These are now retried, and judge calls carry an explicit timeout and retry budget so a stalled provider can no longer block scoring indefinitely.

  • Silent span-conversion failures in the ADK integration are now reported — When converting an ADK LlmResponse or extracting its usage failed, the span was still logged but silently missing output and usage, with nothing surfacing the failure. This is now logged as an error instead of being swallowed.

  • Fixed two production NullPointerExceptions — Deleting or batch-updating dataset items, and querying span metrics, could 500 when a filter used an operator not supported by its field, instead of returning the same 400 other endpoints already gave. Separately, a judge response with no text (for example from a content filter) could crash scoring for the whole trace instead of reporting that nothing was scored.

  • Prompts no longer leak across projects with the same name — A backward-compatibility fallback for looking up a prompt by name, when no project was specified, incorrectly matched a same-named prompt from any project in the workspace instead of only project-less legacy prompts, so a client without a project_name set could read, and unintentionally version, another project’s prompt.

  • Grouped chart series and tag colors are easier to tell apart — Two adjacent colors in the series palette were close enough to be confused in a thin chart line; one has been replaced with a more distinct color, so any label that resolved to that palette slot now renders more legibly wherever it appears — chart series, tag chips, and feedback scores.

  • Dataset item sorting by JSON fields now handles special characters correctly — Sorting by input.*, output.*, or metadata.* keys built the underlying query by interpolating the key text directly, unlike data.* sorting, which could corrupt results for keys containing quotes or other special characters. JSON sort keys are now passed as bound parameters like everywhere else.

Performance Improvements

  • Faster traces table on projects with many feedback score columns — The traces table builds one column per feedback-score name, and every cell — editable or not — was paying the cost of wiring up inline editing. Non-editable score columns now render a lighter read-only cell, cutting a multi-second main-thread block down substantially on projects with hundreds of score columns and rows.

Experiment Traces Are Now a Logs Tab

Experiment traces were previously reachable only through a small “Go to logs” link tucked into the metadata row, easy to miss — users who couldn’t find it would conclude nothing had been logged. They’re now a dedicated Logs tab on the experiment page, always scoped to that experiment or, when comparing, to every experiment in the comparison. Comparing experiments that live in different projects now shows an explicit message instead of silently dropping traces, since logs are read per project.

Bug Fixes & Improvements

  • LLM-as-judge online-eval rules score more consistently — Built-in judge prompt templates and the scoring engine disagreed on whether a judge’s answer should be a flat or nested JSON object, so a rule could non-deterministically return an empty score depending on which shape the model happened to pick. The engine now accepts both shapes, correctly parses quoted numbers ("0.8" no longer becomes 0.0) and common boolean spellings (yes/no, pass/fail, on/off), and reports — rather than silently drops — a score name the rule doesn’t recognize. Judge prompts are also no longer HTML-escaped before being sent, which had been corrupting values like base64 data URLs assembled from template variables.

  • Unsupported LLM features now return a clear 400 instead of a mysterious 500 — Requesting a capability the selected provider doesn’t support (for example, requiring tool choice against a model that can’t do it) previously failed with a generic server error and, for online-scoring rules, burned through the full retry budget before the message was dropped. It’s now reported immediately as a 400 naming the unsupported feature, for both regular and streaming requests.

  • Fixed a Vertex AI thread leak that could exhaust a backend pod — Every LLM call routed through Vertex AI built a brand-new client whose background threads were never released, which could accumulate to thousands of threads per pod over time and eventually require a restart. The client is now created and closed per request instead.

  • Optimization cost totals now include GEPA reflection spend — LLM calls the optimizer makes internally during GEPA reflection belong to no single trial, so they were billed but excluded from a run’s total cost, undercounting real spend. That spend is now traced, tagged with the run, and included in total_optimization_cost (opik-optimizer 3.2.0). Upgrade both packages together — pip install -U "opik-optimizer>=3.2.0" "gepa>=0.1.0" — raising gepa on its own breaks at runtime against the older reflection-template format that 3.2.0 no longer speaks.

  • Self-hosted: per-component image registry override, and a demo-data job fix — The Helm chart now lets you set a separate image registry for the backend, frontend, or Python backend individually, falling back to the chart-wide registry when unset. The demo-data job also now honors a component’s configured image repository instead of always defaulting to opik-python-backend.

Performance Improvements

  • Faster trace counts on projects using guardrails filters — The traces-count query that fires on every Logs page load always joined against guardrails results and sorted every row to deduplicate, even when no guardrails filter was applied. Both are now skipped unless actually needed, cutting rows read roughly in half and cutting measured P99 latency by about 90% on production workspaces.

  • Faster trace lookups by thread ID — Filtering traces or threads by thread ID wasn’t engaging an existing prefilter, so a lookup could scan a project’s entire span history just to find one thread’s traces. The prefilter now engages for thread-scoped lookups, cutting measured production query times from multiple seconds to under 200ms.