feat(feedback): additional_instructions and cot reasons for conversation metrics - #2710
Merged
sfc-gh-jreini merged 1 commit intoAug 20, 2026
Conversation
sfc-gh-jreini
approved these changes
Aug 20, 2026
sfc-gh-jreini
left a comment
Contributor
There was a problem hiding this comment.
This is a fantastic contribution that solves the issue well, is nicely congruent with the codebase and is well documented. Thank you!
3 tasks done
joshreini1
pushed a commit
that referenced
this pull request
Aug 25, 2026
PR #2710 inserted additional_instructions before temperature, so positional temperature calls bound the float to instructions and raised TypeError. Restore temperature to its original slot and make additional_instructions keyword-only. Fixes #2727 Co-authored-by: Yuzhong Zhang <BetterAndBetterII@users.noreply.github.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #2699.
The four conversation-level metrics added in 2.12 (
coherence_across_turns,conversation_helpfulness,topic_adherence,agent_goal_accuracy) had no wayto pass domain context to the judge and returned a bare float. This adds an
additional_instructionsparameter to all four, and adds a_with_cot_reasonsvariant of each that returnsTuple[float, Dict].additional_instructionsmatters more at conversation scope than at turn scopebecause the judge reads a whole transcript at once. When assistant turns are
structured output such as SQL result rows or JSON, the default prompts are
calibrated for prose and score correct output as incoherent. The reasoning
output matters because a single number over a five turn conversation does not
say which turn broke the thread.
Both additions follow the idiom already used by the turn-level metrics in the
same file.
additional_instructionsis threaded through_build_criteria_with_instructions, which is howcoherence,harmfulness,correctnessandconcisenessalready fold instructions into a template'sdefault system prompt. The chain of thought variants append
templates_base.COT_REASONS_TEMPLATEto the user prompt and callgenerate_score_and_reasons, which is howsentiment_with_cot_reasonsand theagent trace metrics do it. Parameter order, type hints, docstring layout and
return types match the turn-level methods. The score ranges are unchanged, so
agent_goal_accuracy_with_cot_reasonsstays binary while the other three stayon the 0 to 3 Likert scale.
Other details good to know for developers
additional_instructionsis placed after the existing domain parameters(
reference_topics,reference_goal) rather than immediately afterrecords,so existing positional calls keep working.
The turn-level metrics also carry a
**kwargsshim that maps the deprecatedcustom_instructionsname ontoadditional_instructions. I left that off hereon purpose. These four methods never accepted
custom_instructions, so a shimwould introduce a deprecated alias rather than preserve one, and
Metricalready rewrites that name before it reaches the provider. Happy to add it if
you would rather have the signatures byte-identical.
tests/unit/static/golden/api.trulens.3.11.yamlis not regenerated in this PR.That golden file already predates the 2.12 and 2.13 provider additions: it is
missing
coherence_across_turns,conversation_helpfulness,topic_adherence,agent_goal_accuracyandcitation_accuracyonmaintoday. Regenerating ithere would fold that unrelated drift into this change, so it seems better for a
maintainer to run
make write-apiseparately.The docs page
component_guides/instrumentation/conversation_evaluation.mdgainstwo short subsections showing
additional_instructionson a conversation metricand the
_with_cot_reasonsvariant.Test coverage
New tests in
tests/unit/test_feedback_criteria_and_additional_instructions.pyreuse the existing
MockLLMProvider, which captures the prompts the providerwould have sent. They assert that
additional_instructionsreaches the systemprompt for all four metrics, that the default system prompt is untouched when
the parameter is omitted, that
reference_topicsandreference_goalsurvivealongside it, and that
Metric(additional_instructions=...)plumbs through to aconversation metric end to end. A second class covers the chain of thought
variants: tuple return shape, reasons dictionary, the reasons template landing
in the user prompt, the transcript landing in the user prompt, and a plain
transcript string being accepted in place of records.
New tests in
tests/unit/test_templates_conversation.pyfollow the autospecstyle already in that file and pin the score arguments the new variants pass to
generate_score_and_reasons, including the binary range foragent_goal_accuracy_with_cot_reasons.What was verified and what was not
Verified on Python 3.11 with
trulens-coreandtrulens-feedbackinstalledfrom this tree:
tests/unit/test_templates_conversation.pyandtests/unit/test_feedback_criteria_and_additional_instructions.py, 32 passed.The wider
tests/unitrun went from 488 passed to 501 passed with the samesingle pre-existing failure in
test_otel_async_concurrency.py::test_cancelled_finalization_failure_preserves_cancelled_error,which reproduces on an unmodified
mainin a full-directory run and passes inisolation on both.
ruff checkandruff format --checkare clean on the threechanged Python files, run with ruff 0.5.5 to match the pinned pre-commit rev.
Not verified:
tests/unit/static/test_api.py, the streamlit and dashboardtests,
test_gepa_fitness.py,test_records_utils.pyandtest_reward.py, allof which need optional dependencies that are not installed here; any test
requiring provider API keys; and live judge behaviour against a real model. The
new prompt text has not been run through an actual LLM, so the wording of the
appended instructions has not been quality-checked end to end.
Type of change
not work as expected)