Skip to content

feat(feedback): additional_instructions and cot reasons for conversation metrics - #2710

Merged
sfc-gh-jreini merged 1 commit into
truera:mainfrom
dylanpulver:feat/conversation-metrics-additional-instructions
Aug 20, 2026
Merged

sfc-gh-jreini merged 1 commit into
truera:mainfrom
dylanpulver:feat/conversation-metrics-additional-instructions

Conversation

@dylanpulver

Copy link
Copy Markdown
Contributor

Description

Fixes #2699.

The four conversation-level metrics added in 2.12 (coherence_across_turns,
conversation_helpfulness, topic_adherence, agent_goal_accuracy) had no way
to pass domain context to the judge and returned a bare float. This adds an
additional_instructions parameter to all four, and adds a
_with_cot_reasons variant of each that returns Tuple[float, Dict].

additional_instructions matters more at conversation scope than at turn scope
because the judge reads a whole transcript at once. When assistant turns are
structured output such as SQL result rows or JSON, the default prompts are
calibrated for prose and score correct output as incoherent. The reasoning
output matters because a single number over a five turn conversation does not
say which turn broke the thread.

Both additions follow the idiom already used by the turn-level metrics in the
same file. additional_instructions is threaded through
_build_criteria_with_instructions, which is how coherence, harmfulness,
correctness and conciseness already fold instructions into a template's
default system prompt. The chain of thought variants append
templates_base.COT_REASONS_TEMPLATE to the user prompt and call
generate_score_and_reasons, which is how sentiment_with_cot_reasons and the
agent trace metrics do it. Parameter order, type hints, docstring layout and
return types match the turn-level methods. The score ranges are unchanged, so
agent_goal_accuracy_with_cot_reasons stays binary while the other three stay
on the 0 to 3 Likert scale.

Other details good to know for developers

additional_instructions is placed after the existing domain parameters
(reference_topics, reference_goal) rather than immediately after records,
so existing positional calls keep working.

The turn-level metrics also carry a **kwargs shim that maps the deprecated
custom_instructions name onto additional_instructions. I left that off here
on purpose. These four methods never accepted custom_instructions, so a shim
would introduce a deprecated alias rather than preserve one, and Metric
already rewrites that name before it reaches the provider. Happy to add it if
you would rather have the signatures byte-identical.

tests/unit/static/golden/api.trulens.3.11.yaml is not regenerated in this PR.
That golden file already predates the 2.12 and 2.13 provider additions: it is
missing coherence_across_turns, conversation_helpfulness, topic_adherence,
agent_goal_accuracy and citation_accuracy on main today. Regenerating it
here would fold that unrelated drift into this change, so it seems better for a
maintainer to run make write-api separately.

The docs page component_guides/instrumentation/conversation_evaluation.md gains
two short subsections showing additional_instructions on a conversation metric
and the _with_cot_reasons variant.

Test coverage

New tests in tests/unit/test_feedback_criteria_and_additional_instructions.py
reuse the existing MockLLMProvider, which captures the prompts the provider
would have sent. They assert that additional_instructions reaches the system
prompt for all four metrics, that the default system prompt is untouched when
the parameter is omitted, that reference_topics and reference_goal survive
alongside it, and that Metric(additional_instructions=...) plumbs through to a
conversation metric end to end. A second class covers the chain of thought
variants: tuple return shape, reasons dictionary, the reasons template landing
in the user prompt, the transcript landing in the user prompt, and a plain
transcript string being accepted in place of records.

New tests in tests/unit/test_templates_conversation.py follow the autospec
style already in that file and pin the score arguments the new variants pass to
generate_score_and_reasons, including the binary range for
agent_goal_accuracy_with_cot_reasons.

What was verified and what was not

Verified on Python 3.11 with trulens-core and trulens-feedback installed
from this tree: tests/unit/test_templates_conversation.py and
tests/unit/test_feedback_criteria_and_additional_instructions.py, 32 passed.
The wider tests/unit run went from 488 passed to 501 passed with the same
single pre-existing failure in
test_otel_async_concurrency.py::test_cancelled_finalization_failure_preserves_cancelled_error,
which reproduces on an unmodified main in a full-directory run and passes in
isolation on both. ruff check and ruff format --check are clean on the three
changed Python files, run with ruff 0.5.5 to match the pinned pre-commit rev.

Not verified: tests/unit/static/test_api.py, the streamlit and dashboard
tests, test_gepa_fitness.py, test_records_utils.py and test_reward.py, all
of which need optional dependencies that are not installed here; any test
requiring provider API keys; and live judge behaviour against a real model. The
new prompt text has not been run through an actual LLM, so the wording of the
appended instructions has not been quality-checked end to end.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to
    not work as expected)
  • New Tests
  • This change includes re-generated golden test results
  • This change requires a documentation update

@dosubot dosubot Bot added size:L This PR changes 100-499 lines, ignoring generated files. documentation Improvements or additions to documentation labels Aug 19, 2026

@sfc-gh-jreini sfc-gh-jreini left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a fantastic contribution that solves the issue well, is nicely congruent with the codebase and is well documented. Thank you!

@sfc-gh-jreini
sfc-gh-jreini merged commit 19a3b9c into truera:main Aug 20, 2026
9 checks passed
joshreini1 pushed a commit that referenced this pull request Aug 25, 2026
PR #2710 inserted additional_instructions before temperature, so
positional temperature calls bound the float to instructions and
raised TypeError. Restore temperature to its original slot and make
additional_instructions keyword-only.

Fixes #2727

Co-authored-by: Yuzhong Zhang <BetterAndBetterII@users.noreply.github.com>
@sfc-gh-jreini sfc-gh-jreini mentioned this pull request Sep 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEAT] additional_instructions and _with_cot_reasons Support for Conversation-Level Metrics

2 participants