Skip to content
Closed
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
86c79d7
feat(evals): add agent tool-use evaluation harness
sudoKrishna Sep 29, 2026
af2e6ca
feat(evals): add live DeepSeek model runs to the agent eval suite
sudoKrishna Sep 29, 2026
433fd0f
test(evals): assert grounded behavior instead of exact phrasing in li…
sudoKrishna Sep 29, 2026
85525a4
feat(evals): run agent scenarios through the DAGExecutor
sudoKrishna Sep 29, 2026
fb7f54c
test(evals): assert executor block retry on agent failure
sudoKrishna Sep 29, 2026
c93426b
test(evals): assert executor model fallback on primary failure
sudoKrishna Sep 29, 2026
2032c4d
feat(evals): record and replay live model transcripts
sudoKrishna Sep 29, 2026
c12342a
feat(evals): add agent context eval suite
sudoKrishna Sep 29, 2026
26937ff
feat(evals): compare agent tool-use across models
sudoKrishna Sep 30, 2026
a22387d
feat(evals): report failed checks in the model comparison
sudoKrishna Oct 1, 2026
9d2ec19
feat(evals): add adversarial agent tool-use scenarios
sudoKrishna Oct 1, 2026
cdebc4d
test(evals): fix adversarial scenario design after the live run
sudoKrishna Oct 1, 2026
fc549db
test(evals): fix chain and near-duplicate scenarios for live models
sudoKrishna Oct 1, 2026
4590af8
feat(evals): add LLM-as-judge scoring
sudoKrishna Oct 1, 2026
031f5ab
Merge branch 'feat/evals-agent-context' into feat/evals-integration
sudoKrishna Oct 1, 2026
ae57d49
Merge branch 'feat/evals-model-compare' into feat/evals-integration
sudoKrishna Oct 1, 2026
6dace97
Merge branch 'feat/evals-adversarial-scenarios' into feat/evals-integ…
sudoKrishna Oct 1, 2026
5235110
Merge branch 'feat/evals-llm-judge' into feat/evals-integration
sudoKrishna Oct 1, 2026
f619479
test(evals): address review findings on the eval harness
sudoKrishna Oct 2, 2026
3d8de59
feat(evals): judge open-ended agent answers
sudoKrishna Oct 2, 2026
74faa92
feat(evals): record judge identity and refuse cross-identity comparisons
sudoKrishna Oct 2, 2026
2023d07
feat(evals): judge the remaining open-ended scenarios
sudoKrishna Oct 2, 2026
8bf9c01
fix(evals): address the judge review findings
sudoKrishna Oct 2, 2026
69af5d7
feat(evals): record and replay judge transcripts
sudoKrishna Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
test(evals): assert executor block retry on agent failure
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
  • Loading branch information
sudoKrishna committed Sep 29, 2026
commit fb7f54c73b9a8dbfe36fa951f3a500f3ec4967e2
15 changes: 9 additions & 6 deletions apps/sim/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,10 @@ Add a case to `EXECUTOR_SCENARIOS` in `executor-harness.ts`:
- `providerResponse` is what the mocked provider returns (`content`,
`toolCalls`, `tokens`).
- `expect` uses the loop's checks plus `resolvedInput` (a substring that must
reach the provider messages) and `succeeds` (expected `ExecutionResult.success`).
reach the provider messages), `succeeds` (expected `ExecutionResult.success`),
and `providerCalls` (exact provider call count).
- Set `agent.retry` to exercise the executor's per-block retry policy; make the
first `providerResponse` a `reject` and the retry lands on the next one.

Both suites write one report, so executor rows appear alongside loop rows.

Expand All @@ -132,8 +135,8 @@ model/tool time, first-response time, and token usage.
## Scope and next steps

Two harnesses share one result shape and report: the tool loop and the
`DAGExecutor`. The executor harness mocks the provider boundary, so the
executor's retry/fallback policy is not yet asserted; add a scenario with a
first-call rejection and a block retry config to cover it. Further expansion
(context/memory, model routing, subagent orchestration) is tracked as
follow-up work.
`DAGExecutor`. The executor suite covers block retry —
`executor-retries-failed-block` makes the first provider call reject, the block
is replayed, and the run completes. Model fallback (`fallbackModels`) is not
asserted yet. Further expansion (context/memory, model routing, subagent
orchestration) is tracked as follow-up work.
60 changes: 54 additions & 6 deletions apps/sim/evals/agent-tool-use/executor-harness.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import {
} from '@sim/testing/factories/serialized-block.factory'
import { providersMockFns } from '@sim/testing/mocks/providers.mock'
import { DAGExecutor } from '@/executor/execution/executor'
import type { SerializedWorkflow } from '@/serializer/types'
import type { SerializedBlock, SerializedWorkflow } from '@/serializer/types'
import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness'
import type {
AgentToolUseExpectations,
Expand All @@ -30,8 +30,10 @@ export interface ExecutorProviderToolCall {
result?: unknown
}

/** The provider response `executeProviderRequest` returns for one model call. */
/** One model call: either a response, or a rejection the block must recover from. */
export interface ExecutorProviderResponse {
/** When set, the call rejects with this message instead of resolving. */
reject?: string
content: string
model?: string
tokens?: { input?: number; output?: number; total?: number }
Expand All @@ -52,26 +54,30 @@ export interface ExecutorScenario {
systemPrompt?: string
userPrompt?: string
temperature?: number
/** Enables the executor's per-block retry policy for the Agent block. */
retry?: { enabled: boolean; maxTries: number; waitBetweenTriesMs: number }
}
/** One entry per model call; the last entry serves any extra fallback calls. */
/** One entry per model call; the last entry serves any extra/retry calls. */
providerResponse: ExecutorProviderResponse | ExecutorProviderResponse[]
expect: AgentToolUseExpectations & {
/** Substring that must appear in the messages sent to the provider. */
resolvedInput?: string
/** Expected `ExecutionResult.success`. */
succeeds?: boolean
/** Exact number of provider calls the executor made. */
providerCalls?: number
}
}

function buildWorkflow(scenario: ExecutorScenario): SerializedWorkflow {
const start = createSerializedBlock({
const start: SerializedBlock = createSerializedBlock({
id: 'start',
type: 'start_trigger',
name: 'Start',
})
/** The trigger handler claims a block whose metadata says it is a trigger. */
if (start.metadata) start.metadata.category = 'triggers'
const agent = createSerializedBlock({
const agent: SerializedBlock = createSerializedBlock({
id: 'agent',
type: 'agent',
name: 'Eval Agent',
Expand All @@ -85,6 +91,7 @@ function buildWorkflow(scenario: ExecutorScenario): SerializedWorkflow {
? { temperature: scenario.agent.temperature }
: {}),
}
if (scenario.agent.retry) agent.retry = scenario.agent.retry

return createSerializedWorkflow([start, agent], [{ source: 'start', target: 'agent' }])
}
Expand All @@ -109,6 +116,7 @@ export async function runExecutorScenario(
requests.push(request)
const response = responses[Math.min(callIndex, responses.length - 1)]
callIndex += 1
if (response.reject) throw new Error(response.reject)
return {
content: response.content,
model: response.model ?? scenario.agent.model,
Expand Down Expand Up @@ -156,7 +164,14 @@ export async function runExecutorScenario(
durationMs: typeof call.duration === 'number' ? call.duration : 0,
}))

const checks = scoreExpectations(scenario.expect, toolCalls, finalContent, 1, runError, mode)
const checks = scoreExpectations(
scenario.expect,
toolCalls,
finalContent,
requests.length,
runError,
mode
)

if (scenario.expect.resolvedInput !== undefined) {
const sent = JSON.stringify(requests)
Expand All @@ -175,6 +190,14 @@ export async function runExecutorScenario(
})
}

if (scenario.expect.providerCalls !== undefined) {
checks.push({
name: 'provider-calls',
passed: requests.length === scenario.expect.providerCalls,
detail: `expected ${scenario.expect.providerCalls}, got ${requests.length}`,
})
}

const tokens = (output.tokens ?? {}) as { input?: number; output?: number; total?: number }

return {
Expand Down Expand Up @@ -259,4 +282,29 @@ export const EXECUTOR_SCENARIOS: ExecutorScenario[] = [
finalContent: /A-1937/,
},
},
{
id: 'executor-retries-failed-block',
name: 'retries a failed Agent block and completes the run',
category: 'recovery',
description:
'The first provider call rejects with a 503. The block has retry enabled, so the executor replays it and the second call succeeds — proving the executor retry policy, not the agent handler, recovered the turn.',
workflowInput: { message: 'What is the API rate limit?' },
agent: {
model: 'gpt-4o',
userPrompt: 'What is the API rate limit?',
retry: { enabled: true, maxTries: 3, waitBetweenTriesMs: 0 },
},
providerResponse: [
{ reject: '503 Service Unavailable', content: '' },
{
content: 'The API rate limit is 100 requests per minute.',
tokens: { input: 10, output: 20, total: 30 },
},
],
expect: {
succeeds: true,
finalContent: '100 requests per minute',
providerCalls: 2,
},
},
]