Skip to content
Closed
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
86c79d7
feat(evals): add agent tool-use evaluation harness
sudoKrishna Sep 29, 2026
af2e6ca
feat(evals): add live DeepSeek model runs to the agent eval suite
sudoKrishna Sep 29, 2026
433fd0f
test(evals): assert grounded behavior instead of exact phrasing in li…
sudoKrishna Sep 29, 2026
85525a4
feat(evals): run agent scenarios through the DAGExecutor
sudoKrishna Sep 29, 2026
fb7f54c
test(evals): assert executor block retry on agent failure
sudoKrishna Sep 29, 2026
c93426b
test(evals): assert executor model fallback on primary failure
sudoKrishna Sep 29, 2026
2032c4d
feat(evals): record and replay live model transcripts
sudoKrishna Sep 29, 2026
c12342a
feat(evals): add agent context eval suite
sudoKrishna Sep 29, 2026
26937ff
feat(evals): compare agent tool-use across models
sudoKrishna Sep 30, 2026
a22387d
feat(evals): report failed checks in the model comparison
sudoKrishna Oct 1, 2026
9d2ec19
feat(evals): add adversarial agent tool-use scenarios
sudoKrishna Oct 1, 2026
cdebc4d
test(evals): fix adversarial scenario design after the live run
sudoKrishna Oct 1, 2026
fc549db
test(evals): fix chain and near-duplicate scenarios for live models
sudoKrishna Oct 1, 2026
4590af8
feat(evals): add LLM-as-judge scoring
sudoKrishna Oct 1, 2026
031f5ab
Merge branch 'feat/evals-agent-context' into feat/evals-integration
sudoKrishna Oct 1, 2026
ae57d49
Merge branch 'feat/evals-model-compare' into feat/evals-integration
sudoKrishna Oct 1, 2026
6dace97
Merge branch 'feat/evals-adversarial-scenarios' into feat/evals-integ…
sudoKrishna Oct 1, 2026
5235110
Merge branch 'feat/evals-llm-judge' into feat/evals-integration
sudoKrishna Oct 1, 2026
f619479
test(evals): address review findings on the eval harness
sudoKrishna Oct 2, 2026
3d8de59
feat(evals): judge open-ended agent answers
sudoKrishna Oct 2, 2026
74faa92
feat(evals): record judge identity and refuse cross-identity comparisons
sudoKrishna Oct 2, 2026
2023d07
feat(evals): judge the remaining open-ended scenarios
sudoKrishna Oct 2, 2026
8bf9c01
fix(evals): address the judge review findings
sudoKrishna Oct 2, 2026
69af5d7
feat(evals): record and replay judge transcripts
sudoKrishna Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
feat(evals): run agent scenarios through the DAGExecutor
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
  • Loading branch information
sudoKrishna committed Sep 29, 2026
commit 85525a412210c54290a9583b060bd4b3aa7db238
32 changes: 27 additions & 5 deletions apps/sim/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,27 @@ A scenario is data, not code — there is no harness change needed for a new cas
string. The loop must not execute the call and must return the parse error to
the model.

## Executor-level scenarios

[`agent-tool-use/executor-harness.ts`](./agent-tool-use/executor-harness.ts)
runs a case through a real `DAGExecutor`: a Start block → Agent block workflow,
with only the provider boundary (`executeProviderRequest`) mocked. This covers
what the loop harness cannot — agent-block input wiring, variable resolution
from Start outputs, and the executor's run/error handling. Tool dispatch stays
covered by the loop suite.

Add a case to `EXECUTOR_SCENARIOS` in `executor-harness.ts`:

- `workflowInput` is exposed on the Start block; reference an output with
`<start.field>` from the Agent prompt.
- `agent` is the Agent block config (`model`, `systemPrompt`, `userPrompt`).
- `providerResponse` is what the mocked provider returns (`content`,
`toolCalls`, `tokens`).
- `expect` uses the loop's checks plus `resolvedInput` (a substring that must
reach the provider messages) and `succeeds` (expected `ExecutionResult.success`).

Both suites write one report, so executor rows appear alongside loop rows.

## Report shape

`report.json` is machine-readable for dashboards and trend tracking; `report.md`
Expand All @@ -110,8 +131,9 @@ model/tool time, first-response time, and token usage.

## Scope and next steps

This suite evaluates the tool loop directly. The next layer is a scenario that
runs the same scripted model through the full `DAGExecutor` so agent block
wiring, variable resolution, and the executor's retry/fallback policy are
measured alongside the loop. The `AgentToolUseResult` shape is deliberately
independent of the harness entry point so both can share scoring and reporting.
Two harnesses share one result shape and report: the tool loop and the
`DAGExecutor`. The executor harness mocks the provider boundary, so the
executor's retry/fallback policy is not yet asserted; add a scenario with a
first-call rejection and a block retry config to cover it. Further expansion
(context/memory, model routing, subagent orchestration) is tracked as
follow-up work.
51 changes: 49 additions & 2 deletions apps/sim/evals/agent-tool-use/agent-tool-use.eval.test.ts
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
import {
permissionCheckMock,
permissionCheckMockFns,
} from '@sim/testing/mocks/permission-check.mock'
import { providersMock } from '@sim/testing/mocks/providers.mock'
import { providersConversationHistoryMock } from '@sim/testing/mocks/providers-conversation-history.mock'
import { providersUtilsMock } from '@sim/testing/mocks/providers-utils.mock'
import { providersUtilsMock, providersUtilsMockFns } from '@sim/testing/mocks/providers-utils.mock'
import { toolsMock } from '@sim/testing/mocks/tools.mock'
import { afterAll, describe, expect, it, vi } from 'vitest'
import { workspaceFileSecretProvenanceMock } from '@sim/testing/mocks/workspace-file-secret-provenance.mock'
import { afterAll, beforeEach, describe, expect, it, vi } from 'vitest'
import { EXECUTOR_SCENARIOS, runExecutorScenario } from '@/evals/agent-tool-use/executor-harness'
import { runScenario } from '@/evals/agent-tool-use/harness'
import { writeEvalReport } from '@/evals/agent-tool-use/report'
import { AGENT_TOOL_USE_SCENARIOS } from '@/evals/agent-tool-use/scenarios'
Expand All @@ -12,9 +18,37 @@ vi.mock('@/providers/conversation-history', () => providersConversationHistoryMo
vi.mock('@/tools', () => toolsMock)
vi.mock('@/providers/utils', () => providersUtilsMock)
vi.mock('@/providers', () => providersMock)
vi.mock('@/ee/access-control/utils/permission-check', () => permissionCheckMock)
vi.mock(
'@/lib/uploads/contexts/workspace/workspace-file-secret-provenance',
() => workspaceFileSecretProvenanceMock
)
vi.mock('@/lib/memory/agent-turn-session', () => ({
openAgentTurnSession: vi.fn(async () => undefined),
}))
vi.mock('@/lib/internal/mcp/discover-tools', () => ({
discoverMcpServerToolsAsExecutor: vi.fn(async () => []),
}))
vi.mock('@/lib/internal/custom-tools/read-available-by-id-or-title', () => ({
readAvailableCustomToolByIdOrTitleAsExecutor: vi.fn(async () => undefined),
}))
vi.mock('@/executor/utils/http', () => ({
buildAuthHeaders: vi.fn(async () => ({ 'Content-Type': 'application/json' })),
buildAPIUrl: vi.fn((path: string) => path),
extractAPIErrorMessage: vi.fn(async () => 'request failed'),
}))
vi.mock('@/lib/execution/cancellation', () => ({
subscribeToExecutionCancellation: vi.fn(async () => () => {}),
isExecutionCancelled: vi.fn(async () => false),
}))

const results: AgentToolUseResult[] = []

beforeEach(() => {
permissionCheckMockFns.mockValidateModelProvider.mockResolvedValue(undefined)
providersUtilsMockFns.mockGetProviderFromModel.mockReturnValue('mock-provider')
})

afterAll(() => {
const reportPath = process.env.EVAL_REPORT_PATH
if (reportPath) writeEvalReport(results, reportPath)
Expand All @@ -32,3 +66,16 @@ describe('agent tool-use eval suite', () => {
).toEqual([])
})
})

describe('agent executor eval suite', () => {
it.each(EXECUTOR_SCENARIOS)('$id: $name', async (scenario) => {
const result = await runExecutorScenario(scenario)
results.push(result)

const failed = result.checks.filter((entry) => !entry.passed)
expect(
failed,
failed.map((entry) => `${entry.name}: ${entry.detail}`).join('; ') || undefined
).toEqual([])
})
})
262 changes: 262 additions & 0 deletions apps/sim/evals/agent-tool-use/executor-harness.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,262 @@
import {
createSerializedBlock,
createSerializedWorkflow,
} from '@sim/testing/factories/serialized-block.factory'
import { providersMockFns } from '@sim/testing/mocks/providers.mock'
import { DAGExecutor } from '@/executor/execution/executor'
import type { SerializedWorkflow } from '@/serializer/types'
import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness'
import type {
AgentToolUseExpectations,
AgentToolUseResult,
EvalCategory,
EvalToolInvocation,
} from './types'

/**
* Executor-level harness.
*
* Drives a real `DAGExecutor` run: Start block → Agent block. The provider
* boundary (`executeProviderRequest`) is the only thing mocked — the Agent
* block handler, input/variable resolution, and the executor run/error handling
* are real. Tool calls are what the mocked provider returns; tool *dispatch* is
* covered by the loop harness.
*/

/** One tool call the mocked provider reports in its response. */
export interface ExecutorProviderToolCall {
name: string
arguments?: Record<string, unknown>
result?: unknown
}

/** The provider response `executeProviderRequest` returns for one model call. */
export interface ExecutorProviderResponse {
content: string
model?: string
tokens?: { input?: number; output?: number; total?: number }
toolCalls?: ExecutorProviderToolCall[]
cost?: unknown
timing?: unknown
}

export interface ExecutorScenario {
id: string
name: string
category: EvalCategory
description: string
/** Exposed on the Start block and referenced from the Agent block. */
workflowInput: Record<string, unknown>
agent: {
model: string
systemPrompt?: string
userPrompt?: string
temperature?: number
}
/** One entry per model call; the last entry serves any extra fallback calls. */
providerResponse: ExecutorProviderResponse | ExecutorProviderResponse[]
expect: AgentToolUseExpectations & {
/** Substring that must appear in the messages sent to the provider. */
resolvedInput?: string
/** Expected `ExecutionResult.success`. */
succeeds?: boolean
}
}

function buildWorkflow(scenario: ExecutorScenario): SerializedWorkflow {
const start = createSerializedBlock({
id: 'start',
type: 'start_trigger',
name: 'Start',
})
/** The trigger handler claims a block whose metadata says it is a trigger. */
if (start.metadata) start.metadata.category = 'triggers'
const agent = createSerializedBlock({
id: 'agent',
type: 'agent',
name: 'Eval Agent',
})
agent.config.tool = 'agent'
agent.config.params = {
model: scenario.agent.model,
systemPrompt: scenario.agent.systemPrompt,
userPrompt: scenario.agent.userPrompt,
...(scenario.agent.temperature !== undefined
? { temperature: scenario.agent.temperature }
: {}),
}

return createSerializedWorkflow([start, agent], [{ source: 'start', target: 'agent' }])
}

/**
* Runs one executor scenario and scores it with the shared scorer, returning
* the same result shape as the loop harness so both land in one report.
*/
export async function runExecutorScenario(
scenario: ExecutorScenario,
options: { mode?: EvalRunMode } = {}
): Promise<AgentToolUseResult> {
const mode = options.mode ?? 'scripted'
const responses = Array.isArray(scenario.providerResponse)
? [...scenario.providerResponse]
: [scenario.providerResponse]
const requests: Array<Record<string, unknown>> = []
let callIndex = 0

providersMockFns.mockExecuteProviderRequest.mockImplementation(
async (_providerId: string, request: Record<string, unknown>) => {
requests.push(request)
const response = responses[Math.min(callIndex, responses.length - 1)]
callIndex += 1
return {
content: response.content,
model: response.model ?? scenario.agent.model,
tokens: response.tokens ?? { input: 0, output: 0, total: 0 },
toolCalls: response.toolCalls ?? [],
cost: response.cost ?? 0,
timing: response.timing ?? { total: 0 },
}
}
)

const executor = new DAGExecutor({
workflow: buildWorkflow(scenario),
workflowInput: scenario.workflowInput,
contextExtensions: {
workspaceId: 'eval-workspace',
executionId: 'eval-execution',
userId: 'eval-user',
},
})

let result: { success?: boolean; output?: Record<string, unknown> } | undefined
let runError: unknown
const startedAt = Date.now()
try {
result = (await executor.execute('eval-workflow')) as typeof result
} catch (error) {
runError = error
}
const latencyMs = Date.now() - startedAt

const output = (result?.output ?? {}) as Record<string, unknown>
const finalContent = typeof output.content === 'string' ? output.content : ''
const rawToolCalls = ((output.toolCalls as { list?: unknown[] } | undefined)?.list ??
[]) as Array<Record<string, unknown>>

const toolCalls: ScoredToolCall[] = rawToolCalls.map((call) => ({
name: typeof call.name === 'string' ? call.name : 'unknown',
success: true,
}))
const toolInvocations: EvalToolInvocation[] = rawToolCalls.map((call) => ({
name: typeof call.name === 'string' ? call.name : 'unknown',
arguments: (call.arguments ?? {}) as Record<string, unknown>,
success: true,
durationMs: typeof call.duration === 'number' ? call.duration : 0,
}))

const checks = scoreExpectations(scenario.expect, toolCalls, finalContent, 1, runError, mode)

if (scenario.expect.resolvedInput !== undefined) {
const sent = JSON.stringify(requests)
checks.push({
name: 'resolved-input',
passed: sent.includes(scenario.expect.resolvedInput),
detail: `looking for ${JSON.stringify(scenario.expect.resolvedInput)} in provider messages`,
})
}

if (scenario.expect.succeeds !== undefined) {
checks.push({
name: 'workflow-success',
passed: result?.success === scenario.expect.succeeds,
detail: `success=${String(result?.success)}`,
})
}

const tokens = (output.tokens ?? {}) as { input?: number; output?: number; total?: number }

return {
id: scenario.id,
name: scenario.name,
category: scenario.category,
passed: checks.every((entry) => entry.passed),
checks,
finalContent,
toolInvocations,
metrics: {
iterations: requests.length,
toolCalls: toolCalls.length,
successfulToolCalls: toolCalls.filter((call) => call.success).length,
erroredToolCalls: 0,
latencyMs,
modelTimeMs: 0,
toolsTimeMs: 0,
firstResponseTimeMs: 0,
inputTokens: tokens.input ?? 0,
outputTokens: tokens.output ?? 0,
totalTokens: tokens.total ?? 0,
},
...(runError ? { error: String(runError) } : {}),
}
}

/**
* Executor-level scenarios. Two cover the wiring the loop suite cannot see:
* Start → Agent execution, and variable resolution from a Start output into the
* Agent's prompt.
*/
export const EXECUTOR_SCENARIOS: ExecutorScenario[] = [
{
id: 'executor-agent-runs',
name: 'runs a Start → Agent workflow and surfaces the Agent output',
category: 'tool-selection',
description:
'The real Agent block handler runs inside the DAG. The mocked provider reports one tool call; the executor result must carry the content and the tool call through.',
workflowInput: { message: 'What is the API rate limit?' },
agent: {
model: 'gpt-4o',
systemPrompt: 'You are a documentation assistant.',
userPrompt: 'What is the API rate limit?',
},
providerResponse: {
content: 'The API rate limit is 100 requests per minute.',
toolCalls: [
{
name: 'search_docs',
arguments: { query: 'api rate limit' },
result: { snippet: 'The API rate limit is 100 requests per minute.' },
},
],
tokens: { input: 10, output: 20, total: 30 },
},
expect: {
succeeds: true,
finalContent: '100 requests per minute',
toolCallSequence: ['search_docs'],
successfulToolCalls: 1,
},
},
{
id: 'executor-resolves-start-input',
name: 'resolves a Start output into the Agent prompt before the provider call',
category: 'planning',
description:
'The Agent userPrompt references <start.message>. The value must be resolved by the executor and reach the provider request, not passed through verbatim.',
workflowInput: { message: 'Summarize order A-1937' },
agent: {
model: 'gpt-4o',
userPrompt: '<start.message>',
},
providerResponse: {
content: 'Order A-1937 shipped via DHL.',
tokens: { input: 8, output: 12, total: 20 },
},
expect: {
succeeds: true,
resolvedInput: 'Summarize order A-1937',
finalContent: /A-1937/,
},
},
]
Loading