Skip to content
Closed
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
86c79d7
feat(evals): add agent tool-use evaluation harness
sudoKrishna Sep 29, 2026
af2e6ca
feat(evals): add live DeepSeek model runs to the agent eval suite
sudoKrishna Sep 29, 2026
433fd0f
test(evals): assert grounded behavior instead of exact phrasing in li…
sudoKrishna Sep 29, 2026
85525a4
feat(evals): run agent scenarios through the DAGExecutor
sudoKrishna Sep 29, 2026
fb7f54c
test(evals): assert executor block retry on agent failure
sudoKrishna Sep 29, 2026
c93426b
test(evals): assert executor model fallback on primary failure
sudoKrishna Sep 29, 2026
2032c4d
feat(evals): record and replay live model transcripts
sudoKrishna Sep 29, 2026
c12342a
feat(evals): add agent context eval suite
sudoKrishna Sep 29, 2026
26937ff
feat(evals): compare agent tool-use across models
sudoKrishna Sep 30, 2026
a22387d
feat(evals): report failed checks in the model comparison
sudoKrishna Oct 1, 2026
9d2ec19
feat(evals): add adversarial agent tool-use scenarios
sudoKrishna Oct 1, 2026
cdebc4d
test(evals): fix adversarial scenario design after the live run
sudoKrishna Oct 1, 2026
fc549db
test(evals): fix chain and near-duplicate scenarios for live models
sudoKrishna Oct 1, 2026
4590af8
feat(evals): add LLM-as-judge scoring
sudoKrishna Oct 1, 2026
031f5ab
Merge branch 'feat/evals-agent-context' into feat/evals-integration
sudoKrishna Oct 1, 2026
ae57d49
Merge branch 'feat/evals-model-compare' into feat/evals-integration
sudoKrishna Oct 1, 2026
6dace97
Merge branch 'feat/evals-adversarial-scenarios' into feat/evals-integ…
sudoKrishna Oct 1, 2026
5235110
Merge branch 'feat/evals-llm-judge' into feat/evals-integration
sudoKrishna Oct 1, 2026
f619479
test(evals): address review findings on the eval harness
sudoKrishna Oct 2, 2026
3d8de59
feat(evals): judge open-ended agent answers
sudoKrishna Oct 2, 2026
74faa92
feat(evals): record judge identity and refuse cross-identity comparisons
sudoKrishna Oct 2, 2026
2023d07
feat(evals): judge the remaining open-ended scenarios
sudoKrishna Oct 2, 2026
8bf9c01
fix(evals): address the judge review findings
sudoKrishna Oct 2, 2026
69af5d7
feat(evals): record and replay judge transcripts
sudoKrishna Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
feat(evals): add agent context eval suite
Drive the Agent block through the executor with conversation memory on. The
memory read is stubbed per conversation id, so the provider request shows what
the handler assembled: prior history, then the new prompt, system prompt
preserved, correct conversation id. A wrong id surfaces as missing history and
fails (checked locally).

- agent-context/scenarios.ts: two context scenarios
- executor-harness.ts: memory seam + assembly/isolation checks
- test:evals:context script; README documents the suite
  • Loading branch information
sudoKrishna committed Sep 29, 2026
commit c12342a6701a570320e93bbf694fb174229f66ed
19 changes: 19 additions & 0 deletions apps/sim/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,25 @@ Add a case to `EXECUTOR_SCENARIOS` in `executor-harness.ts`:

Both suites write one report, so executor rows appear alongside loop rows.

## Context evals

[`agent-context/`](./agent-context/) drives the Agent block through the executor
with conversation memory on. The memory read is stubbed per conversation id, so
the provider request shows exactly what the handler assembled: prior history,
then the new user prompt, with the system prompt preserved, and the conversation
id must match. Windowing inside the memory service (`sliding_window`, token
budgets) is covered by its unit tests; this suite covers the assembly the model
sees.

```sh
cd apps/sim
bun run test:evals:context # writes test-results/evals/agent-context.{json,md}
```

Scenarios live in [`agent-context/scenarios.ts`](./agent-context/scenarios.ts)
and reuse the executor harness, so a case is the same shape as an executor case
plus `agent.memory`.

## Report shape

`report.json` is machine-readable for dashboards and trend tracking; `report.md`
Expand Down
67 changes: 67 additions & 0 deletions apps/sim/evals/agent-context/agent-context.eval.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
import {
permissionCheckMock,
permissionCheckMockFns,
} from '@sim/testing/mocks/permission-check.mock'
import { providersMock } from '@sim/testing/mocks/providers.mock'
import { providersConversationHistoryMock } from '@sim/testing/mocks/providers-conversation-history.mock'
import { providersUtilsMock, providersUtilsMockFns } from '@sim/testing/mocks/providers-utils.mock'
import { toolsMock } from '@sim/testing/mocks/tools.mock'
import { workspaceFileSecretProvenanceMock } from '@sim/testing/mocks/workspace-file-secret-provenance.mock'
import { afterAll, beforeEach, describe, expect, it, vi } from 'vitest'
import { AGENT_CONTEXT_SCENARIOS } from '@/evals/agent-context/scenarios'
import { runExecutorScenario } from '@/evals/agent-tool-use/executor-harness'
import { writeEvalReport } from '@/evals/agent-tool-use/report'
import type { AgentToolUseResult } from '@/evals/agent-tool-use/types'

vi.mock('@/providers/conversation-history', () => providersConversationHistoryMock)
vi.mock('@/tools', () => toolsMock)
vi.mock('@/providers/utils', () => providersUtilsMock)
vi.mock('@/providers', () => providersMock)
vi.mock('@/ee/access-control/utils/permission-check', () => permissionCheckMock)
vi.mock(
'@/lib/uploads/contexts/workspace/workspace-file-secret-provenance',
() => workspaceFileSecretProvenanceMock
)
vi.mock('@/lib/memory/agent-turn-session', () => ({
openAgentTurnSession: vi.fn(async () => undefined),
}))
vi.mock('@/lib/internal/mcp/discover-tools', () => ({
discoverMcpServerToolsAsExecutor: vi.fn(async () => []),
}))
vi.mock('@/lib/internal/custom-tools/read-available-by-id-or-title', () => ({
readAvailableCustomToolByIdOrTitleAsExecutor: vi.fn(async () => undefined),
}))
vi.mock('@/executor/utils/http', () => ({
buildAuthHeaders: vi.fn(async () => ({ 'Content-Type': 'application/json' })),
buildAPIUrl: vi.fn((path: string) => path),
extractAPIErrorMessage: vi.fn(async () => 'request failed'),
}))
vi.mock('@/lib/execution/cancellation', () => ({
subscribeToExecutionCancellation: vi.fn(async () => () => {}),
isExecutionCancelled: vi.fn(async () => false),
}))

const results: AgentToolUseResult[] = []

beforeEach(() => {
permissionCheckMockFns.mockValidateModelProvider.mockResolvedValue(undefined)
providersUtilsMockFns.mockGetProviderFromModel.mockReturnValue('mock-provider')
})

afterAll(() => {
const reportPath = process.env.EVAL_CONTEXT_REPORT_PATH
if (reportPath) writeEvalReport(results, reportPath)
})

describe('agent context eval suite', () => {
it.each(AGENT_CONTEXT_SCENARIOS)('$id: $name', async (scenario) => {
const result = await runExecutorScenario(scenario)
results.push(result)

const failed = result.checks.filter((entry) => !entry.passed)
expect(
failed,
failed.map((entry) => `${entry.name}: ${entry.detail}`).join('; ') || undefined
).toEqual([])
})
})
71 changes: 71 additions & 0 deletions apps/sim/evals/agent-context/scenarios.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
import type { ExecutorScenario } from '@/evals/agent-tool-use/executor-harness'

/**
* Agent context evals.
*
* These drive the real Agent block inside the `DAGExecutor` with conversation
* memory on. The memory read is stubbed per conversation id, so the provider
* request shows exactly what the handler assembled: prior history, then the new
* user prompt, with the system prompt preserved. A handler that passed the
* wrong conversation id, dropped history, or reordered the prompt fails.
*
* Windowing inside the memory service (`sliding_window`, token budgets) is
* covered by its own unit tests; this suite covers the assembly the model sees.
*/
export const AGENT_CONTEXT_SCENARIOS: ExecutorScenario[] = [
{
id: 'context-remembers-prior-turns',
name: 'includes prior conversation memory before the new user prompt',
category: 'retrieval',
description:
'Memory holds two prior turns. The provider request must contain both, in order, followed by the new user prompt, with the system prompt present.',
workflowInput: {},
agent: {
model: 'gpt-4o',
systemPrompt: 'You are a helpful assistant.',
userPrompt: 'What is my name?',
memory: {
conversationId: 'conv-ada',
history: [
{ role: 'user', content: 'My name is Ada.' },
{ role: 'assistant', content: 'Nice to meet you, Ada.' },
],
},
},
providerResponse: {
content: 'Your name is Ada.',
tokens: { input: 30, output: 5, total: 35 },
},
expect: {
succeeds: true,
finalContent: 'Ada',
},
},
{
id: 'context-isolates-conversations',
name: 'reads the conversation named by the block, not another',
category: 'retrieval',
description:
'The memory stub only returns history for the block conversation id; any other id yields a placeholder. A handler that passed the wrong id would surface the placeholder and fail.',
workflowInput: {},
agent: {
model: 'gpt-4o',
userPrompt: 'What did we decide?',
memory: {
conversationId: 'conv-b',
history: [
{ role: 'user', content: 'We decided to ship on Friday.' },
{ role: 'assistant', content: 'Shipping Friday.' },
],
},
},
providerResponse: {
content: 'You decided to ship on Friday.',
tokens: { input: 25, output: 6, total: 31 },
},
expect: {
succeeds: true,
finalContent: 'Friday',
},
},
]
84 changes: 84 additions & 0 deletions apps/sim/evals/agent-tool-use/executor-harness.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,9 @@ import {
createSerializedWorkflow,
} from '@sim/testing/factories/serialized-block.factory'
import { providersMockFns } from '@sim/testing/mocks/providers.mock'
import { vi } from 'vitest'
import { DAGExecutor } from '@/executor/execution/executor'
import { memoryService } from '@/executor/handlers/agent/memory'
import type { SerializedBlock, SerializedWorkflow } from '@/serializer/types'
import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness'
import type {
Expand Down Expand Up @@ -58,6 +60,14 @@ export interface ExecutorScenario {
retry?: { enabled: boolean; maxTries: number; waitBetweenTriesMs: number }
/** Ordered models the Agent handler tries after the primary fails. */
fallbackModels?: Array<{ model: string }>
/**
* Turns on conversation memory. `history` is what the mocked memory read
* returns for `conversationId`, so a wrong id surfaces as a failed check.
*/
memory?: {
conversationId: string
history: Array<{ role: 'user' | 'assistant' | 'system'; content: string }>
}
}
/** One entry per model call; the last entry serves any extra/retry calls. */
providerResponse: ExecutorProviderResponse | ExecutorProviderResponse[]
Expand All @@ -73,6 +83,13 @@ export interface ExecutorScenario {
}
}

function lastMatchingIndex(contents: string[], needle: string): number {
for (let index = contents.length - 1; index >= 0; index--) {
if (contents[index].includes(needle)) return index
}
return -1
}

function buildWorkflow(scenario: ExecutorScenario): SerializedWorkflow {
const start: SerializedBlock = createSerializedBlock({
id: 'start',
Expand All @@ -95,6 +112,9 @@ function buildWorkflow(scenario: ExecutorScenario): SerializedWorkflow {
? { temperature: scenario.agent.temperature }
: {}),
...(scenario.agent.fallbackModels ? { fallbackModels: scenario.agent.fallbackModels } : {}),
...(scenario.agent.memory
? { memoryType: 'conversation', conversationId: scenario.agent.memory.conversationId }
: {}),
}
if (scenario.agent.retry) agent.retry = scenario.agent.retry

Expand Down Expand Up @@ -133,6 +153,17 @@ export async function runExecutorScenario(
}
)

const fetchedConversationIds: unknown[] = []
if (scenario.agent.memory) {
const memory = scenario.agent.memory
vi.spyOn(memoryService, 'fetchMemoryMessages').mockImplementation(async (_ctx, inputs) => {
fetchedConversationIds.push(inputs.conversationId)
return inputs.conversationId === memory.conversationId
? memory.history.map((message) => ({ ...message }))
: [{ role: 'user', content: '__WRONG_CONVERSATION__' }]
})
}

const executor = new DAGExecutor({
workflow: buildWorkflow(scenario),
workflowInput: scenario.workflowInput,
Expand Down Expand Up @@ -212,6 +243,59 @@ export async function runExecutorScenario(
})
}

if (scenario.agent.memory) {
const memory = scenario.agent.memory
const requestMessages = ((requests[0] as { messages?: unknown[] } | undefined)?.messages ??
[]) as Array<{ role?: string; content?: unknown }>
const contents = requestMessages.map((message) =>
typeof message.content === 'string' ? message.content : ''
)

const missingHistory = memory.history.filter(
(message) => !contents.some((content) => content.includes(message.content))
)
checks.push({
name: 'memory-history-in-request',
passed: missingHistory.length === 0,
detail:
missingHistory.length === 0
? `all ${memory.history.length} history messages reached the provider`
: `missing [${missingHistory.map((message) => message.content).join(', ')}]`,
})

const lastHistoryIndex =
memory.history.length === 0
? -1
: Math.max(...memory.history.map((message) => lastMatchingIndex(contents, message.content)))
const promptIndex = scenario.agent.userPrompt
? lastMatchingIndex(contents, scenario.agent.userPrompt)
: -1
Comment thread
sudoKrishna marked this conversation as resolved.
checks.push({
name: 'memory-before-user-prompt',
passed: promptIndex >= 0 && promptIndex > lastHistoryIndex,
detail: `history ends at ${lastHistoryIndex}, user prompt at ${promptIndex}`,
})

if (scenario.agent.systemPrompt) {
const systemPrompt = scenario.agent.systemPrompt
checks.push({
name: 'system-prompt-in-request',
passed: requestMessages.some(
(message) => message.role === 'system' && String(message.content).includes(systemPrompt)
),
detail: 'configured system prompt reached the provider',
})
}

checks.push({
name: 'conversation-id',
passed:
fetchedConversationIds.length > 0 &&
fetchedConversationIds.every((id) => id === memory.conversationId),
detail: `expected ${memory.conversationId}, got [${fetchedConversationIds.join(', ')}]`,
})
}

const tokens = (output.tokens ?? {}) as { input?: number; output?: number; total?: number }

return {
Expand Down
1 change: 1 addition & 0 deletions apps/sim/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@
"test:coverage": "vitest run --coverage",
"test:evals": "EVAL_REPORT_PATH=test-results/evals/agent-tool-use.json vitest run evals/agent-tool-use",
"test:evals:live": "EVAL_LIVE=1 vitest run --mode live evals/agent-tool-use",
"test:evals:context": "EVAL_CONTEXT_REPORT_PATH=test-results/evals/agent-context.json vitest run evals/agent-context",
"email:dev": "email dev --dir components/emails",
"type-check": "tsc --noEmit",
"lint": "biome check --write --unsafe .",
Expand Down