Skip to content

Add deepsense.ai cookbook on harness-aware plugin evaluation - #3129

Open
dwigg-openai wants to merge 3 commits into
mainfrom
codex/deepsense-plugin-eval-cookbook
Open

dwigg-openai wants to merge 3 commits into
mainfrom
codex/deepsense-plugin-eval-cookbook

Conversation

@dwigg-openai

Copy link
Copy Markdown
Contributor

Summary

Adds a cookbook developed with deepsense.ai showing how to evaluate an MCP-backed plugin at three levels: direct tool execution, a controlled model loop, and the Codex harness.

Using a small BLS example, the walkthrough shares test cases and grading logic across the three levels, examines differences in recorded traces, and demonstrates how MCP server instructions can improve behaviour.

Includes the notebook, runnable companion code, setup instructions, dependency lockfiles, recorded traces, and screenshots under examples/partners/harness_aware_plugin_evals/, plus registry and author entries.

Motivation

Help developers distinguish tool implementation failures from model and harness behaviour, choose an appropriate evaluation setup during development, and check whether their results reflect the environment where the plugin will run.

Validation

  • All 3 unit tests passed.
  • Notebook executed successfully in offline mode.
  • Notebook structure, author metadata, asset paths, author order, and whitespace checks passed.

@dwigg-openai
dwigg-openai requested a review from a team as a code owner September 23, 2026 14:04
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T14:09:45.782903Z 43e5433 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 43e543375f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread examples/partners/harness_aware_plugin_evals/assert_codex_result.py Outdated

@kkahadze-oai kkahadze-oai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Most of my comments focus on the human reading experience. The extra context is useful for agents, but they can also get context from exploring the repository. I’d prioritize making the article easy to follow, getting readers to setup quickly, and keeping the visible code focused on what teaches the approach or is useful to copy and adapt.

Comment on lines +19 to +32
"A plugin that scores well in your own evaluation can still fail once it runs inside the product it\n",
"ships in. The harness around the model accounts for much of that gap, since it controls the prompt\n",
"the model sees and the tools it can reach. This article runs one set of test cases at three levels\n",
"of fidelity and shows how to trace a difference back to the level that caused it.\n",
"\n",
"Models often need current, specialized data and actions that training alone cannot provide. A\n",
"[plugin](https://developers.openai.com/plugins/concepts/plugins) can package skills or an MCP server\n",
"with callable tools alongside natural-language instructions. This plugin connects a Codex harness,\n",
"now part of the ChatGPT desktop app, to Bureau of Labor Statistics (BLS) data.\n",
"This plugin resolves each request against the BLS catalog and fetches the observations from the\n",
"agency's own API, so every figure comes from an authoritative source.\n",
"The [BLS API](https://www.bls.gov/developers/) requires known series IDs and does not provide\n",
"natural-language series discovery. To bridge that gap, the connector resolves a user's question\n",
"into ranked candidate series before fetching any data.\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we substantially shorten the opening and get to setup sooner? The first paragraph already mostly answers what the reader will learn. The BLS API background and terminology definitions delay that payoff. Readers shouldn’t need to scroll through roughly two pages before getting started. I’d also save the detailed explanation of each evaluation approach for its dedicated section rather than covering it twice.

Comment on lines +51 to +53
"| **One** Direct Tool Execution | Structured inputs sent directly to the tools, ideal for debugging tool logic | Natural-language interpretation and tool selection | Tool code and a structured corpus | Free | Lowest |\n",
"| **Two** Controlled Loop | Natural-language interpretation and tool selection in a controlled function-calling loop over MCP | Product harness behavior | A running MCP server, model API access, and the controlled loop code | Moderate | Moderate |\n",
"| **Three** Product Harness | The same queries run through the [Codex](https://openai.com/codex/) harness with the plugin installed | Client-specific behavior that the evaluation setup might not replicate. Compare representative samples with the target client to identify gaps | A running MCP server, installed plugin, Codex authentication (here with API key), and harness configuration | Highest | Highest, short of the ChatGPT desktop app itself |\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we use descriptive names consistently instead of “Level One,” “Level Two,” and “Level Three”? These labels recur throughout the article, so a first-time reader has to keep remembering the mapping.

"id": "cell-02",
"metadata": {},
"source": [
"### Evaluation difficulty\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I’d cut most of this section and move a short version toward the end. Keep the parts that teach a generalizable lesson, but we don’t need an academic level of detail. This applies throughout the cookbook.

"metadata": {},
"outputs": [],
"source": [
"import atexit\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we move the server startup, socket polling, cleanup, and path configuration into a helper and show one setup call here? This is a substantial block to read before getting to the evaluation itself.

Comment on lines +392 to +396
"carried as promptfoo metadata. The grader ignores them, but they show up\n",
"in the run report, so a run can be sliced by those fields:\n",
"\n",
"```shell\n",
"npx promptfoo eval -c promptfooconfig.yaml --filter-metadata mechanism=coverage_gap\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’re referring to Promptfoo’s configuration and metadata without first introducing Promptfoo. Please briefly explain what it is and what role it plays in this example before readers encounter its configuration or commands.

"id": "cell-31",
"metadata": {},
"source": [
"### Replaying the three runs\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we put the file-loading and reporting loop in a helper and show a replay command plus its output here? The useful lesson is how to run the comparison and interpret the results. Walking through the JSON loading and printing adds length without helping much with that.

"id": "cell-42",
"metadata": {},
"source": [
"## Practical takeaways\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This section repeats the model gaps, harness differences, and variability we’ve just covered. I’d keep the opening bullets about when to use each approach and the final checklist, then substantially condense the material between them. Preserve any unique takeaway without re-explaining the results.

Comment on lines +1254 to +1264
"The most demanding part is adjusting the evaluation strategy itself:\n",
"\n",
"- The sample corpus is built around ambiguity, coverage gaps, and conventional defaults, since those\n",
" are where a plausible BLS answer turns out to be wrong. Another domain has its own equivalents, worth\n",
" identifying before writing the test cases.\n",
"- The shared grader checks tool names, argument shapes, and expected-result fields, so it has to\n",
" change whenever those do.\n",
"- A standalone loop leaves the system prompt open, while a product harness usually supplies its own.\n",
" Guidance then has to sit in the tool descriptions or the MCP server's instructions, since a fix\n",
" placed elsewhere may not reach the model.\n",
"\n"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The bullets after the table revisit adapting the corpus, grader, and instructions, which the table already covers. Could we fold any additional guidance into the relevant rows and remove the repetition?

Comment on lines +1290 to +1297
"## Contributors\n",
"\n",
"This cookbook serves as a joint collaboration effort between OpenAI and its Advanced Partner [deepsense.ai](https://deepsense.ai/), who built and evaluated the BLS connector.\n",
"\n",
"- [Maciej Domagała](https://www.linkedin.com/in/macdomagala/)\n",
"- [Łukasz Dragan](https://www.linkedin.com/in/%C5%82ukasz-d-a22b3986/)\n",
"- [Michał Rdzany](https://www.linkedin.com/in/micha%C5%82-rdzany-85567a225/)\n",
"- [Danny Wigg](https://www.linkedin.com/in/dannywigg/)"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This reads as promotional, particularly “Advanced Partner,” and the authors already appear at the top with their links. I’d remove this section.

Comment on lines +598 to +601
"This example is a trajectory smoke test. It checks the mechanics\n",
"of a run: that the expected tools ran, that the expected series were fetched, and that no series was\n",
"fetched outside what the case expects, which also catches a fabricated series ID fetched alongside a\n",
"real one. For `employment-ambiguous`, it checks that the model called\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This reads a bit jargony and AI-generated, especially “trajectory smoke test.” Could we explain what the example checks in plain language and simplify the sentence that follows?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants