Add deepsense.ai cookbook on harness-aware plugin evaluation - #3129
dwigg-openai wants to merge 3 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 43e543375f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
kkahadze-oai
left a comment
There was a problem hiding this comment.
Most of my comments focus on the human reading experience. The extra context is useful for agents, but they can also get context from exploring the repository. I’d prioritize making the article easy to follow, getting readers to setup quickly, and keeping the visible code focused on what teaches the approach or is useful to copy and adapt.
| "A plugin that scores well in your own evaluation can still fail once it runs inside the product it\n", | ||
| "ships in. The harness around the model accounts for much of that gap, since it controls the prompt\n", | ||
| "the model sees and the tools it can reach. This article runs one set of test cases at three levels\n", | ||
| "of fidelity and shows how to trace a difference back to the level that caused it.\n", | ||
| "\n", | ||
| "Models often need current, specialized data and actions that training alone cannot provide. A\n", | ||
| "[plugin](https://developers.openai.com/plugins/concepts/plugins) can package skills or an MCP server\n", | ||
| "with callable tools alongside natural-language instructions. This plugin connects a Codex harness,\n", | ||
| "now part of the ChatGPT desktop app, to Bureau of Labor Statistics (BLS) data.\n", | ||
| "This plugin resolves each request against the BLS catalog and fetches the observations from the\n", | ||
| "agency's own API, so every figure comes from an authoritative source.\n", | ||
| "The [BLS API](https://www.bls.gov/developers/) requires known series IDs and does not provide\n", | ||
| "natural-language series discovery. To bridge that gap, the connector resolves a user's question\n", | ||
| "into ranked candidate series before fetching any data.\n", |
There was a problem hiding this comment.
Can we substantially shorten the opening and get to setup sooner? The first paragraph already mostly answers what the reader will learn. The BLS API background and terminology definitions delay that payoff. Readers shouldn’t need to scroll through roughly two pages before getting started. I’d also save the detailed explanation of each evaluation approach for its dedicated section rather than covering it twice.
| "| **One** Direct Tool Execution | Structured inputs sent directly to the tools, ideal for debugging tool logic | Natural-language interpretation and tool selection | Tool code and a structured corpus | Free | Lowest |\n", | ||
| "| **Two** Controlled Loop | Natural-language interpretation and tool selection in a controlled function-calling loop over MCP | Product harness behavior | A running MCP server, model API access, and the controlled loop code | Moderate | Moderate |\n", | ||
| "| **Three** Product Harness | The same queries run through the [Codex](https://openai.com/codex/) harness with the plugin installed | Client-specific behavior that the evaluation setup might not replicate. Compare representative samples with the target client to identify gaps | A running MCP server, installed plugin, Codex authentication (here with API key), and harness configuration | Highest | Highest, short of the ChatGPT desktop app itself |\n", |
There was a problem hiding this comment.
Could we use descriptive names consistently instead of “Level One,” “Level Two,” and “Level Three”? These labels recur throughout the article, so a first-time reader has to keep remembering the mapping.
| "id": "cell-02", | ||
| "metadata": {}, | ||
| "source": [ | ||
| "### Evaluation difficulty\n", |
There was a problem hiding this comment.
I’d cut most of this section and move a short version toward the end. Keep the parts that teach a generalizable lesson, but we don’t need an academic level of detail. This applies throughout the cookbook.
| "metadata": {}, | ||
| "outputs": [], | ||
| "source": [ | ||
| "import atexit\n", |
There was a problem hiding this comment.
Could we move the server startup, socket polling, cleanup, and path configuration into a helper and show one setup call here? This is a substantial block to read before getting to the evaluation itself.
| "carried as promptfoo metadata. The grader ignores them, but they show up\n", | ||
| "in the run report, so a run can be sliced by those fields:\n", | ||
| "\n", | ||
| "```shell\n", | ||
| "npx promptfoo eval -c promptfooconfig.yaml --filter-metadata mechanism=coverage_gap\n", |
There was a problem hiding this comment.
We’re referring to Promptfoo’s configuration and metadata without first introducing Promptfoo. Please briefly explain what it is and what role it plays in this example before readers encounter its configuration or commands.
| "id": "cell-31", | ||
| "metadata": {}, | ||
| "source": [ | ||
| "### Replaying the three runs\n", |
There was a problem hiding this comment.
Could we put the file-loading and reporting loop in a helper and show a replay command plus its output here? The useful lesson is how to run the comparison and interpret the results. Walking through the JSON loading and printing adds length without helping much with that.
| "id": "cell-42", | ||
| "metadata": {}, | ||
| "source": [ | ||
| "## Practical takeaways\n", |
There was a problem hiding this comment.
This section repeats the model gaps, harness differences, and variability we’ve just covered. I’d keep the opening bullets about when to use each approach and the final checklist, then substantially condense the material between them. Preserve any unique takeaway without re-explaining the results.
| "The most demanding part is adjusting the evaluation strategy itself:\n", | ||
| "\n", | ||
| "- The sample corpus is built around ambiguity, coverage gaps, and conventional defaults, since those\n", | ||
| " are where a plausible BLS answer turns out to be wrong. Another domain has its own equivalents, worth\n", | ||
| " identifying before writing the test cases.\n", | ||
| "- The shared grader checks tool names, argument shapes, and expected-result fields, so it has to\n", | ||
| " change whenever those do.\n", | ||
| "- A standalone loop leaves the system prompt open, while a product harness usually supplies its own.\n", | ||
| " Guidance then has to sit in the tool descriptions or the MCP server's instructions, since a fix\n", | ||
| " placed elsewhere may not reach the model.\n", | ||
| "\n" |
There was a problem hiding this comment.
The bullets after the table revisit adapting the corpus, grader, and instructions, which the table already covers. Could we fold any additional guidance into the relevant rows and remove the repetition?
| "## Contributors\n", | ||
| "\n", | ||
| "This cookbook serves as a joint collaboration effort between OpenAI and its Advanced Partner [deepsense.ai](https://deepsense.ai/), who built and evaluated the BLS connector.\n", | ||
| "\n", | ||
| "- [Maciej Domagała](https://www.linkedin.com/in/macdomagala/)\n", | ||
| "- [Łukasz Dragan](https://www.linkedin.com/in/%C5%82ukasz-d-a22b3986/)\n", | ||
| "- [Michał Rdzany](https://www.linkedin.com/in/micha%C5%82-rdzany-85567a225/)\n", | ||
| "- [Danny Wigg](https://www.linkedin.com/in/dannywigg/)" |
There was a problem hiding this comment.
This reads as promotional, particularly “Advanced Partner,” and the authors already appear at the top with their links. I’d remove this section.
| "This example is a trajectory smoke test. It checks the mechanics\n", | ||
| "of a run: that the expected tools ran, that the expected series were fetched, and that no series was\n", | ||
| "fetched outside what the case expects, which also catches a fabricated series ID fetched alongside a\n", | ||
| "real one. For `employment-ambiguous`, it checks that the model called\n", |
There was a problem hiding this comment.
This reads a bit jargony and AI-generated, especially “trajectory smoke test.” Could we explain what the example checks in plain language and simplify the sentence that follows?
Summary
Adds a cookbook developed with deepsense.ai showing how to evaluate an MCP-backed plugin at three levels: direct tool execution, a controlled model loop, and the Codex harness.
Using a small BLS example, the walkthrough shares test cases and grading logic across the three levels, examines differences in recorded traces, and demonstrates how MCP server instructions can improve behaviour.
Includes the notebook, runnable companion code, setup instructions, dependency lockfiles, recorded traces, and screenshots under
examples/partners/harness_aware_plugin_evals/, plus registry and author entries.Motivation
Help developers distinguish tool implementation failures from model and harness behaviour, choose an appropriate evaluation setup during development, and check whether their results reflect the environment where the plugin will run.
Validation