Skip to content

feat(reskinnable-demo): add the Ledgerline automatic-learning skin - #7570

Draft
davidmckayv wants to merge 23 commits into
mainfrom
david/ledgerline-learning-demo
Draft

davidmckayv wants to merge 23 commits into
mainfrom
david/ledgerline-learning-demo

Conversation

@davidmckayv

@davidmckayv davidmckayv commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Adds ledgerline, a fictitious expense and approvals skin for the reskinnable demo. It is the app half of the Automatic Learning demo.

Story. The in-app agent and ChatGPT (over MCP) both try to approve a team-event report. Both stop at 409 POLICY_HOLD POL-114. The fix only appears in the report page's Policy panel. A person then completes it by hand: allocate the report to the events cost center, approve, reimburse. That work is captured as a product trajectory and linked to the failed Threads. Learning turns it into an insight, a skill and eval candidates. Once a reviewer publishes the skill, the agent loads it and succeeds.

What's in it

  • Pages: overview, expense reports (filters and search), report detail (status timeline, Policy panel, Allocate dialog, approve, reimburse, notes), approvals queue, reimbursements, cost centers, people and policies. All seeded and in memory.
  • Frontend tools: each one draws its own tool line, so a run that does not converge stays visible. Reimbursement is a human-in-the-loop confirm card. A "Using learned skill" card appears when the agent loads a skill.
  • Trajectory recorder: emits AG-UI CUSTOM events (page, navigation, click, network, screen.context, thread.linked, expense.*). It is demo-local, shaped like the draft trajectory contract (feat(learning): opt-in session activity capture as AG-UI events #7556).
  • Agent-trace capture: on the skin's BuiltInAgent, with clone() overridden so the runtime's per-run copies keep capturing.
  • Learning API: /api/learning/v1/* with CORS, for the Intelligence screens. /learn uses an LLM and checks its citations against the real eventIds, with a deterministic fallback.
  • MCP server for ChatGPT: /api/ledgerline/mcp. A presenter reset manages a port-scoped tunnel.
  • Shell: registered in both registries, skins-config, LINTED_SKIN_IDS and the eslint selector table. The roster docs name the new skin, and the drift guard gains eleventh. Adds @modelcontextprotocol/sdk to the demo.

Reskin skill impact: checked. The roster lists and counts in CLAUDE.md, README, .env.example and templates.md now name ledgerline, and CLAUDE.md has a skin entry. Nothing else in .claude/skills/reskin/ changes.

Verification: lint, typecheck and test:unit pass (3,286 tests), and next build passes in a separate copy. New unit tests cover the approve gate, what the agent can read, the recorder store, the learn fallback and the MCP fail-then-succeed path. Driven end to end with Playwright and with a ChatGPT-style MCP client: fail in-app, fail over MCP, manual fix captured, learn, publish, succeed in-app and over MCP.

The /intelligence screens are a separate change.

Ledgerline is a fictitious expense and approvals app built to demo
Automatic Learning. The agent fails on a policy hold whose fix only shows
in the report page's Policy panel. The person then fixes it by hand, and
that work is captured as a product trajectory linked to the failed
in-app and ChatGPT Threads. A learning step turns the trajectory into an
insight, a skill and eval candidates. Once a reviewer publishes the
skill, the agent loads it and succeeds.

Includes the overview, reports, approvals, reimbursements, cost
centers, people and policies pages, a trajectory recorder that emits
AG-UI CUSTOM events, agent-trace capture on the skin's BuiltInAgent, the
/api/learning/v1 contract API, an MCP server for ChatGPT at
/api/ledgerline/mcp, and the presenter reset with tunnel management.

Reskin skill checked: the roster lists and counts now name ledgerline,
the eleventh ordinal joins the drift guard, and the eslint selector
table covers the new skin's files.
@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

The Intelligence product screens for the Automatic Learning demo, served by
the same app at a top-level /intelligence route, outside every skin and not
linked from any skin.

- The Intelligence web app's own UI is copied in (src/intelligence-ui, sources
  and commit in SOURCES.md): @cpki/ui, the app shell styles and frame, and the
  Automatic Learning components (Learning spaces rail, latest analysis,
  Insights, Skills and Skill candidates, analysis results, Insight drawer with
  evidence). An adapter implements their LearningApi over /api/learning/v1.
- The trajectory view is Atai's prototype page (CopilotKitCommonAgentWorkspace
  9c9dc16b) served as-is with only its data layer swapped.
- Trajectories, Eval candidates and Fine-tune are new screens built from the
  copied pieces, with the official provider logos.
- Reads fall back to bundled sample data when the API is unreachable.
- The LOCK_SKIN proxy no longer rewrites /intelligence (segment-boundary
  exclusion, with a test).

Reskin skill: checked, no skill impact. The proxy change only excludes the
new non-skin /intelligence segment; no skin uses that segment.
…ve UI

Each tool keeps its status line and adds a card under it. A report table
for listReports pages and folds columns that do not fit into each row, so
it never scrolls sideways. A report card renders getReport. A refused
approval shows a policy-hold card with the code and nothing about the fix.
The success run confirms through one approve-and-reimburse card that shows
the allocation.

ChatGPT gets the same report card and approve card as MCP Apps from
/api/ledgerline/mcp. scripts/build-ledgerline-mcp-app.mjs bundles the
components from one source, the write goes through an app-only confirm
tool, and an identical widget call within 20 seconds collapses to an
empty view, because ChatGPT sends each widget call twice.
…nce shell

The trajectory detail now uses the same Intelligence shell as every other
/intelligence page. Atai's page is embedded from /intelligence/trajectory-view
with only its own sidebar, breadcrumb bar, background and prototype badge
hidden; its content is unchanged. The list and the Insight drawer's Open
trajectory link land on this page, and Show event scrolls the shell to the
event.
… a hold is open

Over MCP, approveAndReimburse now answers an open policy hold with the
same POLICY_HOLD error as approveReport, flagged isError, and draws the
hold card instead of the approve card. ChatGPT therefore visibly tries and
fails, and its Thread records the errors. The approve card opens only once
the hold is cleared.

The duplicate-call collapse now remembers only calls that drew a card, so
a retry after a refusal, or after the hold is cleared, runs again instead
of being answered as a repeat.
A formatter pass reformatted Atai's prototype page wholesale. Put back the
copy that differs from the 9c9dc16b original only by the data-layer patches
listed in SOURCE.md, so a diff against the original stays readable.
…latform

Intelligence is not an eval platform. Its Eval candidates screen keeps only
the candidates (generated from trajectories, check/X review, export) and
drops the customer's suite. The suite now lives in a separate stand-in for
the customer's own eval tool, "Benchline Evals", at /eval-platform: ten
cases, 60% pass, run history, the Priya case failing.

"Export to your eval platform" imports the accepted candidates into it
(POST /api/learning/v1/evals/import, GET for the list); they appear as
"Imported from CopilotKit Intelligence", not yet run. LangSmith and
Braintrust stay file downloads. /reset clears the imported cases. The
LOCK_SKIN proxy leaves /eval-platform alone.

Also fixes a type error in the Skill link added earlier: carry the extra
supporting-Insight ids through a spread so the object literal type-checks.
…erline history

- GET /api/learning/v1/trajectories/:id/agui returns a trajectory with each
  Thread's AG-UI event stream (@ag-ui/core types, ids, order, timestamps),
  rebuilt from the stored trace. ChatGPT's MCP calls map to TOOL_CALL_*
  events with rawEvent.source "mcp". Product events stay AG-UI CUSTOM.
- The trajectory view shows a "Raw AG-UI event" disclosure on every agent
  trace step and product event, labels ChatGPT as "via MCP, tool calls
  mapped to AG-UI", and Export downloads the AG-UI payload.
- Seeded history so no Intelligence screen starts empty: eight earlier
  trajectories, four Insights, three published Skills and a pending
  candidate, four analyses, eval candidates (two exported to Benchline
  Evals earlier) and a fine-tune dataset with its last export. All dated
  before today, none about POL-114, team events or CC-410, and none reach
  the agent, so today's attempts still fail and today's run is newest.
…view

Each agent trace step and product event can now expand to its raw AG-UI
event JSON. The rest of Atai's page is unchanged (see SOURCE.md).
… and restyle the skin

The POL-114 fix is no longer one button. The Policy panel shows only the
hold and a View policy link, Approve is disabled while the hold is open,
the policy text lives on /policies/POL-114, and Cost centers shows which
center owns the events budget. Edit coding recodes lines one by one and
re-checks the policy; a decoy center or recoding every line leaves the
hold open with the reason inline.

recodeLines and listCostCenters replace allocateCostCenter in the in-app
agent and the MCP server. The recorder captures lines_recoded and
policy_rechecked; learning keys the fix on the recode that resolved the
hold and rejects an LLM skill that does not name every moved line.

The skin is restyled in a crisp light direction: Geist, hairline tables,
Ledger Blue for active and primary, Overview charts with a policy-hold
trend, a command palette, and matching chat and MCP App cards.

Shell: skins can set layoutDefaults (chat side, thread rail side, rail
open, chat width). A skin that sets it gets its own stored layout and a
per-skin rail state. Ledgerline docks chat right with the rail closed.
Other skins are unchanged. The reskin skill needs no change; CLAUDE.md
documents the new optional field.
…d Threads-ready count

- Each published Skill dates from the newest Insight it rests on, so the
  seeded Skills keep their older dates and only today's Skill reads today.
- The Skill delivery switch opts out of browser form-state restore, so a
  reload cannot leave it off while the status says enabled.
- "Threads ready" counts the same pending Threads as the run button, so a
  reset space reads 0 / 1 beside No new Threads.
…read rail tidy

The Inspector launcher sat over the chat header's swap and close controls
now that Ledgerline docks chat on the right. While the skin is open, an
undragged launcher moves into the sidebar above Reset, and the shared
stored position goes back to the default when the skin is left.

Without Intelligence the runtime neither names nor deletes threads, so
rehearsals piled up as "New chat" rows. A new optional threadList skin
field titles unnamed threads from their first user message and hides
threads from before a time the skin supplies; Ledgerline passes its last
full demo reset, served on /ledger/version.

Cost centers shows the quarter budget as a compact meter and one line,
with the exact figures and pending share in the tooltip.
The drawer opened on top of the chat, covering the pills and composer.
A new layoutDefaults.inboxPlacement "column" renders the thread rail as
its own column on the chat's outer edge. Opening it animates its width
over 200ms; the chat panel holds its pixel width (preserve-pixel-size),
so the app panel gives up the space and reflows narrower. Closing
reverses it. The rail stays mounted but inert while closed so it slides
with real content.

Ledgerline opts in and drops its overlay-drawer CSS. Its app header now
adapts to the app card's width rather than the viewport, so the narrower
card keeps the breadcrumb and search on one line, and chart headers keep
a gap between title and caption. Other skins keep the rail inside the
chat card.
…as a key moment

- Seeded history no longer touches receipts, card transactions, splits or
  FX (today's month-end reconciliation): its failures are now an approval
  delegate, a card limit, archiving stale drafts and a per diem override,
  with Insights, Skills, eval candidates and fine-tune examples to match.
- The trajectory view lists the person's network requests, in order with
  repeats folded, as a key moment: the API recipe the agent lacked; network
  events carry their summary and any request or response body.
…d close

The agent's task is now matching card charges to receipts. The new Card
close board puts each card's unmatched charges beside its receipts inbox:
drag a receipt onto a charge (or use Match), the board adds the tip
written on a slip or the currency conversion as the pair's adjustment,
Validate matches, then Close the month. A wrong receipt fails with the
reason inline. Under the hood it is a reconciliation session: one pair
per charge, validate, close. The board's recorded calls are the recipe.

The agent gets one generic ledgerlineApi tool with a terse endpoint
index, and receipts as merchant, date and total only, so it fails
visibly in-app and over MCP. Learning turns the trajectory into the
match-card-receipts skill: the exact session recipe plus matching rules
for descriptors, posting lag, tips, currency and split receipts. With
it the agent prepares a different card's matches and hands over a
Review matches card; it never closes a period, and only the person's
Confirm validates and closes. Same over MCP with the app card.

Approve and reimburse leave the demo path and pills; the expense
pages stay.
…e story

- Sample data is now a snapshot of a real card close run: Priya's
  September card, the failed in-app attempt, the board close, and the
  match-card-receipts skill with its eval candidates and fine-tune preview.
- Trajectory key moments: the agent's failure, what it was missing (the
  Card close workflow and the receipt details), the API recipe read from
  the network requests in order, and the person's steps with matches
  folded into a count.
- Seeded history drops its approve and reimburse items for read-only ones.
… card close

The customer's eval suite still described the old policy-hold story. Its
ten cases now cover the month-end card close and everyday expense
questions, still six passing at a 0.6 pass rate, with matching the
September card transactions failing on the tip and currency pairs.
…card close

The agent announced the learned skill and ended its turn after one call, and
the runtime's 10-step cap with a skill loaded is below a full close (~12 calls).
…lose about its exceptions

Receipts now auto-match when a close opens (tips, euros, two-receipt charges),
the way spend platforms do. What is left on each card are four exceptions a
person clears in their own workflow on the Card close board: split an offsite
by attendees (allocation draft, lines, commit), reclass a miscoded software
charge in a soft-locked month, mark a personal charge for payroll repayment,
and request a missing-receipt affidavit. Editing a charge is refused with the
rule for its exception, never the workflow; the review card lists each cleared
exception.
… close

The agent prompt, pills and learning step follow the exceptions story. The
skill (renamed close-card-exceptions) is the recipe the board recorded:
session, allocation split by attendees, reclass entry, repayment, affidavit,
validate, then hand over for review. Its rules come from what the person chose
on each exception. Insight, eval candidates, the fine-tune examples and the
seeded Benchline suite tell the same story.
…mple data

Title a trajectory by the card the agent worked on, not the one the screen
showed, so Marcus's run no longer reads as Priya's. Keep the card tabs' counts
current when a close lands from the chat's review card, record one labelled
click per saved split, and regenerate the Intelligence screens' offline sample
data from a live run of the new close.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant