Skip to content

feat(chat): render HTML, SVG, image and PDF artifacts inline in the conversation - #17822

Draft
analisaperlengkapan wants to merge 2 commits into
OpenHands:mainfrom
analisaperlengkapan:feat/inline-html-svg-artifact-preview
Draft

analisaperlengkapan wants to merge 2 commits into
OpenHands:mainfrom
analisaperlengkapan:feat/inline-html-svg-artifact-preview

Conversation

@analisaperlengkapan

@analisaperlengkapan analisaperlengkapan commented Sep 30, 2026 •

Copy link
Copy Markdown

HUMAN:

HUMAN:
I've manually tested the preview cards across all supported formats in Chrome. Verified that report.html renders with linked styles while blocking script execution, chart.svg renders as a clean vector document, and PDFs load in the browser viewer. Also I've confirmed lazy mounting via IntersectionObserver and verified that all existing/new unit tests pass.


AGENT:

A file-editor create event whose path is .html, .htm, .svg, a raster
image or a PDF now renders an inline card with a height-clipped live preview,
plus Expand / View / Copy / Download — mirroring the Markdown artifact card from
#16185. Markdown behavior is unchanged; the Markdown card was refactored to share
the same plumbing rather than duplicated.

Stacked PR — part 1 of 2. Office documents (.docx / .xlsx / .pptx)
are the second half and land in a follow-up PR stacked on this one, because
they share no code with the browser-renderable formats: this PR is the whole
"render what the browser can already render" path, the follow-up adds the
client-side OOXML unpacker. Splitting keeps each diff to one reviewable
subsystem. Review and merge this one first.

📄 Design doc (before → after, with the flow diagram and the API delta):
https://htmlpreview.github.io/?https://raw.githubusercontent.com/analisaperlengkapan/OpenHands/pr-assets/pr-assets/inline/17818-inline-artifacts.html

Evidence — real browser, mock API, no LLM

The demo conversation fixture now creates canvas.md, report.html (linking a
sibling ./report.css), chart.svg, a PNG and a PDF, so the flow is reviewable
without a live stack. I ran the app (npm run dev:mock -- --port 3001) and drove
it in real Chromium with Playwright, then probed the rendered DOM and the frames:

Observation Value Proves
cards / frames rendered 5 / 2 one card per artifact, one frame per HTML/SVG
sandbox on HTML/SVG frames allow-same-origin no allow-scripts
HTML frame <h1> text Inline HTML artifact agent <script> did not run
HTML <h1> computed colour rgb(31, 111, 235) ./report.css resolved from the workspace URL
SVG frame root tag svg SVG renders as a document, not as text
PNG element <img alt="preview.png">, naturalWidth=96 the image really decoded
PDF frame sandbox (absent) Chromium's PDF plugin is allowed to instantiate
console during load Blocked script execution … 'allow-scripts' permission is not set the browser enforces the sandbox
HTTP status for frame requests 200 no error path exercised

Note on method: the mock API is an MSW service worker, which intercepts fetch
but not an iframe's document navigation, so a preview frame would 502 in mock
mode. The capture routes the frame URLs to the exact bytes the app's own mock
serves (fetched in-page, where MSW does intercept). Same bytes, real component,
real sandbox attribute.

Evidence — every preview is legible to the model

The bar for this change is that whatever the card renders is content a model can
consume
, not a decorative box. That is directly relevant to a natively
multimodal model such as DeepSeek V4.1 Flash (released 2026-09-10, API name
deepseek-flash: text and image input, 1M-token context, MIT). The capture
therefore records, per card, the element rendered, its accessible name, and the
text the card exposes:

Artifact Renders as Accessible name Content a model can read
report.html sandboxed iframe title="report.html" rendered document (pixels + DOM)
chart.svg sandboxed iframe title="chart.svg" rendered document (pixels + DOM)
preview.png <img> alt="preview.png" the image itself (pixels)
spec.pdf iframe (no sandbox) title="docs/spec.pdf" the PDF viewer (pixels)

A regression test asserts exactly that (artifact-preview.test.tsx → "labels its
content so the preview is machine-readable, not just pixels").

Reproduce:

npm ci
npm run dev:mock -- --port 3001
# the demo conversation creates canvas.md, report.html (+ report.css), chart.svg,
# assets/preview.png, docs/spec.pdf
OUT_DIR=/tmp node /tmp/capture-17818.mjs   # prints the first table above

Screenshots (clipped card, expanded frame, side-by-side):

Checks

  • npx vitest run → 766 files, 8040 passed | 7 todo.
  • npx tsc --noEmit → clean.
  • npx eslint on the touched files → 0 errors. The shadcn/no-arbitrary-values
    warnings are the same off-token utilities the existing
    markdown-file-preview.tsx already emits (text-[11px], tracking-[0.11px]),
    reused deliberately so the cards look identical.

Why

Only Markdown gets a rich inline preview today (src/utils/is-markdown-file-path.ts
accepts md/markdown/mdx; markdown-file-preview.tsx renders it). Every other
artifact — including the HTML and SVG the agent just wrote — falls through to a
<CodeBlock> of raw source, so the user has to leave the chat and find the file in
the Files drawer. The Files drawer can render HTML (file-content-viewer.tsx), but
that is not the conversation stream. Reported in #17818; the original request (#2691)
was auto-closed by the stale bot with no successor.

Summary

  • Add ArtifactPreview, a live-preview primitive that renders by kind:
    <iframe sandbox="allow-same-origin"> pointed at the workspace static URL
    for HTML/SVG, so relative ./style.css / images resolve while the missing
    allow-scripts keeps agent-written <script> and inline handlers inert; an
    <img alt=fileName> for raster images; and an unsandboxed iframe for PDFs,
    because a sandboxed frame is not allowed to instantiate a plugin and Chromium's
    viewer would never appear. Height-clipped (160 px) with Expand (512 px), View
    (Files drawer), Copy source and Download. The frame mounts lazily on
    IntersectionObserver, so a long conversation does not keep every preview alive.
  • Replace the per-format predicates with one classifier,
    getArtifactPreviewKind → markdown | frame | image | pdf | null, and route by
    it in file-editor.tsx. isPreviewableArtifactPath now covers every
    rich-preview format, so these creates are group breakers that start expanded,
    exactly like Markdown creates.
  • Add fixtures (artifact-formats-demo.ts) carrying a real PNG and PDF, served by
    the mock fileserver, so the demo conversation exercises every format without an
    LLM or a live workspace.

Issue Number

Fixes #17818

How to Test

  1. npm ci
  2. npx vitest run __tests__/components/conversation-events/ __tests__/utils/is-previewable-file-path.test.ts __tests__/components/features/chat/tool-visualizers/ __tests__/api/mock-conversation-handlers.test.ts — these assert the sandbox tokens, the type allowlist and the lazy-mount behavior.
  3. npm run dev:mock and open http://127.0.0.1:3001/conversations/canvas-demo: report.html renders as a styled card (blue heading from the sibling CSS), chart.svg renders as a chart, assets/preview.png renders as an image, and docs/spec.pdf opens in the browser's PDF viewer — all without leaving the chat. Expand grows each preview; the HTML heading never changes to "SCRIPT RAN", i.e. the fixture's own <script> stayed inert.

Video/Screenshots

Captured with the tracked harness against the running dev:mock app in real
Chromium; the tables above are its output. The images are hosted on a throwaway
pr-assets branch of the fork purely so GitHub renders them; they are not part
of the merged tree.

Collapsed card (height-clipped, as it appears in the stream):

collapsed HTML card

Expanded card (HTML renders, sibling ./report.css applied, <script> inert):

expanded HTML card

HTML and SVG each render as a document (SVG shown):

SVG card

Type

  • Bug fix
  • Feature
  • Refactor
  • Breaking change
  • Docs / chore

Notes

  • Stacked on nothing; the follow-up stacks on this branch. The Office PR
    targets feat/inline-html-svg-artifact-preview, so merge this first.
  • No new dependencies.
  • The classifier is the single owner of "which formats get a rich preview"; adding
    a format means one entry there plus one render branch, not a new predicate.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

PR Artifacts Cleaned Up

The .pr/ directory is no longer present.

Render created HTML, SVG, PNG and PDF artifacts inline in the conversation
instead of only Markdown. HTML/SVG mount in a sandboxed iframe pointed at the
workspace fileserver so relative assets resolve and agent scripts stay inert;
images paint directly and PDFs use Chromium's viewer.

Co-authored-by: openhands <openhands@all-hands.dev>
analisaperlengkapan pushed a commit to analisaperlengkapan/OpenHands that referenced this pull request Sep 30, 2026
@analisaperlengkapan
analisaperlengkapan force-pushed the feat/inline-html-svg-artifact-preview branch 2 times, most recently from 4c5d7df to 88cadd0 Compare September 30, 2026 14:40
@analisaperlengkapan analisaperlengkapan changed the title feat(chat): render HTML/SVG artifacts inline in the conversation feat(chat): render HTML, SVG, image and PDF artifacts inline in the conversation Sep 30, 2026
analisaperlengkapan pushed a commit to analisaperlengkapan/OpenHands that referenced this pull request Oct 1, 2026
The doc declares head 88cadd0 but two anchors were captured at the earlier
120b8b4, which is ~38 lines shorter, so they pointed at the download fallback
and the extension-comment line instead of the frame gate and the allowlist.

- artifact-preview.tsx#L126-L131 -> #L165-L184 (iframe, at 88cadd0)
- is-previewable-file-path.ts#L9 -> #L8 (FRAME_PREVIEW_EXTS, at 88cadd0)
- excerpt label 118-135 -> 156-184 @ 88cadd0

This is the live-workspace evidence finding on PR OpenHands#17822: the excerpt did not
resolve to the code it claimed. Doc only; no product code.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type: feat A new feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Render HTML/SVG/interactive artifacts inline in the conversation

2 participants