Convert MHTML files — or any HTML document with its resources — to a single standalone HTML string: CSS inlined, images and fonts embedded as data URIs, active content sanitized for safe display.
pip install unmhtmlConvert an MHTML archive:
from pathlib import Path
from unmhtml import load_mhtml, to_standalone_html
document = load_mhtml(Path("saved_page.mhtml").read_bytes())
html = to_standalone_html(document)Or bring your own HTML and resources — a rendered page plus the bodies captured in a HAR archive, for example:
from unmhtml import Document, Resource, to_standalone_html
document = Document(
html=rendered_html,
resources={
"https://example.com/style.css": Resource(css_bytes, "text/css"),
"https://example.com/logo.png": Resource(png_bytes, "image/png"),
},
base_url="https://example.com/page.html",
)
html = to_standalone_html(document)References resolve against base_url and match resource keys exactly; CSS
references inside a stylesheet resolve against the stylesheet's own URL.
Anything unresolvable is neutralized, so the result never makes network
requests.
Conversion is secure by default: scripts, event handlers and dangerous URLs
are removed, forms and meta redirects are defused, and CSS that could still
fetch or execute is stripped. Each neutralization has its own flag on
Security (remove_javascript, disable_forms, remove_meta_redirects,
sanitize_css), all defaulting to on:
from unmhtml import Security, to_standalone_html
html = to_standalone_html(
document,
security=Security(remove_javascript=False), # flags are independent
)The contract is reduction, not isolation: the output makes no network
requests, but this is defense in depth, not a sandbox — display untrusted
output in a sandboxed context of your own, such as a sandboxed iframe
without allow-scripts, combined with a Content-Security-Policy.
load_mhtml(data) -> Document— parse MHTML content; raisesValueErrorwhen no HTML part existsDocument(html, resources, base_url)— a parsed document; resources map absolute URLs toResourcevaluesResource(data, mime_type)— one resource; the MIME type is guessed from the reference URL when omittedto_standalone_html(document, *, security=Security()) -> str— the document as one self-contained HTML string
The output is an HTML fragment: html/head/body wrappers are removed
and the original doctype is re-emitted. HTML comments are removed, and
every anchor carries rel="noopener noreferrer".
Python 3.8+. The single dependency is nh3 (Python bindings to the Rust ammonia sanitizer) — zero transitive runtime dependencies, wheels for all major platforms.
MIT