Skip to content

vmem: Add stronger tlb invalidation barrier - #1876

Open
syntactically wants to merge 1 commit into
mainfrom
lm/extra-tlbi
Open

syntactically wants to merge 1 commit into
mainfrom
lm/extra-tlbi

Conversation

@syntactically

Copy link
Copy Markdown
Member

No description provided.

@hyperlight-gh-bot

This comment has been minimized.

@syntactically syntactically added kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. ready-for-review PR is ready for (re-)review labels Oct 2, 2026
Copilot AI balanced review requested due to automatic review settings October 2, 2026 12:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unaligned ranges can leave stale permissions cached, and the new architecture-specific behavior lacks tests.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 2 Low severity

Open (4)
What changed in this PR

Adds a guest paging barrier for invalidating cached translations after permission downgrades.

Changes:

  • Exposes downgrade_in_place.
  • Uses invlpg on amd64.
  • Uses DSB, TLBI, and ISB on AArch64.
File Description
src/​hyperlight_guest_bin/​src/​paging.rs Exposes the barrier API.
src/​hyperlight_guest_bin/​src/​arch/​amd64/​paging.rs Adds amd64 invalidation.
src/​hyperlight_guest_bin/​src/​arch/​aarch64/​paging.rs Adds AArch64 invalidation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/hyperlight_guest_bin/src/paging.rs
Comment thread src/hyperlight_guest_bin/src/paging.rs
Comment thread src/hyperlight_guest_bin/src/arch/aarch64/paging.rs
Comment thread src/hyperlight_guest_bin/src/arch/amd64/paging.rs
@hyperlight-gh-bot

This comment has been minimized.

Base automatically changed from lm/modify-mappings to main October 2, 2026 15:45
Signed-off-by: Lucy Menon <168595099+syntactically@users.noreply.github.com>
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: 3423312ebafc
Baseline commit: 43cf3539bd10

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 766.32 ns (➖ 1.01x faster)
vec_bytes 587.62 ns (➖ 1.02x slower)
375.62 µs (➖ 1.00x slower)

payload_allocation

slot_pool_segmented
262144 543.74 ns (➖ 1.02x slower)
65536 141.37 ns (➖ 1.02x faster)

sandboxes

create_initialized_and_drop
medium 77.52 ms (➖ 1.03x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.85 ns (➖ 1.00x slower) 7.73 ns (➖ 1.01x slower) 7.78 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 93.76 µs (➖ 1.01x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.26 µs (➖ 1.01x slower) 7.31 µs (➖ 1.00x faster)
65536 2.07 µs (➖ 1.00x slower) 2.08 µs (➖ 1.08x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.32 µs (➖ 1.01x faster) 6.25 µs (➖ 1.02x faster)
8192 1.06 µs (➖ 1.00x faster) 1.08 µs (➖ 1.01x slower)
262144 27.43 µs (➖ 1.01x slower)
kvm / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 683.51 ns (➖ 1.09x faster)
vec_bytes 518.06 ns (➖ 1.05x faster)
691.40 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 515.66 ns (➖ 1.11x faster)
65536 139.34 ns (➖ 1.11x faster)

sandboxes

create_initialized_and_drop
medium 76.25 ms (➖ 1.04x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.06 ns (➖ 1.06x faster) 7.05 ns (➖ 1.06x faster) 7.06 ns (➖ 1.06x faster)

snapshot_files

load_snapshot_unverified
small 46.96 µs (➖ 1.05x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.76 µs (➖ 1.10x faster) 7.81 µs (➖ 1.10x faster)
65536 2.18 µs (➖ 1.10x faster) 2.20 µs (➖ 1.08x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.31 µs (➖ 1.08x faster) 7.17 µs (➖ 1.11x faster)
8192 760.88 ns (➖ 1.09x faster) 737.95 ns (➖ 1.08x faster)
262144 30.11 µs (➖ 1.11x faster)
mshv3 / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 958.90 ns (➖ 1.01x slower)
vec_bytes 712.22 ns (➖ 1.00x faster)
336.21 µs (➖ 1.05x faster)

payload_allocation

slot_pool_segmented
262144 727.24 ns (➖ 1.01x faster)
65536 194.19 ns (➖ 1.01x faster)

sandboxes

create_initialized_and_drop
medium 58.16 ms (➖ 1.05x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.77 ns (➖ 1.02x slower) 9.74 ns (➖ 1.00x slower) 9.72 ns (➖ 1.06x faster)

snapshot_files

load_snapshot_unverified
small 83.84 µs (➖ 1.00x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.05 µs (➖ 1.07x faster) 8.97 µs (➖ 1.01x faster)
65536 2.29 µs (➖ 1.01x slower) 2.46 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.24 µs (➖ 1.00x slower) 8.70 µs (➖ 1.08x slower)
8192 1.35 µs (➖ 1.06x slower) 1.34 µs (➖ 1.01x slower)
262144 36.64 µs (➖ 1.00x slower)
mshv3 / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 926.64 ns (➖ 1.00x faster)
vec_bytes 653.34 ns (➖ 1.01x faster)
685.95 µs (➖ 1.02x slower)

payload_allocation

slot_pool_segmented
262144 626.43 ns (➖ 1.03x faster)
65536 165.46 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 66.12 ms (➖ 1.06x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.75 ns (➖ 1.00x faster) 8.75 ns (➖ 1.00x faster) 8.76 ns (➖ 1.00x faster)

snapshot_files

load_snapshot_unverified
small 43.22 µs (➖ 1.02x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.74 µs (➖ 1.00x faster) 7.97 µs (➖ 1.00x faster)
65536 2.21 µs (➖ 1.01x faster) 2.27 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.71 µs (➖ 1.03x slower) 7.55 µs (➖ 1.00x faster)
8192 853.74 ns (➖ 1.00x faster) 880.94 ns (➖ 1.01x faster)
262144 36.91 µs (➖ 1.03x faster)
hyperv-ws2025 / amd (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.21 µs (➖ 1.07x slower)
vec_bytes 820.84 ns (➖ 1.15x faster)
2.19 ms (➖ 1.14x faster)

payload_allocation

slot_pool_segmented
262144 828.97 ns (➖ 1.02x slower)
65536 231.90 ns (➖ 1.02x slower)

sandboxes

create_initialized_and_drop
medium 87.66 ms (➖ 1.26x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.04 ns (➖ 1.38x faster) 10.12 ns (➖ 1.04x faster) 10.12 ns (➖ 1.04x faster)

snapshot_files

load_snapshot_unverified
small 714.64 µs (➖ 1.11x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.13 µs (➖ 1.05x slower) 9.16 µs (➖ 1.02x faster)
65536 2.38 µs (➖ 1.06x faster) 2.32 µs (➖ 1.16x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.82 µs (➖ 1.09x slower) 8.75 µs (➖ 1.00x slower)
8192 1.26 µs (➖ 1.13x faster) 1.28 µs (➖ 1.10x faster)
262144 38.89 µs (➖ 1.09x slower)
hyperv-ws2025 / intel (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.21 µs (➖ 1.00x slower)
vec_bytes 800.22 ns (➖ 1.04x slower)
3.14 ms (➖ 1.01x slower)

payload_allocation

slot_pool_segmented
262144 745.59 ns (➖ 1.01x slower)
65536 213.97 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 108.79 ms (➖ 1.11x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.01 ns (➖ 1.00x slower) 10.20 ns (➖ 1.01x slower) 10.66 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 615.62 µs (➖ 1.01x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.79 µs (➖ 1.02x slower) 7.71 µs (➖ 1.01x faster)
65536 2.31 µs (➖ 1.03x slower) 2.32 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.55 µs (➖ 1.02x slower) 10.00 µs (➖ 1.18x slower)
8192 1.16 µs (➖ 1.08x faster) 1.21 µs (➖ 1.08x slower)
262144 52.15 µs (➖ 1.22x slower)

Reported by cargo ci bench-report --candidate run:37029632479 --baseline run:36945564806 --config-file bench_report.toml.

@github-actions github-actions Bot removed the ready-for-review PR is ready for (re-)review label Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/enhancement For PRs adding features, improving functionality, docs, tests, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants