Skip to content

arm64: Port multikernel to arm64 - #40

Open
congwang-mk wants to merge 24 commits into
masterfrom
multikernel-arm
Open

congwang-mk wants to merge 24 commits into
masterfrom
multikernel-arm

Conversation

@congwang-mk

@congwang-mk congwang-mk commented Sep 25, 2026 •

Copy link
Copy Markdown

This series ports multikernel to arm64, following the plan in the tracking issue #21.

PSCI does the CPU lifecycle on arm64: a pool CPU is simply OFF, a spawn starts with CPU_ON, and a spawn halts by turning every CPU off. There is no arm64 counterpart of the x86 trampolines or park loop. Most of the work is in the interrupt controller. A spawn is a tenant of the host's GICv3 and never initialises the distributor. The host proxies ITS commands, so a spawn's MSIs work through an ITS that the host owns.

What is in it

Foundation and CPU lifecycle

Boot data and messaging

Interrupts

Robustness and isolation

Documentation

Testing

  • QEMU virt (GICv3 with ITS): the mk-smoke scenarios for lifecycle, hotplug, PCI, NVMe, offline and wedge pass with DT boot, with TF-A (the host entered at EL2), and with UEFI/ACPI. PCI also passes behind SMMUv3 and behind a PCIe root port.
  • kerf on arm64 (kerf arm64 branch): init, create, load, exec, console, CPU and memory update, graceful and forced kill.
  • Bare metal: an arm64 server boots a spawn, and its console works through kerf console.

Not tested yet: more than one ITS on real hardware (QEMU has a single ITS), crash healing on arm64, and PCIe switches with devices of several owners below one port.

Left open

Closes #9
Closes #10
Closes #11
Closes #12
Closes #13
Closes #14
Closes #15
Closes #16
Closes #18
Closes #20

🤖 Generated with Claude Code

Multikernel only builds on x86 so far. Give arm64 the pieces the
generic core needs from an architecture, so that the port can grow
one function at a time in a tree that links.

arm64 selects ARCH_SUPPORTS_MULTIKERNEL when it has PSCI and
kexec_file: pool CPUs will be held by firmware in the PSCI OFF state
and started with CPU_ON, and instances will be loaded as Image files.
It does not select ARCH_HAS_MK_HOST_PARK, as there is no software park
area to keep in pool memory.

asm/multikernel.h names CPUs by MPIDR through the logical map and
sizes the control block for the boot device tree alone: without a
trampoline, a park page or identity page tables there is nothing else
to carve out of instance memory.

Every function of the architecture interface gets a stub. Spawning
fails with -EOPNOTSUPP, so CONFIG_MULTIKERNEL=y is harmless on arm64
until the follow-up changes fill these in.

Signed-off-by: Cong Wang <cwang@multikernel.io>
CPU_ON hands its third argument, the context ID, to the started CPU
in x0. Secondary boot has no use for it and the driver hardcodes 0.

Multikernel starts a CPU on a different kernel image, straight at
that Image's entry point, and the arm64 boot protocol wants the
device tree address in x0 there. CPU_ON can deliver it without any
trampoline if the caller gets to choose the context ID.

Add psci_cpu_on_context() for that, using whichever CPU_ON function
ID the firmware was probed with. psci_ops.cpu_on() is unchanged.

Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 has no firmware that controls CPUs after boot, so its multikernel
code is mostly a software CPU lifecycle: trampolines, identity page
tables, a park loop in host-owned memory and mailboxes to move parked
CPUs between them. On arm64 firmware does all of that.

A pool CPU is a CPU in the PSCI OFF state. The host gets it there with
the stock hotplug path, cpu_psci_cpu_die() and cpu_psci_cpu_kill(),
which depart_cpu() already ends in. Spawning is CPU_ON at the Image's
entry point with the device tree address as context ID, which is the
state head.S asks for, after cleaning what the host wrote to the PoC
because the new kernel reads it with the caches off.

An off CPU belongs to whoever turns it on next, so there is no slot to
repark between. What remains of that bookkeeping is the question the
core asks before it rewrites an image or frees instance memory: could
a CPU still be executing in there? The instance's kernel turns on its
secondaries without telling the host, so every CPU of a spawned
instance counts as started, a hot-added one joins them, and a CPU
leaves the set once AFFINITY_INFO reports it off. That reuses
cpus_on_slot, whose comment claimed it stays empty on such
architectures; say what it means there instead.

Waiting for OFF matters on hot-remove too: the instance acknowledges
the offline before the CPU has reached firmware, and CPU_ON of a CPU
that is not off yet fails with ALREADY_ON.

Sending the doorbell IPI, forced stop and the halt path of a spawn are
still stubs.

Signed-off-by: Cong Wang <cwang@multikernel.io>
Stock arm64 halts by sending the other CPUs to local_cpu_stop(), which
leaves them in cpu_park_loop(), and spinning on the last one. All of
them keep executing the kernel image, in a WFI loop with DAIF masked.
That is fine when the machine resets next. For a spawn kernel it is
the hazard the x86 park page exists to avoid: the host rewrites that
image for the next spawn and the CPUs run whatever lands under them.
PSCI cannot take a CPU away from outside either, so such a CPU is
lost to the host for good.

A spawn kernel returns every CPU to firmware instead:

 - machine_halt(), machine_power_off() and machine_restart() notify
   the host and turn all CPUs off through mk_halt_to_pool(). A spawn
   must not reach the PSCI SYSTEM_OFF or SYSTEM_RESET behind them,
   which would take the whole machine down, host included.
 - mk_enter_pool_state() is cpu_die() of the calling CPU, the boot CPU
   included: the host's CPUs are on, so firmware never sees the last
   CPU go.
 - local_cpu_stop(), where IPI_CPU_STOP and a parallel panic() put a
   CPU, turns it off instead of parking it.
 - panic_timeout defaults to reboot, which is the halt above, as on
   x86.

The crash stop path already ends in cpu_die(). Secondaries need
nothing: a spawn brings them up with the stock cpu_psci_cpu_boot()
and hot-removes them with cpu_psci_cpu_die(), so there is no
counterpart to x86's wakeup_secondary_cpu_64 hook, and the generic
code never assumed one.

The host learns that a halted instance's CPUs have arrived from
AFFINITY_INFO, in mk_arch_confirm_parked().

mk_spawned() tells a kernel whether it booted with a manifest. The
CPU_OFF path needs cpu_operations::cpu_die, so arm64 multikernel now
depends on HOTPLUG_CPU, which the generic core needs to move CPUs into
the pool anyway.

Signed-off-by: Cong Wang <cwang@multikernel.io>
kexec_file on arm64 already does what loading an instance takes: it
relocates an Image to any 2 MB aligned address, and the generic code
picks that address inside the instance's memory grant for a
KEXEC_TYPE_MULTIKERNEL image. image->start is the Image's entry point
and image->arch.dtb_mem the device tree, the two things CPU_ON needs.
There never was a purgatory to skip.

What has to go is the preparation for replacing the running kernel,
which an instance never does:

 - machine_kexec_post_load() builds the relocation code, a copy of the
   linear map and the EL2 vectors. A multikernel image is loaded in
   place and started on another CPU, so it needs none of them. It is
   not flushed here either: every spawn copies the segments in again,
   and mk_arch_spawn_instance() cleans them to the PoC afterwards, the
   one place that covers the first exec and every later one.
 - machine_kexec_prepare() refuses to load while CPUs are stuck in the
   kernel, which only matters to a kernel that is about to be left.

The device tree the loader builds is a copy of this kernel's own with
a new /chosen. It names every CPU, all of memory and every device, so
an instance must not boot from it, and what an instance owns can
change between load and exec anyway. Its boot tree is therefore
rebuilt in place on every spawn. Size the segment for that, and refuse
to spawn until the code that builds the tree exists, instead of
handing the machine to the new kernel.

With the tree a kimage segment, nothing of arm64's lives in the
instance control block, so drop the fields reserved for it.

Signed-off-by: Cong Wang <cwang@multikernel.io>
On x86 a spawn gets its memory map in boot_params and everything else
in the manifest, a device tree on a channel of its own. On arm64 the
tree in x0 is the only boot channel, and the boot path reads memory,
CPUs, PSCI, the timer and the interrupt controller from it long before
any multikernel code runs. So the two become one tree: the instance
tree the generic code writes, with /resources and the multikernel
/chosen, is the base, and arm64 adds what its boot path needs:

 - a /memory node per region of the grant, the role mk_e820_fill() has
   on x86;
 - /cpus, the instance's own CPUs first so that the boot CPU leads and
   they get the low logical numbers, then every other CPU of the
   machine. A kernel can only online CPUs it enumerated at boot, and
   the generic code already prunes the ones the instance does not own
   from the present mask before smp_init(), so mk_arch_register_cpu()
   stays empty;
 - bootargs, the initrd and the seeds, taken from the /chosen of the
   tree the kexec_file loader made;
 - PSCI, the arch timer and the GICv3.

The last three come from whichever description the host booted with.
A DT host copies its own nodes, translating the GIC frames to CPU
addresses as the node moves to the root and leaving its ITS child
behind. An ACPI host, the SBSA server case, writes the generic
bindings from its tables: the PSCI conduit from the FADT, the timer
PPIs from the GTDT, the distributor and redistributor frames from the
MADT. Both produce the same nodes on QEMU's virt machine. Either way
the spawn is a DT platform: a populated tree disables ACPI on arm64
without any quirk, and no UEFI is advertised.

There is no UART node on purpose. earlycon takes an address on the
command line, and a probed pl011 would claim an SPI the host owns.
ECAM windows for granted PCI segments are left to the device work;
without them a spawn sees no PCI bus at all.

The tree is written into the kimage's device tree segment on every
spawn, since the grant can change between load and exec, and that
segment grows to hold the manifest with the boot nodes on top. On the
other side, setup_arch() accepts the boot tree as the manifest when
its root says so, which is what makes mk_spawned() true.

An instance now boots on QEMU virt from a DT and from a UEFI/ACPI
host: both of its CPUs come up, it halts by turning them off, and a
second exec finds them off and boots again. Until the GIC driver
learns not to reset the shared distributor, that boot takes the host's
SPIs away.

Signed-off-by: Cong Wang <cwang@multikernel.io>
gic_ipi_send_mask() takes a mask of logical CPUs, which can only name
CPUs of the running kernel. The hardware has no such limit: ICC_SGI1R
addresses a redistributor by affinity, and an SGI is per-target state
there, so it reaches a CPU that a different kernel image runs on with
nothing shared between the two.

Multikernel needs exactly that for its doorbell. Add
gic_v3_send_sgi_to_mpidr(), which refuses a target whose Aff0 needs
the range selector when the GIC has none.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The message ring between kernels needs a doorbell, the role
MULTIKERNEL_VECTOR has on x86. On arm64 that is an SGI sent to the
target's MPIDR with gic_v3_send_sgi_to_mpidr().

Its number is ABI between kernels, since the sender picks the INTID
the receiver sees, so it cannot simply be the next free enumerator.
Nor is there a free one: Linux uses all eight non-secure SGIs, and
8-15 belong to the secure world on the TF-A platforms this port
targets. The doorbell therefore takes SGI 7 from the KGDB roundup
whenever CONFIG_MULTIKERNEL is set. KGDB keeps working: without the
arch kgdb_roundup_cpus() it uses the generic one, which needs no IPI
of its own. Unlike the roundup, the doorbell is never an NMI, as it
drains a message ring.

There is no foreign doorbell to filter on the receiving side. All the
handler does is drain the kernel's own ring, which a kernel without an
instance relationship cannot have written to.

With this a spawn's HALTED notification reaches the host, which moves
the instance from active back to loaded, and a graceful halt requested
by the host reaches the spawn. The ring itself needed nothing: both
sides map it cacheable, as memremap(MEMREMAP_WB) is ioremap_cache() on
arm64.

Signed-off-by: Cong Wang <cwang@multikernel.io>
A GICv3 has one distributor. A kernel that multikernel spawns on some
of the machine's CPUs finds it in use by the kernel that spawned it,
and gic_dist_init() is the wrong thing to do to it: it disables the
distributor, puts every SPI into group 1, disables and deactivates
them all, and routes them to the boot CPU. The host loses its
interrupts the moment the new kernel probes the GIC.

Redistributors and CPU interfaces are per CPU, and nothing about them
changes: a kernel initialises those of the CPUs it runs on, and the
region walk only reads the frames of the others.

Add a tenant mode, selected by "multikernel,tenant" in the GIC node of
the tree such a kernel boots from:

 - The distributor is left as it is. It has to be enabled with
   affinity routing, which is checked, since nothing would work
   otherwise and the tenant is not the one to fix it.
 - Only the SPIs granted in "multikernel,spis", <first count> pairs of
   SPI numbers, can be mapped; every other SPI and all ESPIs fail with
   -EPERM. What the driver does to a mapped SPI is per-interrupt state
   (enable, priority, routing), apart from the read-modify-write of
   its ICFGR word, which it shares with 15 neighbours.
 - No MBIs, which are SPIs the tenant was not granted.

SGIs and PPIs are per redistributor and unchanged. There is no ITS in
such a tree, so no LPIs either.

Signed-off-by: Cong Wang <cwang@multikernel.io>
Mark the GIC node of every instance tree "multikernel,tenant", so the
spawn's GIC driver leaves the distributor to the host. Until now a
spawn took all SPIs away from the host as it booted.

An instance that owns no device owns no SPI. One that was given a
platform device gets the device's node, copied from the host tree with
reg translated to CPU addresses and the interrupts rewritten as plain
GIC specifiers, and each of those SPIs is granted in
"multikernel,spis". A spawn probes what its tree describes, so on
arm64 the node is the grant; the name based allowlist that x86 applies
in platform_device_add() never sees an OF device.

The host gives the SPIs up by unbinding its driver, which disables
them at the distributor, and takes them back when a driver binds
again, which sets them up from scratch. A device still bound at exec
is refused. Providers a node points to (clocks, resets, power domains,
pinctrl) are not followed, which limits this to self-contained devices
for now, and an ACPI host has no node to copy at all.

On QEMU virt a spawn now runs a virtio-mmio disk on SPI 47 while the
host keeps its UART and its own virtio devices, and the host binds the
disk again once the instance has halted.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The instance tree already describes the root buses above an instance's
PCI devices in the devicetree PCI binding, with the devices as the
only children, and the spawn's PCI core already skips every function
without a node. What a spawn on arm64 lacks is a way to generate
config cycles and a driver for the bridge node:

 - mk_dt_bridge_ecam() only knew x86's MMCONFIG regions. A root bus on
   the generic ECAM accessors, which is what both pci-host-generic and
   an ACPI host on arm64 use, keeps its window in sysdata; take it
   from there so the node gets its reg.
 - With a reg the node is an ordinary ECAM host bridge. arm64 names
   pci-host-ecam-generic as its second compatible, so the stock driver
   binds; x86 creates the bus from the node by hand because it has no
   such driver.
 - Where I/O space is memory mapped, an I/O window's resource holds
   logical port numbers, and the ranges entry claimed the window sits
   at CPU address 0. Translate it with pci_pio_to_address(); x86, whose
   ports are an address space of their own, is unaffected.
 - BARs and bridge windows stay as the host assigned them: the bridge
   is shared and the host keeps running on the other devices. Say so
   with linux,pci-probe-only in /chosen.

A spawn on QEMU virt now enumerates the virtio NIC it was given, and
nothing else, behind windows identical to the host's. The NIC does not
probe yet: INTx is wired to SPIs the host shares between devices, and
there is no MSI controller in the tree until the ITS can be used from
a spawn.

Signed-off-by: Cong Wang <cwang@multikernel.io>
An ITS has one command queue, owned by the kernel that booted the
machine. A kernel that multikernel spawns on some of the CPUs cannot
run the ITS driver against it, so it cannot map the MSIs of the PCI
devices it was given. The host has to do it for it.

Add the small API that takes:

 - its_foreign_map() maps an event of a DeviceID to an LPI from this
   kernel's allocator and routes it to the collection of a given CPU,
   or moves it there if it is mapped already. The CPU is one the other
   kernel runs on. Its collection is the one this kernel mapped while
   the CPU was its own, and stays valid after the CPU was given away.
   The same goes for the LPI tables: that redistributor keeps using
   the ones this kernel enabled, which cannot be taken back, so the
   configuration byte is written here too and the LPI is enabled from
   the start. No Linux interrupt exists for such an LPI on this side;
   it never reaches one of this kernel's CPU interfaces.
 - its_foreign_release() discards everything mapped for one owner and
   frees its ITTs and LPIs, for when that kernel is gone, whether it
   went in order or not.
 - its_foreign_doorbell() is the address such a device writes to.

A device that is mapped for another kernel is not an alias to share an
ITT with, so its_msi_prepare() refuses it.

Systems with more than one ITS are left out: picking the ITS behind a
device, or giving a whole ITS away, can come later.

Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 needs nothing to give an instance the MSIs of its PCI devices: the
message address names the destination LAPIC. On a GICv3 they go
through the ITS, which only the host can program, so an instance asks
the host over the message ring.

MK_IO_MSI_MAP names a device by its PCI address, an event and one of
the instance's CPUs. The host checks that the instance owns both,
looks up the DeviceID behind the device, maps the event with
its_foreign_map() from a work item, since messages arrive in interrupt
context and the ITS driver sleeps, and answers with the LPI in
MK_IO_MSI_ACK. Mapping an event again moves it, which is how an
affinity change is sent from under the descriptor lock without waiting
for an answer. There is no unmap: the host drops everything an
instance had when it halts, however it halts, and once more before it
is started, so a crashed instance leaves nothing behind and the host
can bind the device again right after a halt.

The instance's side is irq-gic-v3-its-mk.c, what is left of an ITS
driver when someone else runs the command queue: an MSI parent domain
on the GIC's LPI range, reusing the ITS MSI parent ops, that composes
messages from the ITS's doorbell address and the event. LPIs are not
masked there. The redistributors of the instance's CPUs run on the
host's LPI tables, so the enable bit is the host's to write, and the
MSI capability of the device can be masked without asking anyone.

arm64 puts the proxy node into the instance tree, with the doorbell
as its reg, and points the msi-map of the instance's root buses at it
with an identity mapping.

On QEMU virt a spawn now runs a virtio-net PCI device: it gets a DHCP
lease and pings through MSI-X vectors delivered as LPIs, while the
host keeps its own virtio devices on the same ITS. After a halt the
host binds the NIC and uses it with vectors of its own, gives it back,
and the next spawn takes it again.

Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 forces a wedged spawn down with an NMI per APIC ID. PSCI has no
call to turn another CPU off, so arm64 has to reach the CPU in band
and have it turn itself off.

The stop IPI already does that in a spawn: IPI_CPU_STOP lands in
local_cpu_stop(), which returns the CPU to firmware there. So
mk_force_stop_cpu() sends IPI_CPU_STOP_NMI to the MPIDR, and the
generic force halt, which arms its marker in the shared ring area
first, works unchanged. A kernel running with pseudo-NMIs takes that
SGI with interrupts masked; to any other kernel it is an ordinary
interrupt, and a CPU spinning with interrupts off is out of reach.
That case is not silent: the next exec finds the CPU still on through
AFFINITY_INFO and refuses to reload the image under it.

SGIs cross kernel boundaries, so the receiving side now tells its own
stop from a foreign one: smp_send_stop() and the crash path mark
theirs, and anything else is obeyed only if the force-halt marker is
up, which only a kernel with an instance relationship can set. x86
filters stop NMIs it did not ask for the same way. Obeying means
turning the CPU off, in a host that is being fenced as well, not
parking it in an image that is about to be replaced.

Tested on QEMU virt with a spawn CPU spinning in lkdtm's HARDLOCKUP:
with irqchip.gicv3_pseudo_nmi=1 in the spawn both CPUs go off and the
next exec boots; without it the exec is refused with the CPU named.

SDEI, which would reach a CPU regardless of how its kernel was booted,
is not implemented: it takes an SDEI_EVENT_SIGNAL call the firmware
driver does not have yet and TF-A to test against.

Signed-off-by: Cong Wang <cwang@multikernel.io>
An instance drives its PCI devices without an IOMMU of its own. On
arm64 the SMMU is shared and stays with the host, the instance tree
does not mention it, and the instance hands physical addresses to its
devices. Behind a translating SMMU that DMA faults, which is what a
spawn's virtio NIC does on QEMU virt with iommu=smmuv3: F_TRANSLATION
for its stream ID on the first buffer. Behind a bypassing one it
works, and a buggy driver in the instance can reach all of the host.

While an instance runs, attach each of its devices to a paging domain
of its own that maps the instance's memory grant onto itself, plus the
page of the MSI doorbell, which a device behind the SMMU cannot write
to otherwise. The device works untranslated and reaches nothing else.
This is the generic IOMMU API, the way VFIO would do it: claim DMA
ownership, which fails while a host driver is bound, allocate, map,
attach.

The domain is built on every exec, since the grant can change between
load and exec, and an instance whose devices cannot be confined is not
started. It follows the grant of a running instance: memory is mapped
before the instance is told about it and unmapped once the instance
has let go of it. It goes away when the instance halts, however it
halts, which returns the device to the host's default domain for a
host driver to bind. A device without an IOMMU is left alone, its DMA
as unrestricted as it always was.

Only arm64 gets this for now: x86 runs instance devices without an
IOMMU, and changing that is a change of its own.

arm64 also marks a root bus dma-coherent in the instance tree when the
host treats the devices behind it that way. Without the property every
device of the instance is non-coherent, which is correct on SBSA
hardware but pays for cache maintenance nobody needs.

On QEMU virt with an SMMUv3 the NIC of a spawn now gets its lease and
pings with no SMMU event logged, the host binds and uses it after a
halt, and the next spawn takes it again.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The user interface is the same on both architectures, so usage.rst
only gains a pointer. What differs is below it: firmware holds the
pool CPUs, the instance tree is the boot tree, the GIC is shared with
the host as a tenant, MSIs go through the host's ITS and the SMMU
confines a spawn's devices from the host's side. Document that, the
device tree properties it introduces, the command line a spawn wants,
and what the platform has to provide.

Signed-off-by: Cong Wang <cwang@multikernel.io>
…e device

The MSI routing for instances only worked on a machine with a single
ITS: its_foreign_map() took "the" ITS and gave up when there was more
than one, and the instance tree carried that one ITS's doorbell as the
reg of the proxy node. Servers have an ITS per socket at least, so on
those the PCI devices of an instance got no interrupts at all.

The ITS behind a device is named by the device's MSI domain, which the
host has for every PCI device whether a driver is bound or not. Pass
that domain to its_foreign_map() and its_foreign_doorbell() and take
the ITS, with its own DeviceID space, collections and lock, from
there; a domain that is not an ITS's yields none. The single-ITS
shortcut is gone rather than kept as a fast path, so every machine
runs the code a multi-ITS machine needs. its_foreign_release() walks
all of them. An ITS with erratum 23144 only reaches the CPUs of its
own node, which is checked when the collection is picked.

Only the host knows which ITS that is, so the doorbell moves from the
device tree into the answer: MK_IO_MSI_ACK carries the address next to
the LPI, the spawn's MSI domain keeps it per interrupt, and the proxy
node loses its reg. That also covers a root bus whose devices sit
behind different ITSs, and the pre-ITS doorbell, since the address now
comes from the ITS's own get_msi_base(). The generic pending-message
helper returns a single int, so the proxy tracks its few in-flight
requests itself to get the whole answer back to the waiter.

The instance tree now decides per root bus: a bus gets its msi-map when
a device on the host's side of it has an ITS behind it, and a warning
otherwise. The SMMU containment maps the doorbell of the ITS behind
each device instead of a global one.

QEMU has no machine with more than one ITS, so that case itself is
untested. The PCI smoke test passes on QEMU virt booted from a device
tree, under TF-A, from UEFI/ACPI where the domain is found through
IORT, and behind an SMMUv3; with its=off the instance boots, its NIC
fails to probe for lack of interrupts, and the host is unaffected.

Signed-off-by: Cong Wang <cwang@multikernel.io>
Two things stand in the way of a caller that has to exchange a message
with another kernel from a context that cannot sleep, such as an
interrupt chip callback under the descriptor lock.

Sending looks the instance up by ID, and mk_instance_find() takes a
mutex. Everything else on that path is fit for atomic context, down to
the GFP_ATOMIC allocation, and the console already gets there with
interrupts off. Add mk_send_message_to() and
multikernel_send_ipi_data_to() for a caller that holds a reference to
the instance, and build the ID variants on top of them.

Receiving depends on the doorbell interrupt, which such a caller may
be unable to take: it may be the doorbell CPU itself, or stop_machine()
may hold every CPU, as it does while a dying CPU migrates its
interrupts. Add mk_ipi_poll() to drain the ring from the caller's
context. The ring has a single consumer, which used to be implied by
the doorbell going to one CPU, so the drain now runs under a lock. The
interrupt waits for that lock instead of giving up: a poller that is
just leaving may not look at the ring again.

No change for existing callers.

Signed-off-by: Cong Wang <cwang@multikernel.io>
its_proxy.c sat in kernel/multikernel/ and called straight into the
GICv3 ITS driver. Nothing about it is generic: the problem it solves,
one command queue per ITS, belongs to the GIC, and the only other
interrupt controller with the same shape is GICv5's, which is arm64 as
well. x86, and anything that writes its MSIs straight at a CPU, never
needs it.

Move it to arch/arm64/multikernel/, together with its Kconfig symbol
and the message payload, which only its two ends read. The message
numbers stay in the generic header, which is the one registry for
them.

What the generic code needs from it becomes part of the architecture
interface, behind ARCH_HAS_MK_MSI_PROXY the way the host park area sits
behind ARCH_HAS_MK_HOST_PARK: mk_arch_msi_release(), called when an
instance halts or is about to be started, and mk_arch_msi_doorbell(),
which tells the IOMMU containment what a device has to be able to
write to. The request function the instance's MSI domain calls is
declared in asm/multikernel.h.

No functional change.

Signed-off-by: Cong Wang <cwang@multikernel.io>
its_foreign_map() moves an event that is mapped already, but it takes
the device allocation mutex, so it can only run in process context.
The kernel that owns the interrupt does not have that luxury: its
irq_set_affinity() callback runs with the descriptor locked, possibly
on a CPU that is on its way out and gets turned off right after. This
kernel's own its_set_affinity() is done when it returns, MOVI sent and
a pending interrupt taken along, and the other kernel needs the same
from its host.

Add its_foreign_move() for that. It may run in an interrupt handler of
this kernel, so what keeps the device from going away under it is a
raw spinlock: a move holds it from the lookup to the MOVI, and
its_foreign_release() takes it to disown a device before tearing it
down, after which no move finds it and none is under way. The mapping
path holds it around its ITS commands as well, so that a map and a
move of the same event do not interleave.

Signed-off-by: Cong Wang <cwang@multikernel.io>
An instance asked its host to move an MSI without waiting for the
answer, because irq_set_affinity() runs with the descriptor locked.
That is wrong in the case that matters most. A CPU going offline
migrates its interrupts from inside stop_machine() and is turned off
moments later, quite possibly before the host got to the request. LPIs
aimed at a redistributor whose CPU firmware has put to sleep are not
something the architecture promises to deliver later. A failure on
the host's side was invisible too, and sending looked the host up
under a mutex from atomic context.

Give the move its own request, MK_IO_MSI_MOVE, and make it synchronous
on both ends:

 - The instance sends it with mk_send_message_to() and polls the ring
   for the answer with mk_ipi_poll(), up to the second the ITS driver
   allows one of its own commands. Nobody else could take the
   doorbell: under stop_machine() every CPU has interrupts off. The
   host's verdict is what irq_set_affinity() returns. An interrupt
   stays where it is when the new mask allows that, which saves the
   round trip for most calls.
 - The host answers from its doorbell interrupt with
   its_foreign_move(). Neither the PCI device nor the instance can be
   looked up there, so MAP leaves a small route behind, the device's
   MSI domain and DeviceID and a reference to the instance, which
   MOVE finds under a raw spinlock and mk_arch_msi_release() drops.

Requests now live on the requester's stack and echo what they asked
for, so an answer finds its request whether it sleeps or polls.

The host also sizes the device's ITT by the vectors the device has,
MSI-X or MSI, instead of taking the number from the instance. The MSI
core happens to pass the table size today, so nothing was broken, but
how much the host allocates is not for an instance to decide.

Tested on QEMU virt: with the NIC's interrupts pinned to its second
CPU, a spawn takes that CPU offline and keeps its network, the
interrupts continuing on CPU 0, booted from a device tree, under TF-A
and from UEFI/ACPI. An NVMe disk, which frees and reallocates its
vectors while probing, works in a spawn as well.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The text offered irqchip.gicv3_pseudo_nmi=1 as an option. It is the
only way force halt reaches a spawn CPU that has interrupts masked,
and a CPU that cannot be stopped is lost to the host until reboot, so
say that it is required, what it takes to build and boot with it, that
kerf load adds the parameter, and where its reach ends.

Signed-off-by: Cong Wang <cwang@multikernel.io>
On a server every PCIe device sits behind a root port. A spawn has to
see that port, or it could not reach the bus behind it, so the
instance tree describes the bridges on the way down to a device. But
the port still belongs to the kernel that spawned it, which has its
port driver bound to it and may have other kernels' devices below it.

A spawn given a NIC behind a root port on QEMU virt treated the port
as its own:

 - pcieport bound to it a second time and asked the host for an MSI
   for a device the instance does not own;
 - the PCI core enabled the bridge on behalf of the NIC and rewrote
   its command register, which cleared the INTx disable bit the host
   had set, and it rewrites bridge windows and the port's side of ASPM
   the same way;
 - it cleared the port's AER status and enabled error reporting on it
   while it scanned the bus.

Decide once, when a device is set up and has its node, whether the
function is this kernel's: mk_pci_foreign() says it is not when the
spawn's tree describes it without naming a device, which is how a
bridge on the way down differs from a device that was given away. The
answer is kept in struct pci_dev as "foreign", before the first config
write, which is BAR sizing. The PCI core then

 - binds no driver to a foreign function, so none of the port services
   (AER, PME, hotplug, DPC) exist for it here;
 - drops config writes to it in pci_write_config_*(), next to the
   check that already drops them for a disconnected device, as if the
   function were read-only. That covers the core's own writes to
   parent bridges, the AER setup included, without chasing each of
   them: its owner configured it, and keeps it configured. A write
   from user space fails with -EPERM.

Reads are untouched, so enumeration works as before. The spawn's host
bridge keeps claiming the PCIe services as native: on a device the
spawn owns that is right, its error reporting should be on so that the
errors reach the port, where the host handles them.

With this the root port's config space is byte for byte the same at
boot, while a spawn uses the NIC behind it and after the spawn is
killed. The NIC works in the spawn, behind an SMMUv3 too, the host
binds and uses it in between, and a spawn CPU can go offline under its
interrupts.

Signed-off-by: Cong Wang <cwang@multikernel.io>
A server can describe its GIC redistributors one CPU at a time, with
an address in each MADT GICC entry instead of a GICR range. The host's
ACPI driver knows that each of those mappings holds one redistributor
and stops there. A spawn boots from a device tree, so the host turns
those addresses into separate reg entries in its GIC node.

The DT driver did not keep that boundary. It walked from one
redistributor to the next until GICR_TYPER.Last was set, even when
only one redistributor had been mapped. If Last is clear, the next
read is outside the mapping. This happens during interrupt setup,
before the spawn's console or the stop IPI can help. QEMU's usual
GICv3 redistributor ranges do not expose the lost boundary.

Keep the size alongside each redistributor mapping, for both DT and
ACPI, and stop the walk when it reaches that size. Last still ends a
walk early, and ACPI's single-redistributor mappings still stop after
one frame. Separate DT mappings no longer depend on the hardware
marking each of their redistributors as the last one.

The GICC conversion also assumed that every redistributor occupied
128 KiB. Read the distributor's architecture ID, as the native ACPI
driver does, so a GICv4 mapping covers its full 256 KiB, including the
virtual LPI frames.

A simulated MMIO test of the driver's walk reads beyond an individual
mapping before this change and stays inside it afterwards. Contiguous
ranges, GICv4 frames, explicit strides and the existing stopping rules
pass too. Both changed files compile for arm64. The server's first
boot and force-kill/re-exec cycle still need hardware validation.

Signed-off-by: Cong Wang <cwang@multikernel.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment