arm64: Port multikernel to arm64 - #40
Open
congwang-mk wants to merge 24 commits into
Open
congwang-mk wants to merge 24 commits into
congwang-mk wants to merge 24 commits into
Conversation
Multikernel only builds on x86 so far. Give arm64 the pieces the generic core needs from an architecture, so that the port can grow one function at a time in a tree that links. arm64 selects ARCH_SUPPORTS_MULTIKERNEL when it has PSCI and kexec_file: pool CPUs will be held by firmware in the PSCI OFF state and started with CPU_ON, and instances will be loaded as Image files. It does not select ARCH_HAS_MK_HOST_PARK, as there is no software park area to keep in pool memory. asm/multikernel.h names CPUs by MPIDR through the logical map and sizes the control block for the boot device tree alone: without a trampoline, a park page or identity page tables there is nothing else to carve out of instance memory. Every function of the architecture interface gets a stub. Spawning fails with -EOPNOTSUPP, so CONFIG_MULTIKERNEL=y is harmless on arm64 until the follow-up changes fill these in. Signed-off-by: Cong Wang <cwang@multikernel.io>
CPU_ON hands its third argument, the context ID, to the started CPU in x0. Secondary boot has no use for it and the driver hardcodes 0. Multikernel starts a CPU on a different kernel image, straight at that Image's entry point, and the arm64 boot protocol wants the device tree address in x0 there. CPU_ON can deliver it without any trampoline if the caller gets to choose the context ID. Add psci_cpu_on_context() for that, using whichever CPU_ON function ID the firmware was probed with. psci_ops.cpu_on() is unchanged. Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 has no firmware that controls CPUs after boot, so its multikernel code is mostly a software CPU lifecycle: trampolines, identity page tables, a park loop in host-owned memory and mailboxes to move parked CPUs between them. On arm64 firmware does all of that. A pool CPU is a CPU in the PSCI OFF state. The host gets it there with the stock hotplug path, cpu_psci_cpu_die() and cpu_psci_cpu_kill(), which depart_cpu() already ends in. Spawning is CPU_ON at the Image's entry point with the device tree address as context ID, which is the state head.S asks for, after cleaning what the host wrote to the PoC because the new kernel reads it with the caches off. An off CPU belongs to whoever turns it on next, so there is no slot to repark between. What remains of that bookkeeping is the question the core asks before it rewrites an image or frees instance memory: could a CPU still be executing in there? The instance's kernel turns on its secondaries without telling the host, so every CPU of a spawned instance counts as started, a hot-added one joins them, and a CPU leaves the set once AFFINITY_INFO reports it off. That reuses cpus_on_slot, whose comment claimed it stays empty on such architectures; say what it means there instead. Waiting for OFF matters on hot-remove too: the instance acknowledges the offline before the CPU has reached firmware, and CPU_ON of a CPU that is not off yet fails with ALREADY_ON. Sending the doorbell IPI, forced stop and the halt path of a spawn are still stubs. Signed-off-by: Cong Wang <cwang@multikernel.io>
Stock arm64 halts by sending the other CPUs to local_cpu_stop(), which leaves them in cpu_park_loop(), and spinning on the last one. All of them keep executing the kernel image, in a WFI loop with DAIF masked. That is fine when the machine resets next. For a spawn kernel it is the hazard the x86 park page exists to avoid: the host rewrites that image for the next spawn and the CPUs run whatever lands under them. PSCI cannot take a CPU away from outside either, so such a CPU is lost to the host for good. A spawn kernel returns every CPU to firmware instead: - machine_halt(), machine_power_off() and machine_restart() notify the host and turn all CPUs off through mk_halt_to_pool(). A spawn must not reach the PSCI SYSTEM_OFF or SYSTEM_RESET behind them, which would take the whole machine down, host included. - mk_enter_pool_state() is cpu_die() of the calling CPU, the boot CPU included: the host's CPUs are on, so firmware never sees the last CPU go. - local_cpu_stop(), where IPI_CPU_STOP and a parallel panic() put a CPU, turns it off instead of parking it. - panic_timeout defaults to reboot, which is the halt above, as on x86. The crash stop path already ends in cpu_die(). Secondaries need nothing: a spawn brings them up with the stock cpu_psci_cpu_boot() and hot-removes them with cpu_psci_cpu_die(), so there is no counterpart to x86's wakeup_secondary_cpu_64 hook, and the generic code never assumed one. The host learns that a halted instance's CPUs have arrived from AFFINITY_INFO, in mk_arch_confirm_parked(). mk_spawned() tells a kernel whether it booted with a manifest. The CPU_OFF path needs cpu_operations::cpu_die, so arm64 multikernel now depends on HOTPLUG_CPU, which the generic core needs to move CPUs into the pool anyway. Signed-off-by: Cong Wang <cwang@multikernel.io>
kexec_file on arm64 already does what loading an instance takes: it relocates an Image to any 2 MB aligned address, and the generic code picks that address inside the instance's memory grant for a KEXEC_TYPE_MULTIKERNEL image. image->start is the Image's entry point and image->arch.dtb_mem the device tree, the two things CPU_ON needs. There never was a purgatory to skip. What has to go is the preparation for replacing the running kernel, which an instance never does: - machine_kexec_post_load() builds the relocation code, a copy of the linear map and the EL2 vectors. A multikernel image is loaded in place and started on another CPU, so it needs none of them. It is not flushed here either: every spawn copies the segments in again, and mk_arch_spawn_instance() cleans them to the PoC afterwards, the one place that covers the first exec and every later one. - machine_kexec_prepare() refuses to load while CPUs are stuck in the kernel, which only matters to a kernel that is about to be left. The device tree the loader builds is a copy of this kernel's own with a new /chosen. It names every CPU, all of memory and every device, so an instance must not boot from it, and what an instance owns can change between load and exec anyway. Its boot tree is therefore rebuilt in place on every spawn. Size the segment for that, and refuse to spawn until the code that builds the tree exists, instead of handing the machine to the new kernel. With the tree a kimage segment, nothing of arm64's lives in the instance control block, so drop the fields reserved for it. Signed-off-by: Cong Wang <cwang@multikernel.io>
On x86 a spawn gets its memory map in boot_params and everything else in the manifest, a device tree on a channel of its own. On arm64 the tree in x0 is the only boot channel, and the boot path reads memory, CPUs, PSCI, the timer and the interrupt controller from it long before any multikernel code runs. So the two become one tree: the instance tree the generic code writes, with /resources and the multikernel /chosen, is the base, and arm64 adds what its boot path needs: - a /memory node per region of the grant, the role mk_e820_fill() has on x86; - /cpus, the instance's own CPUs first so that the boot CPU leads and they get the low logical numbers, then every other CPU of the machine. A kernel can only online CPUs it enumerated at boot, and the generic code already prunes the ones the instance does not own from the present mask before smp_init(), so mk_arch_register_cpu() stays empty; - bootargs, the initrd and the seeds, taken from the /chosen of the tree the kexec_file loader made; - PSCI, the arch timer and the GICv3. The last three come from whichever description the host booted with. A DT host copies its own nodes, translating the GIC frames to CPU addresses as the node moves to the root and leaving its ITS child behind. An ACPI host, the SBSA server case, writes the generic bindings from its tables: the PSCI conduit from the FADT, the timer PPIs from the GTDT, the distributor and redistributor frames from the MADT. Both produce the same nodes on QEMU's virt machine. Either way the spawn is a DT platform: a populated tree disables ACPI on arm64 without any quirk, and no UEFI is advertised. There is no UART node on purpose. earlycon takes an address on the command line, and a probed pl011 would claim an SPI the host owns. ECAM windows for granted PCI segments are left to the device work; without them a spawn sees no PCI bus at all. The tree is written into the kimage's device tree segment on every spawn, since the grant can change between load and exec, and that segment grows to hold the manifest with the boot nodes on top. On the other side, setup_arch() accepts the boot tree as the manifest when its root says so, which is what makes mk_spawned() true. An instance now boots on QEMU virt from a DT and from a UEFI/ACPI host: both of its CPUs come up, it halts by turning them off, and a second exec finds them off and boots again. Until the GIC driver learns not to reset the shared distributor, that boot takes the host's SPIs away. Signed-off-by: Cong Wang <cwang@multikernel.io>
gic_ipi_send_mask() takes a mask of logical CPUs, which can only name CPUs of the running kernel. The hardware has no such limit: ICC_SGI1R addresses a redistributor by affinity, and an SGI is per-target state there, so it reaches a CPU that a different kernel image runs on with nothing shared between the two. Multikernel needs exactly that for its doorbell. Add gic_v3_send_sgi_to_mpidr(), which refuses a target whose Aff0 needs the range selector when the GIC has none. Signed-off-by: Cong Wang <cwang@multikernel.io>
The message ring between kernels needs a doorbell, the role MULTIKERNEL_VECTOR has on x86. On arm64 that is an SGI sent to the target's MPIDR with gic_v3_send_sgi_to_mpidr(). Its number is ABI between kernels, since the sender picks the INTID the receiver sees, so it cannot simply be the next free enumerator. Nor is there a free one: Linux uses all eight non-secure SGIs, and 8-15 belong to the secure world on the TF-A platforms this port targets. The doorbell therefore takes SGI 7 from the KGDB roundup whenever CONFIG_MULTIKERNEL is set. KGDB keeps working: without the arch kgdb_roundup_cpus() it uses the generic one, which needs no IPI of its own. Unlike the roundup, the doorbell is never an NMI, as it drains a message ring. There is no foreign doorbell to filter on the receiving side. All the handler does is drain the kernel's own ring, which a kernel without an instance relationship cannot have written to. With this a spawn's HALTED notification reaches the host, which moves the instance from active back to loaded, and a graceful halt requested by the host reaches the spawn. The ring itself needed nothing: both sides map it cacheable, as memremap(MEMREMAP_WB) is ioremap_cache() on arm64. Signed-off-by: Cong Wang <cwang@multikernel.io>
A GICv3 has one distributor. A kernel that multikernel spawns on some of the machine's CPUs finds it in use by the kernel that spawned it, and gic_dist_init() is the wrong thing to do to it: it disables the distributor, puts every SPI into group 1, disables and deactivates them all, and routes them to the boot CPU. The host loses its interrupts the moment the new kernel probes the GIC. Redistributors and CPU interfaces are per CPU, and nothing about them changes: a kernel initialises those of the CPUs it runs on, and the region walk only reads the frames of the others. Add a tenant mode, selected by "multikernel,tenant" in the GIC node of the tree such a kernel boots from: - The distributor is left as it is. It has to be enabled with affinity routing, which is checked, since nothing would work otherwise and the tenant is not the one to fix it. - Only the SPIs granted in "multikernel,spis", <first count> pairs of SPI numbers, can be mapped; every other SPI and all ESPIs fail with -EPERM. What the driver does to a mapped SPI is per-interrupt state (enable, priority, routing), apart from the read-modify-write of its ICFGR word, which it shares with 15 neighbours. - No MBIs, which are SPIs the tenant was not granted. SGIs and PPIs are per redistributor and unchanged. There is no ITS in such a tree, so no LPIs either. Signed-off-by: Cong Wang <cwang@multikernel.io>
Mark the GIC node of every instance tree "multikernel,tenant", so the spawn's GIC driver leaves the distributor to the host. Until now a spawn took all SPIs away from the host as it booted. An instance that owns no device owns no SPI. One that was given a platform device gets the device's node, copied from the host tree with reg translated to CPU addresses and the interrupts rewritten as plain GIC specifiers, and each of those SPIs is granted in "multikernel,spis". A spawn probes what its tree describes, so on arm64 the node is the grant; the name based allowlist that x86 applies in platform_device_add() never sees an OF device. The host gives the SPIs up by unbinding its driver, which disables them at the distributor, and takes them back when a driver binds again, which sets them up from scratch. A device still bound at exec is refused. Providers a node points to (clocks, resets, power domains, pinctrl) are not followed, which limits this to self-contained devices for now, and an ACPI host has no node to copy at all. On QEMU virt a spawn now runs a virtio-mmio disk on SPI 47 while the host keeps its UART and its own virtio devices, and the host binds the disk again once the instance has halted. Signed-off-by: Cong Wang <cwang@multikernel.io>
The instance tree already describes the root buses above an instance's PCI devices in the devicetree PCI binding, with the devices as the only children, and the spawn's PCI core already skips every function without a node. What a spawn on arm64 lacks is a way to generate config cycles and a driver for the bridge node: - mk_dt_bridge_ecam() only knew x86's MMCONFIG regions. A root bus on the generic ECAM accessors, which is what both pci-host-generic and an ACPI host on arm64 use, keeps its window in sysdata; take it from there so the node gets its reg. - With a reg the node is an ordinary ECAM host bridge. arm64 names pci-host-ecam-generic as its second compatible, so the stock driver binds; x86 creates the bus from the node by hand because it has no such driver. - Where I/O space is memory mapped, an I/O window's resource holds logical port numbers, and the ranges entry claimed the window sits at CPU address 0. Translate it with pci_pio_to_address(); x86, whose ports are an address space of their own, is unaffected. - BARs and bridge windows stay as the host assigned them: the bridge is shared and the host keeps running on the other devices. Say so with linux,pci-probe-only in /chosen. A spawn on QEMU virt now enumerates the virtio NIC it was given, and nothing else, behind windows identical to the host's. The NIC does not probe yet: INTx is wired to SPIs the host shares between devices, and there is no MSI controller in the tree until the ITS can be used from a spawn. Signed-off-by: Cong Wang <cwang@multikernel.io>
An ITS has one command queue, owned by the kernel that booted the machine. A kernel that multikernel spawns on some of the CPUs cannot run the ITS driver against it, so it cannot map the MSIs of the PCI devices it was given. The host has to do it for it. Add the small API that takes: - its_foreign_map() maps an event of a DeviceID to an LPI from this kernel's allocator and routes it to the collection of a given CPU, or moves it there if it is mapped already. The CPU is one the other kernel runs on. Its collection is the one this kernel mapped while the CPU was its own, and stays valid after the CPU was given away. The same goes for the LPI tables: that redistributor keeps using the ones this kernel enabled, which cannot be taken back, so the configuration byte is written here too and the LPI is enabled from the start. No Linux interrupt exists for such an LPI on this side; it never reaches one of this kernel's CPU interfaces. - its_foreign_release() discards everything mapped for one owner and frees its ITTs and LPIs, for when that kernel is gone, whether it went in order or not. - its_foreign_doorbell() is the address such a device writes to. A device that is mapped for another kernel is not an alias to share an ITT with, so its_msi_prepare() refuses it. Systems with more than one ITS are left out: picking the ITS behind a device, or giving a whole ITS away, can come later. Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 needs nothing to give an instance the MSIs of its PCI devices: the message address names the destination LAPIC. On a GICv3 they go through the ITS, which only the host can program, so an instance asks the host over the message ring. MK_IO_MSI_MAP names a device by its PCI address, an event and one of the instance's CPUs. The host checks that the instance owns both, looks up the DeviceID behind the device, maps the event with its_foreign_map() from a work item, since messages arrive in interrupt context and the ITS driver sleeps, and answers with the LPI in MK_IO_MSI_ACK. Mapping an event again moves it, which is how an affinity change is sent from under the descriptor lock without waiting for an answer. There is no unmap: the host drops everything an instance had when it halts, however it halts, and once more before it is started, so a crashed instance leaves nothing behind and the host can bind the device again right after a halt. The instance's side is irq-gic-v3-its-mk.c, what is left of an ITS driver when someone else runs the command queue: an MSI parent domain on the GIC's LPI range, reusing the ITS MSI parent ops, that composes messages from the ITS's doorbell address and the event. LPIs are not masked there. The redistributors of the instance's CPUs run on the host's LPI tables, so the enable bit is the host's to write, and the MSI capability of the device can be masked without asking anyone. arm64 puts the proxy node into the instance tree, with the doorbell as its reg, and points the msi-map of the instance's root buses at it with an identity mapping. On QEMU virt a spawn now runs a virtio-net PCI device: it gets a DHCP lease and pings through MSI-X vectors delivered as LPIs, while the host keeps its own virtio devices on the same ITS. After a halt the host binds the NIC and uses it with vectors of its own, gives it back, and the next spawn takes it again. Signed-off-by: Cong Wang <cwang@multikernel.io>
x86 forces a wedged spawn down with an NMI per APIC ID. PSCI has no call to turn another CPU off, so arm64 has to reach the CPU in band and have it turn itself off. The stop IPI already does that in a spawn: IPI_CPU_STOP lands in local_cpu_stop(), which returns the CPU to firmware there. So mk_force_stop_cpu() sends IPI_CPU_STOP_NMI to the MPIDR, and the generic force halt, which arms its marker in the shared ring area first, works unchanged. A kernel running with pseudo-NMIs takes that SGI with interrupts masked; to any other kernel it is an ordinary interrupt, and a CPU spinning with interrupts off is out of reach. That case is not silent: the next exec finds the CPU still on through AFFINITY_INFO and refuses to reload the image under it. SGIs cross kernel boundaries, so the receiving side now tells its own stop from a foreign one: smp_send_stop() and the crash path mark theirs, and anything else is obeyed only if the force-halt marker is up, which only a kernel with an instance relationship can set. x86 filters stop NMIs it did not ask for the same way. Obeying means turning the CPU off, in a host that is being fenced as well, not parking it in an image that is about to be replaced. Tested on QEMU virt with a spawn CPU spinning in lkdtm's HARDLOCKUP: with irqchip.gicv3_pseudo_nmi=1 in the spawn both CPUs go off and the next exec boots; without it the exec is refused with the CPU named. SDEI, which would reach a CPU regardless of how its kernel was booted, is not implemented: it takes an SDEI_EVENT_SIGNAL call the firmware driver does not have yet and TF-A to test against. Signed-off-by: Cong Wang <cwang@multikernel.io>
An instance drives its PCI devices without an IOMMU of its own. On arm64 the SMMU is shared and stays with the host, the instance tree does not mention it, and the instance hands physical addresses to its devices. Behind a translating SMMU that DMA faults, which is what a spawn's virtio NIC does on QEMU virt with iommu=smmuv3: F_TRANSLATION for its stream ID on the first buffer. Behind a bypassing one it works, and a buggy driver in the instance can reach all of the host. While an instance runs, attach each of its devices to a paging domain of its own that maps the instance's memory grant onto itself, plus the page of the MSI doorbell, which a device behind the SMMU cannot write to otherwise. The device works untranslated and reaches nothing else. This is the generic IOMMU API, the way VFIO would do it: claim DMA ownership, which fails while a host driver is bound, allocate, map, attach. The domain is built on every exec, since the grant can change between load and exec, and an instance whose devices cannot be confined is not started. It follows the grant of a running instance: memory is mapped before the instance is told about it and unmapped once the instance has let go of it. It goes away when the instance halts, however it halts, which returns the device to the host's default domain for a host driver to bind. A device without an IOMMU is left alone, its DMA as unrestricted as it always was. Only arm64 gets this for now: x86 runs instance devices without an IOMMU, and changing that is a change of its own. arm64 also marks a root bus dma-coherent in the instance tree when the host treats the devices behind it that way. Without the property every device of the instance is non-coherent, which is correct on SBSA hardware but pays for cache maintenance nobody needs. On QEMU virt with an SMMUv3 the NIC of a spawn now gets its lease and pings with no SMMU event logged, the host binds and uses it after a halt, and the next spawn takes it again. Signed-off-by: Cong Wang <cwang@multikernel.io>
The user interface is the same on both architectures, so usage.rst only gains a pointer. What differs is below it: firmware holds the pool CPUs, the instance tree is the boot tree, the GIC is shared with the host as a tenant, MSIs go through the host's ITS and the SMMU confines a spawn's devices from the host's side. Document that, the device tree properties it introduces, the command line a spawn wants, and what the platform has to provide. Signed-off-by: Cong Wang <cwang@multikernel.io>
…e device The MSI routing for instances only worked on a machine with a single ITS: its_foreign_map() took "the" ITS and gave up when there was more than one, and the instance tree carried that one ITS's doorbell as the reg of the proxy node. Servers have an ITS per socket at least, so on those the PCI devices of an instance got no interrupts at all. The ITS behind a device is named by the device's MSI domain, which the host has for every PCI device whether a driver is bound or not. Pass that domain to its_foreign_map() and its_foreign_doorbell() and take the ITS, with its own DeviceID space, collections and lock, from there; a domain that is not an ITS's yields none. The single-ITS shortcut is gone rather than kept as a fast path, so every machine runs the code a multi-ITS machine needs. its_foreign_release() walks all of them. An ITS with erratum 23144 only reaches the CPUs of its own node, which is checked when the collection is picked. Only the host knows which ITS that is, so the doorbell moves from the device tree into the answer: MK_IO_MSI_ACK carries the address next to the LPI, the spawn's MSI domain keeps it per interrupt, and the proxy node loses its reg. That also covers a root bus whose devices sit behind different ITSs, and the pre-ITS doorbell, since the address now comes from the ITS's own get_msi_base(). The generic pending-message helper returns a single int, so the proxy tracks its few in-flight requests itself to get the whole answer back to the waiter. The instance tree now decides per root bus: a bus gets its msi-map when a device on the host's side of it has an ITS behind it, and a warning otherwise. The SMMU containment maps the doorbell of the ITS behind each device instead of a global one. QEMU has no machine with more than one ITS, so that case itself is untested. The PCI smoke test passes on QEMU virt booted from a device tree, under TF-A, from UEFI/ACPI where the domain is found through IORT, and behind an SMMUv3; with its=off the instance boots, its NIC fails to probe for lack of interrupts, and the host is unaffected. Signed-off-by: Cong Wang <cwang@multikernel.io>
Two things stand in the way of a caller that has to exchange a message with another kernel from a context that cannot sleep, such as an interrupt chip callback under the descriptor lock. Sending looks the instance up by ID, and mk_instance_find() takes a mutex. Everything else on that path is fit for atomic context, down to the GFP_ATOMIC allocation, and the console already gets there with interrupts off. Add mk_send_message_to() and multikernel_send_ipi_data_to() for a caller that holds a reference to the instance, and build the ID variants on top of them. Receiving depends on the doorbell interrupt, which such a caller may be unable to take: it may be the doorbell CPU itself, or stop_machine() may hold every CPU, as it does while a dying CPU migrates its interrupts. Add mk_ipi_poll() to drain the ring from the caller's context. The ring has a single consumer, which used to be implied by the doorbell going to one CPU, so the drain now runs under a lock. The interrupt waits for that lock instead of giving up: a poller that is just leaving may not look at the ring again. No change for existing callers. Signed-off-by: Cong Wang <cwang@multikernel.io>
its_proxy.c sat in kernel/multikernel/ and called straight into the GICv3 ITS driver. Nothing about it is generic: the problem it solves, one command queue per ITS, belongs to the GIC, and the only other interrupt controller with the same shape is GICv5's, which is arm64 as well. x86, and anything that writes its MSIs straight at a CPU, never needs it. Move it to arch/arm64/multikernel/, together with its Kconfig symbol and the message payload, which only its two ends read. The message numbers stay in the generic header, which is the one registry for them. What the generic code needs from it becomes part of the architecture interface, behind ARCH_HAS_MK_MSI_PROXY the way the host park area sits behind ARCH_HAS_MK_HOST_PARK: mk_arch_msi_release(), called when an instance halts or is about to be started, and mk_arch_msi_doorbell(), which tells the IOMMU containment what a device has to be able to write to. The request function the instance's MSI domain calls is declared in asm/multikernel.h. No functional change. Signed-off-by: Cong Wang <cwang@multikernel.io>
its_foreign_map() moves an event that is mapped already, but it takes the device allocation mutex, so it can only run in process context. The kernel that owns the interrupt does not have that luxury: its irq_set_affinity() callback runs with the descriptor locked, possibly on a CPU that is on its way out and gets turned off right after. This kernel's own its_set_affinity() is done when it returns, MOVI sent and a pending interrupt taken along, and the other kernel needs the same from its host. Add its_foreign_move() for that. It may run in an interrupt handler of this kernel, so what keeps the device from going away under it is a raw spinlock: a move holds it from the lookup to the MOVI, and its_foreign_release() takes it to disown a device before tearing it down, after which no move finds it and none is under way. The mapping path holds it around its ITS commands as well, so that a map and a move of the same event do not interleave. Signed-off-by: Cong Wang <cwang@multikernel.io>
An instance asked its host to move an MSI without waiting for the answer, because irq_set_affinity() runs with the descriptor locked. That is wrong in the case that matters most. A CPU going offline migrates its interrupts from inside stop_machine() and is turned off moments later, quite possibly before the host got to the request. LPIs aimed at a redistributor whose CPU firmware has put to sleep are not something the architecture promises to deliver later. A failure on the host's side was invisible too, and sending looked the host up under a mutex from atomic context. Give the move its own request, MK_IO_MSI_MOVE, and make it synchronous on both ends: - The instance sends it with mk_send_message_to() and polls the ring for the answer with mk_ipi_poll(), up to the second the ITS driver allows one of its own commands. Nobody else could take the doorbell: under stop_machine() every CPU has interrupts off. The host's verdict is what irq_set_affinity() returns. An interrupt stays where it is when the new mask allows that, which saves the round trip for most calls. - The host answers from its doorbell interrupt with its_foreign_move(). Neither the PCI device nor the instance can be looked up there, so MAP leaves a small route behind, the device's MSI domain and DeviceID and a reference to the instance, which MOVE finds under a raw spinlock and mk_arch_msi_release() drops. Requests now live on the requester's stack and echo what they asked for, so an answer finds its request whether it sleeps or polls. The host also sizes the device's ITT by the vectors the device has, MSI-X or MSI, instead of taking the number from the instance. The MSI core happens to pass the table size today, so nothing was broken, but how much the host allocates is not for an instance to decide. Tested on QEMU virt: with the NIC's interrupts pinned to its second CPU, a spawn takes that CPU offline and keeps its network, the interrupts continuing on CPU 0, booted from a device tree, under TF-A and from UEFI/ACPI. An NVMe disk, which frees and reallocates its vectors while probing, works in a spawn as well. Signed-off-by: Cong Wang <cwang@multikernel.io>
The text offered irqchip.gicv3_pseudo_nmi=1 as an option. It is the only way force halt reaches a spawn CPU that has interrupts masked, and a CPU that cannot be stopped is lost to the host until reboot, so say that it is required, what it takes to build and boot with it, that kerf load adds the parameter, and where its reach ends. Signed-off-by: Cong Wang <cwang@multikernel.io>
On a server every PCIe device sits behind a root port. A spawn has to see that port, or it could not reach the bus behind it, so the instance tree describes the bridges on the way down to a device. But the port still belongs to the kernel that spawned it, which has its port driver bound to it and may have other kernels' devices below it. A spawn given a NIC behind a root port on QEMU virt treated the port as its own: - pcieport bound to it a second time and asked the host for an MSI for a device the instance does not own; - the PCI core enabled the bridge on behalf of the NIC and rewrote its command register, which cleared the INTx disable bit the host had set, and it rewrites bridge windows and the port's side of ASPM the same way; - it cleared the port's AER status and enabled error reporting on it while it scanned the bus. Decide once, when a device is set up and has its node, whether the function is this kernel's: mk_pci_foreign() says it is not when the spawn's tree describes it without naming a device, which is how a bridge on the way down differs from a device that was given away. The answer is kept in struct pci_dev as "foreign", before the first config write, which is BAR sizing. The PCI core then - binds no driver to a foreign function, so none of the port services (AER, PME, hotplug, DPC) exist for it here; - drops config writes to it in pci_write_config_*(), next to the check that already drops them for a disconnected device, as if the function were read-only. That covers the core's own writes to parent bridges, the AER setup included, without chasing each of them: its owner configured it, and keeps it configured. A write from user space fails with -EPERM. Reads are untouched, so enumeration works as before. The spawn's host bridge keeps claiming the PCIe services as native: on a device the spawn owns that is right, its error reporting should be on so that the errors reach the port, where the host handles them. With this the root port's config space is byte for byte the same at boot, while a spawn uses the NIC behind it and after the spawn is killed. The NIC works in the spawn, behind an SMMUv3 too, the host binds and uses it in between, and a spawn CPU can go offline under its interrupts. Signed-off-by: Cong Wang <cwang@multikernel.io>
A server can describe its GIC redistributors one CPU at a time, with an address in each MADT GICC entry instead of a GICR range. The host's ACPI driver knows that each of those mappings holds one redistributor and stops there. A spawn boots from a device tree, so the host turns those addresses into separate reg entries in its GIC node. The DT driver did not keep that boundary. It walked from one redistributor to the next until GICR_TYPER.Last was set, even when only one redistributor had been mapped. If Last is clear, the next read is outside the mapping. This happens during interrupt setup, before the spawn's console or the stop IPI can help. QEMU's usual GICv3 redistributor ranges do not expose the lost boundary. Keep the size alongside each redistributor mapping, for both DT and ACPI, and stop the walk when it reaches that size. Last still ends a walk early, and ACPI's single-redistributor mappings still stop after one frame. Separate DT mappings no longer depend on the hardware marking each of their redistributors as the last one. The GICC conversion also assumed that every redistributor occupied 128 KiB. Read the distributor's architecture ID, as the native ACPI driver does, so a GICv4 mapping covers its full 256 KiB, including the virtual LPI frames. A simulated MMIO test of the driver's walk reads beyond an individual mapping before this change and stays inside it afterwards. Contiguous ranges, GICv4 frames, explicit strides and the existing stopping rules pass too. Both changed files compile for arm64. The server's first boot and force-kill/re-exec cycle still need hardware validation. Signed-off-by: Cong Wang <cwang@multikernel.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This series ports multikernel to arm64, following the plan in the tracking issue #21.
PSCI does the CPU lifecycle on arm64: a pool CPU is simply OFF, a spawn starts with
CPU_ON, and a spawn halts by turning every CPU off. There is no arm64 counterpart of the x86 trampolines or park loop. Most of the work is in the interrupt controller. A spawn is a tenant of the host's GICv3 and never initialises the distributor. The host proxies ITS commands, so a spawn's MSIs work through an ITS that the host owns.What is in it
Foundation and CPU lifecycle
asm/multikernel.hand arch stubs (arm64: Kconfig, asm/multikernel.h and arch stubs #9)firmware/psci(arm64: PSCI-based spawn, confirm-parked and release #10)kexec_fileImage loader for an instance (arm64: KEXEC_TYPE_MULTIKERNEL Image loader, placement in the grant, PoC cleaning #12)Boot data and messaging
Interrupts
Robustness and isolation
Documentation
Documentation/multikernel/arm64.rst(arm64: QEMU virt + TF-A test environment, kerf support and docs #20)Testing
virt(GICv3 with ITS): themk-smokescenarios for lifecycle, hotplug, PCI, NVMe, offline and wedge pass with DT boot, with TF-A (the host entered at EL2), and with UEFI/ACPI. PCI also passes behind SMMUv3 and behind a PCIe root port.arm64branch): init, create, load, exec, console, CPU and memory update, graceful and forced kill.kerf console.Not tested yet: more than one ITS on real hardware (QEMU has a single ITS), crash healing on arm64, and PCIe switches with devices of several owners below one port.
Left open
Closes #9
Closes #10
Closes #11
Closes #12
Closes #13
Closes #14
Closes #15
Closes #16
Closes #18
Closes #20
🤖 Generated with Claude Code