Skip to content

unshare: support for systemd-nsresourced, --map-foreign option - #4560

Draft
Skyb0rg007 wants to merge 2 commits into
util-linux:masterfrom
Skyb0rg007:unshare-nsresourced
Draft

Skyb0rg007 wants to merge 2 commits into
util-linux:masterfrom
Skyb0rg007:unshare-nsresourced

Conversation

@Skyb0rg007

Copy link
Copy Markdown
Contributor

Adds the --map-foreign option, which uses systemd-nsresourced.service(8) to setup the uid and gid maps, specifically for asking the service to map the foreign UID range. The --map-current-user and --map-root-user options are also compatible: --map-current-user is the default, and --map-root-user translates into "target": 0.

systemd-nsresourced has other options, but the primary use case that I envision unshare to be used for is to manage files owned by the foreign-0-foreign-65534 UIDs and GIDs.
Specifically, running unshare --map-foreign --map-root-user -- rm -rf ./dir to delete local files that were created using systemd-nspawn(1) or systemd-mountfsd.service(8)'s io.systemd.MountFileSystem.MakeDirectory Varlink method.

The important thing to note is that the foreign UID range is a systemd concept, so newuidmap is not usually able to map those uids into the user namespace.

This is a draft PR to garner feedback, as this is a new idea.

@Skyb0rg007
Skyb0rg007 marked this pull request as draft August 13, 2026 23:14
Comment thread meson.build Outdated
Comment thread meson.build Outdated
Comment thread sys-utils/unshare.c Outdated
Comment thread sys-utils/unshare.c Outdated
Comment thread sys-utils/unshare.c Outdated
Comment thread sys-utils/unshare.c Outdated
unshare_flags |= CLONE_NEWUSER;
mapuser = real_euid;
mapgroup = real_egid;
target_mode = TARGET_CURRENT;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we really need target_mode? It's also a variable set in multiple getopt cases, so the final content depends on the order of the command-line options. Seems fragile.

What about adding maplocal for OPT_MAPUSER/MAPGROUP and rejecting its use with map_foreign? And the final switch with target_mode y can be replaced with

  nsresourced_target = (mapuser == 0) ? 0 : -1

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nsresourced lets you configure a different target user when using the "type": "managed" option. I stripped out all that code for this PR so some of these choices probably seem odd.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you assume we will add some other types, then it will make more sense. (Note that I don't know anything about nsresourced.)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I’m honestly not sure — that’s the main reason I marked this PR as draft. I would use the mode I implemented here; the other modes are more specialized so I’d personally not use them so I need others to make that determination.

@karelzak

Copy link
Copy Markdown
Collaborator

BTW, in include/optutils.h, we have a way to detect collisions between command-line options.

static const ul_excl_t excl[] = {
      { OPT_MAP_FOREIGN, OPT_MAPUSERS },
      { OPT_MAP_FOREIGN, OPT_MAPGROUPS },
      { OPT_MAP_FOREIGN, OPT_MAPAUTO },
      { OPT_MAP_FOREIGN, OPT_MAPSUBIDS },
      { OPT_MAP_FOREIGN, OPT_OWNER },
      { 0 }
  };
  int excl_st[ARRAY_SIZE(excl)] = UL_EXCL_STATUS_INIT;

and err_exclusive_options(c, longopts, excl, excl_st) in the getopt's while(). It will help with code incompatibility with options. The --map-user/--map-group needs to be handled more carefully (the maplocal flag?).

Ah, now I see OPT_MAP_FOREIGN vs. OPT_MAPFOREIGN name mess ;-)

Adds the `--map-foreign` option, which uses systemd-nsresourced to
setup the uid and gid maps, additionally asking the service to map the
foreign UID range. The `--map-current-user` and `--map-root-user`
options are also compatible.
Use the builtin command-line collision helpers, and properly follow
dynamic library loading conventions.
@karelzak

karelzak commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Thanks for the patch, and sorry for the slow follow-up. The approach looks right to me: the varlink payload matches the io.systemd.NamespaceResource IDL exactly (size: 1 and target 0-or-unset are what type: "self" requires), the dlopen wrapper mirrors the existing dl-systemd pair faithfully, and the ul_excl_t table is correctly ordered. Build systems, bash completion, man page and tests are all there too — nice work.

A few things to fix before this can leave draft state.

major

1. --map-foreign --setgroups=deny can never work — sys-utils/unshare.c:1344, sys-utils/unshare.1.adoc:150

setgroups_control(setgrpcmd) runs after nsresourced has installed the gid map, and the kernel refuses to write deny once gid_map.nr_extents != 0 (kernel/user_namespace.c:proc_setgroups_write), so the write fails with EPERM and setgroups_control() calls err(EXIT_FAILURE, _("write failed %s")). The man page explicitly recommends this combination:

Since they are written by a privileged process, --setgroups=deny is not implied; use --setgroups=deny explicitly if that is required.

It is also unnecessary — for type: "self" allocations nsresourced already denies setgroups through the BPF-LSM, deliberately not via /proc/self/setgroups (see the comment at src/nsresourced/nsresourcework.c:1527). Suggest either dropping that sentence from the man page, or calling setgroups_control(SETGROUPS_DENY) before the varlink call when --setgroups=deny was requested.

2. The map-root-user subtest fails when the test suite runs as root — tests/ts/unshare/map-foreign:36

normalize_map() does s/\<$1\>/SELF/g with $1 = $(id -ru). As root that is 0, which also rewrites the inner uid column of 0 0 1, producing SELF SELF 1 while tests/expected/unshare/map-foreign-map-root-user contains 0 SELF 1. make check is commonly run as root. Anchoring the substitution to the second column, e.g.

sed -e "s/^\([0-9]\+\) $1 /\1 SELF /"

fixes this and issue 7 below.

3. No guard for "user namespaces unavailable" — tests/ts/unshare/map-foreign:26-32

The two top-level guards cover "built without varlink" and "no /run/systemd/io.systemd.NamespaceResource", but not the case where unshare(CLONE_NEWUSER) itself fails (restricted sysctl, sandboxed CI container). The resulting unshare failed matches neither guard nor check_nsresource(), so all four subtests fail instead of skipping. tests/ts/unshare/forward-signals:29-31 has the probe to copy:

"$TS_CMD_UNSHARE" --user --map-root-user /bin/true &>/dev/null || ts_skip "user namespaces not supported"

minor

4. meson.build:462 — cc.has_header_symbol() is not gated on lib_systemd.found(), unlike the neighbouring cc.has_function() checks. With -Dsystemd=disabled on a host that has systemd-devel, HAVE_DECL_SD_VARLINK_CONNECT_ADDRESS ends up 1 in meson but undefined in autotools. Harmless for the build (the header also requires HAVE_LIBSYSTEMD), but it makes tools/compare-buildsys.sh diff. have_systemd_varlink = lib_systemd.found() and cc.has_header_symbol(...) is enough.

5. tests/ts/unshare/map-foreign:42 — the skip pattern matches only two of the five errx() paths in allocate_user_range(). unable to allow fd passing, unable to push fd and unable to make varlink call would be reported as test failures against an old or partial nsresourced.

6. tests/ts/unshare/map-foreign:24 — FOREIGN_BASE=2147352576 is systemd's default foreign-uid-base, but that is a build-time meson option. Deriving it from id -u foreign-0 (as the new man-page example does) would be more robust.

7. tests/ts/unshare/map-foreign:36 — same unanchored substitution as issue 2: a real uid/gid of 1 or 65536 also rewrites the trailing count/size column.

8. Commits — neither commit carries a Signed-off-by line (please use git commit -s), and "unshare: code cleanup" is a cleanup of the preceding commit, so the two should be squashed.

nit

9. sys-utils/unshare.c:1424 — the foreign parameter of allocate_user_range() is always true at its only call site.

10. sys-utils/unshare.c:1447 — systemd's own client passes mangleName: true and special-cases the io.systemd.NamespaceResource.UserNamespaceInterfaceNotSupported error id for a readable message (src/shared/nsresource.c:nsresource_allocate_userns_full()). Worth copying; unshare-<pid> can also collide after PID reuse while an older registration is still alive.

11. sys-utils/unshare.c:1433-1458 — errno = -r; err(...) is the more usual idiom in the tree than errx(..., strerror(-r)).

12. sys-utils/unshare.c:1327 — wrapping the existing map_id() block in if (!mapforeign) { ... } re-indents it; folding !mapforeign into the two existing conditions would keep the diff smaller.

13. po/POTFILES.in is not updated for the two new files (they contain no translatable strings, but include/dl-systemd.h and lib/dl-systemd.c are both listed), and neither the configure nor the meson build summary reports varlink support.

One last note: unshare is in UL_STATIC_PROGRAMS and no distro ships libsystemd.a, so --enable-static-programs=unshare now needs --without-systemd (which tools/config-gen.d/static.conf already passes). That is fine — static unshare is a specialized use case — just worth knowing.

— assisted by Claude Code

@Skyb0rg007

Copy link
Copy Markdown
Contributor Author

Don't feel bad taking a long time to review -- I marked it as draft in hopes that someone else would be interested in this functionality; I'm not that interested in merging this until at least one other person has some input into what the design would look like. I guess I should've pinged @brauner since he liked my discussion post.

But I'll try to go over those comments when I have the time.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants