Capabilities
Every Motor OS process has an immutable 64-bit capability mask, set when the process is created. Each set bit allows a class of operations, checked by the kernel or by the service that performs the operation. This page lists every capability, how a process passes capabilities to its children, and how rush and rmux use them.
A process cannot change its own capabilities. It can only ask for a mask for a new
child, and the kernel decides whether the parent may grant it. The definitions and the
default child policy live in moto_sys::caps
(src/sys/lib/moto-sys/src/caps.rs); the design notes are in
docs/caps.md and docs/process-roles.md. For the process model around
them, see Processes and roles.
The capability bits
Masks are written in hexadecimal and combined with bitwise OR.
| Capability | Bit | Mask | What it allows |
|---|---|---|---|
CAP_SYS | 0 | 0x01 | The System role: protection from being killed by ordinary userspace processes, and broad authority to grant capabilities to children. |
CAP_IO_MANAGER | 1 | 0x02 | IO-manager operations: mapping device memory (MMIO) and the serial console. Held by sys-io. |
CAP_SPAWN | 2 | 0x04 | Creating processes. The kernel checks it when a child address space is created. |
CAP_LOG | 3 | 0x08 | Writing records to the kernel log and to strobe, the logging service. |
CAP_SHUTDOWN | 4 | 0x10 | Shutting the system down. |
CAP_SPAWN_DETACHED | 5 | 0x20 | Creating detached children, which outlive their parent. |
CAP_INTERACTIVE | 6 | 0x40 | Acting with the logged-in user's authority: the Interactive role, unless CAP_SYS is also set. |
CAP_VSOCK | 7 | 0x80 | Creating and listening on native virtio-vsock streams. Needs CAP_NET as well. |
CAP_NET | 8 | 0x100 | Using sys-io's network API: TCP, UDP, ICMP, loopback, and vsock. |
CAP_FS_WRITE | 9 | 0x200 | Changing the filesystem and using file locks through sys-io. |
Two things these bits do not do:
CAP_SYSdoes not imply the other bits. A System process still needsCAP_SPAWNto create a child,CAP_SHUTDOWNto shut down,CAP_NETto use the network, andCAP_FS_WRITEto change files. Its broader authority applies only to granting capabilities, below.CAP_LOGdoes not open/system/logs. Reading log files is a matter of filesystem permissions;CAP_LOGonly allows sending records. A panic message or a backtrace goes to the process's stderr and needs no capability.
Roles
Two bits also decide the process's role, which is all the filesystem knows about who
is asking. CAP_SYS means System; otherwise
CAP_INTERACTIVE means Interactive; otherwise the role is
None. A None-role process can still hold individual capabilities such as
CAP_SPAWN or CAP_NET, and the Interactive role alone grants no
logging, spawning, or shutdown authority. The roles and the per-role file permissions are
described under Processes and roles and
Filesystem.
Granting capabilities to a child
The kernel checks the requested mask when it creates the child:
- A parent without
CAP_SYSmay grant only capabilities it holds, and neverCAP_SYSorCAP_IO_MANAGER. - A None-role parent may not grant
CAP_LOG, even if it holds it. An Interactive parent that holdsCAP_LOGmay pass it on explicitly. - A System parent may grant capabilities it does not hold, except
CAP_VSOCK,CAP_NET, andCAP_FS_WRITE. Every parent must hold these three to grant them, so they can only be narrowed on the way down the process tree. - Spawning a detached child needs
CAP_SPAWN_DETACHEDin the parent, even when the parent is System.
A request that breaks these rules fails with E_NOT_ALLOWED
(PermissionDenied in std). The kernel never quietly drops the
forbidden bits. Below System, a child can never get a higher role than its parent.
The default mask
Without an explicit mask, the runtime computes one from the parent's role with
default_child_capabilities:
| Parent's role | Default child mask |
|---|---|
| System | CAP_SPAWN and CAP_LOG, plus whichever of CAP_VSOCK, CAP_NET, and CAP_FS_WRITE the parent holds. The child has the None role. |
| Interactive | CAP_INTERACTIVE, plus whichever of CAP_SPAWN, CAP_VSOCK, CAP_NET, and CAP_FS_WRITE the parent holds. |
| None | Only the parent's CAP_SPAWN, CAP_VSOCK, CAP_NET, and CAP_FS_WRITE bits. |
CAP_SYS, CAP_IO_MANAGER, CAP_SHUTDOWN, and
CAP_SPAWN_DETACHED never pass on by default, and CAP_LOG does
only from a System parent. A System parent's default child has the None role even if the
parent also holds CAP_INTERACTIVE: a system service has to say explicitly
when it creates a user session. Defaults from non-System parents only ever contain bits
the parent holds.
An explicit mask
Setting MOTOR_OS_CAPS in a child's environment replaces the whole default.
The value is hexadecimal, with an optional lowercase 0x: 44 and
0x44 both mean CAP_SPAWN | CAP_INTERACTIVE, and 0
asks for nothing. A value that is not valid hexadecimal, or does not fit in 64 bits, fails
the spawn with InvalidInput rather than falling back to the default. The
parent's runtime consumes the variable, so the child never sees it.
use moto_sys::caps::{CAP_SPAWN, MOTOR_OS_CAPS_ENV_KEY};
use std::process::Command;
// A None-role worker that may spawn, but not write files or use the network.
let worker = Command::new(program)
.env(MOTOR_OS_CAPS_ENV_KEY, format!("{CAP_SPAWN:#x}"))
.spawn()?;
An explicit mask is a replacement, not an addition. Leaving out
CAP_INTERACTIVE drops the Interactive role; leaving out CAP_NET
or CAP_FS_WRITE denies that access even though the parent holds it. The
child's own children then use their own defaults or explicit masks.
Detached children
Setting MOTOR_OS_DETACHED to exactly true or TRUE
asks for a detached child; other values do not. The runtime consumes the variable either
way. A detached child is owned by the kernel and survives its parent; an ordinary
non-System child is killed when its parent is reaped. Detaching is separate from the
child's mask:
- the parent needs
CAP_SPAWN_DETACHED; - the child does not need that bit to be detached;
- giving the child
CAP_SPAWN_DETACHEDlets it detach its own children, but does not detach it.
A detached worker is best started with null standard streams, so that it does not depend on its launcher for I/O.
Network access
sys-io authorizes the process at the other end of each connection, using the mask the kernel reports for it; nothing a client sends can widen it. If sys-io cannot get that mask, it drops the connection rather than treat the client as unprivileged.
Without CAP_NET, sys-io drops a network connection as soon as it accepts
it, before serving any request. That covers TCP, UDP, ICMP, loopback, and vsock, including
vsock discovery. The IPC connect itself can succeed, so the refusal shows up as a closed
connection: native network calls fail with NotConnected, and a raw
io_channel client sees its channel close. CAP_VSOCK alone grants
nothing; vsock needs both bits.
The DNS resolver checks its own callers for CAP_NET and refuses the others
with PermissionDenied, so a process cannot borrow the resolver's network
access. Numeric addresses and localhost are answered in the runtime and need
no capability.
Filesystem writes
Without CAP_FS_WRITE, a process still connects to the filesystem and can
stat, read, and list what its role allows. Every other request fails with
PermissionDenied, and the connection stays usable. That includes creating,
writing, resizing, copying into, deleting, moving, and changing permissions of entries,
flushing the filesystem, and every file-lock operation, unlock included. Denying locks is
a provisional policy. The runtime refuses opens with write, append, create, or truncate
intent early, even of an existing file, but sys-io enforces the rule for any client that
bypasses the runtime.
CAP_FS_WRITE never overrides filesystem permissions: a change needs both
the capability and the role's permission. Without the capability, flush on a
read-only file does nothing, and flushing a writable file is refused.
The checks follow the process that talks to sys-io:
- A file handed to a child as a standard stream with
Stdio::from(file)is used through the child's own connection, so a child withoutCAP_FS_WRITEcannot write to it. - A child that inherits a file-backed standard stream writes through its parent's relay, with the parent's authority.
- Pipes carry no filesystem authority at all.
Rust's standard streams report a denied write as PermissionDenied, and
println! and eprintln! panic when a write fails. A closed
standard stream reads as end of file and discards writes without an error.
Reading a mask
A process reads its own mask without a syscall, from a read-only page the kernel maps into every process:
let caps = moto_sys::ProcessStaticPage::get().capabilities;
let role = moto_sys::caps::ProcessRole::from_caps(caps);
let can_spawn = caps & moto_sys::caps::CAP_SPAWN != 0;
For the process at the other end of a connection,
SysObj::get_capabilities(handle) returns that process's mask, and
SysObj::get_pid(handle) its PID. Both answers come from the kernel, so they
cannot be forged; servers should authorize with them, as sys-io, strobe, and the DNS
resolver do. Process statistics (ps, top) show only the derived
role, not the mask, and are for observation, not authorization.
Named IPC services
A server registers a service name by creating a listener for it, usually through
moto_ipc::sync::LocalServer. The kernel gives each name one owner process:
- The first process to register a name owns it; its later listeners share it.
- The owner keeps the name while it has any listening or connected endpoint for it,
and loses it when it closes the last one or exits. A running
LocalServernever gives its name up: when memory runs low it keeps serving its open connections and retries adding listeners every 100 ms. - While the owner is alive, another process cannot register the name. Once the owner has exited, another process can take it.
Only the sys-io name is reserved. Any other name belongs to whoever
registers it first, so a name alone does not prove who the server is. Check both ends
through the connection: a server checks each client's mask before serving it, and a client
that sends anything sensitive checks the server's mask first. Never trust a mask or a PID
that the other side sends as data.
Capabilities in rush
An Interactive shell's commands are Interactive by the default rule. A System console
shell explicitly grants CAP_SYS to the ordinary commands it runs, because the
default rule never passes System on. Programs listed by name under
spawn-detached in /user/cfg/rush.toml (the image lists
rmux) also get CAP_SPAWN_DETACHED, so they can leave a server
running after the session ends.
Giving a command an explicit mask
MOTOR_OS_CAPS works in rush both as a command assignment and as an exported
variable, and the assignment wins over the export. Either way it reaches the runtime
unchanged and suppresses both of rush's automatic grants, so
MOTOR_OS_CAPS=0x2ec rmux runs rmux without CAP_NET and without
the detach grant added on top. Assignments before command and
exec reach the program they run in the same way.
A mask needs a real child process. Rush refuses a masked function, eval,
. (source), any other builtin, a compound command, or a redirection-only
command with status 126, before running its body or opening its redirections, rather
than run it with the shell's own authority. command and exec can
forward a mask to an external program, but cannot make a builtin honor it. Rush only
checks that the variable is present; an empty, zero, or malformed value is refused the
same way.
| Invocation | What happens |
|---|---|
MOTOR_OS_CAPS=0x44 sh | Starts a shell with the requested capabilities. |
MOTOR_OS_CAPS=0x44 ./script.sh | Runs the script in a new, restricted rush process. |
MOTOR_OS_CAPS=0x44 eval '...' | Refused with status 126. |
MOTOR_OS_CAPS=0x44 . ./script.sh | Refused with status 126. |
MOTOR_OS_CAPS=0x44 function_name | Refused with status 126. |
An exported mask follows the same rule, for every builtin: while a mask is exported,
exit, return, break, continue,
wait, unset, trap, if, loops, and
other compounds are all refused with status 126. A refusal is an ordinary failure: without
set -e, a script continues past a refused exit, and a refused
wait does not wait. So prefer a command assignment. Rush's emulated
subshells do not isolate exported variables; use a separate shell process when an export
must not affect the caller. Bare assignments without redirections are ordinary shell
setup.
Refusal happens when rush reaches the command; it does not scan a script in advance.
Command-line expansions, and the redirections of a permitted child launch, still run
with the calling shell's authority: with MOTOR_OS_CAPS=0x44 prog > out,
the shell creates or truncates out, but prog's writes to it are
denied. To restrict a whole shell body, run it as
MOTOR_OS_CAPS=0x44 sh -c '...'. Inside that shell the mask is already
consumed, so its builtins, functions, and sourcing work, with the restricted
capabilities; unsetting MOTOR_OS_CAPS there cannot bring back what was
left out.
Rush stages pipelines and command substitutions through temporary files in
$TMPDIR, so they need CAP_FS_WRITE. If staging fails, the
pipeline or substitution returns status 2; it does not fall back to the shell's own
streams or pretend to succeed with empty output. For example,
MOTOR_OS_CAPS=0x44 sh -c 'x=$(printf hi)' fails with status 2.
Capabilities in rmux
An rmux server often has more capabilities than a program that might try to talk to it. On Motor OS the rmux client and server therefore meet over a named IPC service, not TCP, so each side can ask the kernel for the other's mask.
- The server's mask. The server is spawned with the default mask of
the process that starts it. Its service name is that mask in hex:
rmux/<mask>, orrmux/<mask>/<TMPDIR>whenTMPDIRis set. - Finding a server. A client computes the mask its own server would
get,
default_child_capabilities(own mask), and looks up that name. Before it sends anything, it checks that the server holding the name has exactly that mask, so a less privileged impostor cannot collect keystrokes. If a process with another mask holds the name, the client names that process and gives up. - Refusing clients. The server serves a client only if the client holds every bit the server holds, or would grant it to its own children by default. A client that fails the check is refused before any request is processed, so it cannot list, create, attach to, type into, or kill a session.
- Roles.
CAP_SYSandCAP_INTERACTIVEare part of the mask, so a None-role process never reaches an Interactive server, even if it holds every other bit. - Narrower servers. A client with a different mask finds or starts its own server, with its own sessions. A restricted program gets an equally restricted rmux, never a way into a more privileged one. The match is exact: a client never uses a narrower server, whose new panes would silently lose capabilities.
- Panes. A pane's program gets the server's default child mask, which for an Interactive server is the server's own mask.
- Starting a server needs
CAP_SPAWNandCAP_SPAWN_DETACHED, because the server is detached so that it outlives the session. A client without them fails with a message naming the missing capability. rush grants the detach bit to rmux through itsspawn-detachedlist; an explicitMOTOR_OS_CAPSmust include0x20itself.
How some launchers map to servers:
| Launched as | Client mask | Server and pane mask |
|---|---|---|
rmux from an SSH shell | 0x3e4 | 0x3c4 (Interactive) |
rmux attach typed in one of that server's panes | 0x3c4 | 0x3c4, the same server |
MOTOR_OS_CAPS=0x2ec rmux (no CAP_NET) | 0x2ec | 0x2c4, a separate server |
MOTOR_OS_CAPS=0x3a4 rmux (None role) | 0x3a4 | 0x384, a separate None-role server |
MOTOR_OS_CAPS=0x44 rmux (no detach) | 0x44 | none: it can join a 0x44 server, not start one |
One gap is accepted for now: for a System user the name changes inside a pane. A
System client with every bit set gets a 0x38c server, whose None-role panes
run as 0x384, so an rmux attach typed in such a pane looks for a
different server. The design is in src/bin/rmux/details.md, section 4.2.1.
Masks in the shipped image
/system/cfg/sys-init.cfg starts each service with an explicit mask, and
services get only what they need. russhd runs with 0x3fc, and each
authenticated SSH shell gets the Interactive part of it, 0x3ec. The DNS
resolver runs with 0x108 (CAP_LOG | CAP_NET), so it has the
network but cannot write files; strobe has the reverse. These launch chains pass
CAP_NET and CAP_FS_WRITE on only where they hold them. How a
session gets its role is described under
Processes and roles.