diff --git a/aci_edge_sandboxes/README.md b/aci_edge_sandboxes/README.md index 846e0e43..dbab1859 100644 --- a/aci_edge_sandboxes/README.md +++ b/aci_edge_sandboxes/README.md @@ -117,8 +117,9 @@ compatible kernel and static edge-agent initramfs, a prepared GPT image, the abs to `nvxhost.dll` (`libnvxhost.so` on Linux), and the independently approved SHA-256 of that library. The Rust crate does **not** build, download, or publish the private library. `NvxHostBackend::new` verifies the file digest and ABI version and requires the additive -`nvx_build_ramfs_launch_arguments` and `nvx_session_connect_verified` exports. A missing, -wrong-version, or wrong-digest library fails rather than falling back. +`nvx_plan_sandbox`, `nvx_build_ramfs_launch_arguments_with_plan`, and +`nvx_session_connect_verified` exports. A missing, wrong-version, or wrong-digest library fails +rather than falling back. The image is attached read-only in OpenVMM's distro block slot. The edge guest validates its GPT and p2+ ext4 layers and uses a RAM-backed tmpfs upper/work for its overlay; it @@ -165,19 +166,21 @@ and `guestBuildId`; stop reports `forced` and, after a failed graceful shutdown, `gracefulError`. A capture failure never fails stop or deprovision: both report it as `consoleError`. -This first backend supports provision/start/exec/stop/deprovision with shell commands or -argv and a fixed caller-provided image. Positive execution timeouts must be whole seconds, -matching the guest RPC's precision; finer-grained timeouts fail validation rather than -silently extending execution. It explicitly rejects host file mappings, network -configuration beyond deny-all, piped stdin, execution cancellation, custom working -directories and environments. Snapshot/restore, image selection, and richer guest -operations are not part of this profile. See +This backend supports provision/start/exec/stop/deprovision with shell commands or argv, a +fixed caller-provided image, and the host paths and network rules described below. Positive +execution timeouts must be whole seconds, matching the guest RPC's precision; finer-grained +timeouts fail validation rather than silently extending execution. It explicitly rejects piped +stdin, execution cancellation, custom working directories and environments. Snapshot/restore, +image selection, and richer guest operations are not part of this profile. See [`examples/nvxhost_lifecycle.rs`](examples/nvxhost_lifecycle.rs) for a run requiring `--openvmm`, `--kernel`, `--initrd`, `--image`, `--host-library`, `--host-sha256`, `--state-root`, `--hypervisor`, and a command after `--`; an optional `--image-sha256` -requires the image's digest when it is registered. The ignored -`tests/nvxhost_guest.rs` exercises an actual WHP guest when the corresponding -`NVXHOST_TEST_*` paths and approved DLL digest are set. +requires the image's digest when it is registered. The repeatable `--readonly`, `--readwrite`, +and `--denied` options map host paths, and `--egress allow|deny` with the repeatable +`--egress-allow` and `--egress-deny` options, each taking `CIDR` or `CIDR:tcp|udp:PORT`, attach a +network. The ignored `tests/nvxhost_guest.rs` exercises an actual guest when the corresponding +`NVXHOST_TEST_*` paths and approved library digest are set: under WHP on Windows, or under the +hypervisor that `NVXHOST_TEST_HYPERVISOR` names, such as `mshv` on Linux. For example, from `aci_edge_sandboxes` on a Windows WHP host, set the following paths to compatible, separately built artifacts and a caller-prepared GPT disk @@ -199,12 +202,56 @@ cargo run --release --locked --features nvxhost --example nvxhost_lifecycle -- ` --state-root $stateRoot --hypervisor whp -- 'printf READY' ``` +To map `C:\work\src` read-only and `C:\work\out` read-write, and let the guest reach only +`192.0.2.10` on TCP port 443, add +`--readonly C:\work\src --readwrite C:\work\out --egress deny --egress-allow 192.0.2.10:tcp:443` +before `--`. The example prints where each mapped path appears in the guest, here +`/mnt/c/work/src` and `/mnt/c/work/out`. + Build the native library and the static edge initramfs separately from their matching private sources; this example neither fetches nor builds them. Use the pinned OpenVMM, which includes the scratchless RAM-overlay topology, and a kernel compatible with that OpenVMM and guest revision. On Linux, supply a matching -`libnvxhost.so` and OpenVMM build and select `--hypervisor mshv`; the Linux edge -lifecycle has not yet been verified end to end. +`libnvxhost.so` and OpenVMM build and select `--hypervisor mshv`. + +### Native host paths and network + +The native backend accepts the same `filesystem` and `network` policies as the direct +backend: it maps host paths to the same guest paths under the same rules and expands network +rules the same way; see [Host paths](#host-paths) and [Network rules](#network-rules). It also +refuses mappings that would show workloads its state root, described below. Its workloads run +as the guest's root, so it ignores `map_host_identity`. The private library plans the +policies. Provision passes them to `nvx_plan_sandbox`, which resolves the mapped and denied +paths, chooses the export, pins the objects inside read-write mappings, and expands the +network rules; the sandbox record keeps the resulting plan. Start passes the plan back, and +the library checks it again against every planning rule that needs no host access, and checks +the pinned objects: if one was replaced, start fails with `backend_error`, and the sandbox +must be deprovisioned and provisioned again to accept the change. + +- The state root holds the record, and so the plan, of every sandbox, which decides what the + next start exports and which egress it allows. Provision therefore refuses, with + `policy_validation`, a mapped path inside the state root, and one that contains it unless a + denied path inside the mapping hides it. +- OpenVMM exports the deepest directory that contains every mapped path through one + virtio-fs device, read-only unless a path is read-write, and hides the denied paths. The + guest mounts the export where workloads cannot reach it and bind-mounts each mapped path + into the RAM overlay at its guest path, read-only where requested. The bind mounts travel on + the guest's 1024-byte kernel command line, which leaves room for roughly a dozen typical + paths; provision rejects a policy whose mounts do not fit with `policy_validation`. +- Workloads run as the guest's root with only the default container capabilities (`CHOWN`, + `DAC_OVERRIDE`, `FOWNER`, `FSETID`, `KILL`, `SETGID`, `SETUID`, `NET_BIND_SERVICE`, and + `AUDIT_WRITE`), with `no_new_privs`, and without user namespaces, so they cannot remount or + unmount the mapped paths or mount the export again. Host-side changes through read-write + mappings, including modes and ownership, happen with the credentials of the account that + runs OpenVMM, so run OpenVMM unprivileged. A host hard link that already joins a file in a + read-write mapping to one in a read-only mapping stays writable through the read-write + path. +- A policy that denies egress without allow rules attaches no network device. Otherwise the + guest gets `OpenVmmConfig::guest_network` (`10.0.0.2/24` by default) behind OpenVMM's NAT + gateway, the network's first address, and names that gateway as its DNS server in + `/etc/resolv.conf` when the policy allows TCP or UDP port 53 to it. Choose a guest network + that contains no address the guest must reach. Ingress and host-loopback access must be + `deny`. The caller must independently approve and protect the native asset. Checking a caller-supplied digest does not make a writable path or a self-declared digest trustworthy; use this profile @@ -284,7 +331,9 @@ Defaults: | `stop_timeout`, `exec_response_grace` | 30 s | Each timeout must be at most 30 days (`OpenVmmConfig::MAX_TIMEOUT`), and all -but `exec_response_grace` must be positive. +but `exec_response_grace` must be positive. `guest_network` must be an address +that OpenVMM accepts: a /1 to /30 prefix, and neither the network's own address, +its broadcast address, nor its first address, which is the gateway. Choose a `state_root` that only the current user can access. `OpenVmmConfig::default_state_root` returns `%LOCALAPPDATA%\nvx\sandboxes` on diff --git a/aci_edge_sandboxes/examples/nvxhost_lifecycle.rs b/aci_edge_sandboxes/examples/nvxhost_lifecycle.rs index d24c7bea..574c7d5e 100644 --- a/aci_edge_sandboxes/examples/nvxhost_lifecycle.rs +++ b/aci_edge_sandboxes/examples/nvxhost_lifecycle.rs @@ -6,9 +6,12 @@ use std::process::ExitCode; use std::sync::Arc; use aci_edge_sandboxes::openvmm::{ - Hypervisor, ImageDigest, NvxHostBackend, NvxHostConfig, OpenVmmConfig, + Hypervisor, ImageDigest, NvxHostBackend, NvxHostConfig, OpenVmmConfig, resolve_guest_path, +}; +use aci_edge_sandboxes::{ + Access, AciEdgeSandbox, EgressPolicy, ExecOutcome, ExecRequest, FilesystemPolicy, + NetworkPolicy, NetworkRule, Protocol, ProvisionRequest, }; -use aci_edge_sandboxes::{AciEdgeSandbox, ExecOutcome, ExecRequest, ProvisionRequest}; fn main() -> ExitCode { match run() { @@ -30,6 +33,10 @@ fn run() -> Result<(), String> { let mut image_digest = ImageDigest::Compute; let mut root = None; let mut hypervisor = None; + let mut filesystem = FilesystemPolicy::default(); + let mut egress = None; + let mut allow = Vec::new(); + let mut deny = Vec::new(); let mut command = None; let mut args = std::env::args().skip(1); while let Some(option) = args.next() { @@ -49,6 +56,12 @@ fn run() -> Result<(), String> { } "--state-root" => root = Some(value()?), "--hypervisor" => hypervisor = Some(value()?.parse::().map_err(describe)?), + "--readonly" => filesystem.readonly_paths.push(value()?.into()), + "--readwrite" => filesystem.readwrite_paths.push(value()?.into()), + "--denied" => filesystem.denied_paths.push(value()?.into()), + "--egress" => egress = Some(parse_access(&option, &value()?)?), + "--egress-allow" => allow.push(parse_rule(&option, &value()?)?), + "--egress-deny" => deny.push(parse_rule(&option, &value()?)?), "--" => { command = Some(args.by_ref().collect::>().join(" ")); break; @@ -69,6 +82,33 @@ fn run() -> Result<(), String> { let command = command .filter(|command| !command.is_empty()) .ok_or("pass the guest command after --")?; + let mut request = ProvisionRequest::new(); + if !filesystem.is_empty() { + for path in filesystem + .readonly_paths + .iter() + .chain(&filesystem.readwrite_paths) + { + if let Some(guest) = resolve_guest_path(path) { + eprintln!("{} is {guest} in the guest", path.display()); + } + } + request = request.with_filesystem(filesystem); + } + match egress { + Some(default) => { + request = request.with_network(NetworkPolicy { + egress: EgressPolicy { + default, + allow, + deny, + }, + ..NetworkPolicy::deny_all() + }); + } + None if allow.is_empty() && deny.is_empty() => {} + None => return Err("--egress-allow and --egress-deny require --egress".to_owned()), + } let config = OpenVmmConfig::new(openvmm, kernel, initrd, hypervisor, PathBuf::from(root)); let backend = Arc::new( @@ -78,10 +118,7 @@ fn run() -> Result<(), String> { .map_err(describe)?, ); let client = AciEdgeSandbox::from_shared(backend.clone()); - let id = client - .provision(&ProvisionRequest::new()) - .map_err(describe)? - .sandbox_id; + let id = client.provision(&request).map_err(describe)?.sandbox_id; let started = client.start(&id); let succeeded = started.is_ok(); let executed = started.and_then(|_| { @@ -155,6 +192,38 @@ fn parse_digest(option: &str, hex: &str) -> Result<[u8; 32], String> { Ok(digest) } +fn parse_access(option: &str, text: &str) -> Result { + match text { + "allow" => Ok(Access::Allow), + "deny" => Ok(Access::Deny), + _ => Err(format!("{option} must be allow or deny")), + } +} + +/// Parses `CIDR` or `CIDR:tcp|udp:PORT`, such as `192.0.2.0/24` or `192.0.2.1:tcp:443`. +fn parse_rule(option: &str, text: &str) -> Result { + let malformed = || format!("{option} must be CIDR or CIDR:tcp|udp:PORT"); + let mut parts = text.split(':'); + let rule = NetworkRule::to( + parts + .next() + .filter(|cidr| !cidr.is_empty()) + .ok_or_else(malformed)?, + ); + match (parts.next(), parts.next(), parts.next()) { + (None, ..) => Ok(rule), + (Some(protocol), Some(port), None) => { + let protocol = match protocol { + "tcp" => Protocol::Tcp, + "udp" => Protocol::Udp, + _ => return Err(malformed()), + }; + Ok(rule.on_port(protocol, port.parse().map_err(|_| malformed())?)) + } + _ => Err(malformed()), + } +} + fn describe(error: aci_edge_sandboxes::Error) -> String { error.to_string() } @@ -171,4 +240,28 @@ mod tests { assert!(parse(&"é".repeat(32)).is_err()); assert!(parse(&"gg".repeat(32)).is_err()); } + + #[test] + fn egress_rules_name_a_network_and_optionally_one_port() { + let parse = |text: &str| parse_rule("--egress-allow", text); + assert_eq!( + parse("192.0.2.0/24").unwrap(), + NetworkRule::to("192.0.2.0/24") + ); + assert_eq!( + parse("192.0.2.1:udp:53").unwrap(), + NetworkRule::to("192.0.2.1").on_port(Protocol::Udp, 53) + ); + for malformed in [ + "", + ":tcp:1", + "192.0.2.1:tcp", + "192.0.2.1:icmp:1", + "192.0.2.1:tcp:x", + ] { + assert!(parse(malformed).is_err(), "{malformed}"); + } + assert_eq!(parse_access("--egress", "deny").unwrap(), Access::Deny); + assert!(parse_access("--egress", "Deny").is_err()); + } } diff --git a/aci_edge_sandboxes/src/nvxhost.rs b/aci_edge_sandboxes/src/nvxhost.rs index e3c0593d..8c74091a 100644 --- a/aci_edge_sandboxes/src/nvxhost.rs +++ b/aci_edge_sandboxes/src/nvxhost.rs @@ -29,6 +29,10 @@ const COMPLETION_SEND: u32 = 2; const COMPLETION_RECV: u32 = 3; const COMPLETION_DISPOSE: u32 = 6; const ERROR_CONNECT: u32 = 12; +const ERROR_ARGUMENT: u32 = 10; +const ERROR_ARGUMENT_NULL: u32 = 14; +const ERROR_NOT_SUPPORTED: u32 = 15; +const ERROR_HOST_CHANGED: u32 = 18; const CALL_UNARY: u32 = 0; const CALL_SERVER_STREAM: u32 = 1; const FRAME_RESPONSE: u8 = 2; @@ -37,6 +41,7 @@ const FLAG_REMOTE_CLOSED: u8 = 1; const FLAG_NO_DATA: u8 = 4; const LAUNCH_GUEST_DEBUG: u32 = 1; const LAUNCH_HAS_MEMORY: u32 = 4; +const PLAN_VALIDATE_ONLY: u32 = 1; #[repr(C)] #[derive(Clone, Copy, Default)] @@ -139,8 +144,20 @@ type CallSendRequest = unsafe extern "C" fn(u64, *const u8, usize, u8, i64, u32, type CallRecv = unsafe extern "C" fn(u64, u64) -> i32; type CallTryComplete = unsafe extern "C" fn(u64) -> i32; type CallRelease = unsafe extern "C" fn(u64) -> i32; -type BuildRamfsArguments = unsafe extern "C" fn( +type BuildRamfsArgumentsWithPlan = unsafe extern "C" fn( *const NvxLaunchRequest, + *const u8, + usize, + *mut *mut u8, + *mut usize, + *mut *mut c_void, +) -> i32; +type PlanSandbox = unsafe extern "C" fn( + *const u8, + usize, + *const u8, + usize, + u32, *mut *mut u8, *mut usize, *mut *mut c_void, @@ -164,7 +181,8 @@ struct Exports { call_recv: CallRecv, call_try_complete: CallTryComplete, call_release: CallRelease, - build_ramfs_arguments: BuildRamfsArguments, + build_ramfs_arguments_with_plan: BuildRamfsArgumentsWithPlan, + plan_sandbox: PlanSandbox, buffer_free: BufferFree, } @@ -244,7 +262,11 @@ impl HostLibrary { call_recv: resolve(&library, b"nvx_call_recv\0")?, call_try_complete: resolve(&library, b"nvx_call_try_complete\0")?, call_release: resolve(&library, b"nvx_call_release\0")?, - build_ramfs_arguments: resolve(&library, b"nvx_build_ramfs_launch_arguments\0")?, + build_ramfs_arguments_with_plan: resolve( + &library, + b"nvx_build_ramfs_launch_arguments_with_plan\0", + )?, + plan_sandbox: resolve(&library, b"nvx_plan_sandbox\0")?, buffer_free: resolve(&library, b"nvx_buffer_free\0")?, }; Ok(Arc::new(Self { @@ -253,6 +275,62 @@ impl HostLibrary { })) } + /// Plans the host paths and network of `policy`, the JSON of a request's `filesystem` and + /// `network` sections, for a guest at `guest_network`. + /// + /// Returns the plan's JSON, which the caller stores and passes back at launch, or `None` + /// when `validate_only`, which checks the policy without touching the host. Policies that + /// the library cannot enforce fail with `policy_validation`. + pub(crate) fn plan_sandbox( + &self, + policy: &str, + guest_network: &str, + validate_only: bool, + ) -> Result> { + let mut buffer = ptr::null_mut(); + let mut length = 0; + let mut error = ptr::null_mut(); + let status = unsafe { + (self.exports.plan_sandbox)( + policy.as_ptr(), + policy.len(), + guest_network.as_ptr(), + guest_network.len(), + if validate_only { PLAN_VALIDATE_ONLY } else { 0 }, + &mut buffer, + &mut length, + &mut error, + ) + }; + if status != STATUS_OK || !error.is_null() { + let failure = self.failure(error); + if !error.is_null() { + unsafe { (self.exports.error_free)(error) }; + } + return Err( + if matches!( + failure.kind, + ERROR_ARGUMENT | ERROR_ARGUMENT_NULL | ERROR_NOT_SUPPORTED + ) { + Error::policy_validation(failure.message) + } else { + failure.as_error("planning host paths and network", status) + }, + ); + } + if validate_only { + return Ok(None); + } + if buffer.is_null() || length == 0 { + return Err(Error::backend_error("nvxhost returned no sandbox plan")); + } + let plan = unsafe { slice::from_raw_parts(buffer, length).to_vec() }; + unsafe { (self.exports.buffer_free)(buffer, length) }; + String::from_utf8(plan).map(Some).map_err(|error| { + Error::backend_error("nvxhost returned a non-UTF-8 sandbox plan").with_source(error) + }) + } + pub(crate) fn launch_arguments(&self, inputs: &LaunchInputs<'_>) -> Result> { fn path_text(path: &Path) -> Result<&str> { path.to_str().ok_or_else(|| { @@ -287,10 +365,31 @@ impl HostLibrary { let mut buffer = ptr::null_mut(); let mut length = 0; let mut error = ptr::null_mut(); + let plan = inputs.plan.unwrap_or_default(); let status = unsafe { - (self.exports.build_ramfs_arguments)(&request, &mut buffer, &mut length, &mut error) + (self.exports.build_ramfs_arguments_with_plan)( + &request, + plan.as_ptr(), + plan.len(), + &mut buffer, + &mut length, + &mut error, + ) }; - self.check_status("building OpenVMM arguments", status, error)?; + if status != STATUS_OK || !error.is_null() { + let failure = self.failure(error); + if !error.is_null() { + unsafe { (self.exports.error_free)(error) }; + } + return Err(if failure.kind == ERROR_HOST_CHANGED { + Error::backend_error(format!( + "{}; deprovision the sandbox and provision it again", + failure.message + )) + } else { + failure.as_error("building OpenVMM arguments", status) + }); + } if buffer.is_null() || length == 0 { return Err(Error::backend_error( "nvxhost returned no OpenVMM launch arguments", @@ -377,6 +476,8 @@ pub(crate) struct LaunchInputs<'a> { pub(crate) hypervisor: &'a str, pub(crate) memory_mb: u32, pub(crate) guest_debug: bool, + /// JSON of the sandbox plan from [`HostLibrary::plan_sandbox`], if the sandbox has one. + pub(crate) plan: Option<&'a str>, } #[derive(Debug)] @@ -881,18 +982,18 @@ mod tests { .message() .contains("does not match the approved SHA-256") ); - let args = host - .launch_arguments(&LaunchInputs { - kernel: Path::new(r"C:\test\vmlinux"), - initrd: Path::new(r"C:\test\edge-initramfs.cpio.gz"), - image: Path::new(r"C:\test\image.gpt"), - control: "//./pipe/openvmm-microvm-edge-control", - boot: "//./pipe/openvmm-microvm-edge-boot", - hypervisor: "whp", - memory_mb: 256, - guest_debug: false, - }) - .unwrap(); + let inputs = |plan| LaunchInputs { + kernel: Path::new(r"C:\test\vmlinux"), + initrd: Path::new(r"C:\test\edge-initramfs.cpio.gz"), + image: Path::new(r"C:\test\image.gpt"), + control: "//./pipe/openvmm-microvm-edge-control", + boot: "//./pipe/openvmm-microvm-edge-boot", + hypervisor: "whp", + memory_mb: 256, + guest_debug: false, + plan, + }; + let args = host.launch_arguments(&inputs(None)).unwrap(); let args: Vec<_> = args .iter() .map(|argument| argument.to_str().unwrap()) @@ -905,5 +1006,37 @@ mod tests { .count(), 1 ); + assert!(!args.contains(&"--mount") && !args.contains(&"--net")); + + // Validation consults nothing on the host; unenforceable policies are policy errors. + assert_eq!(host.plan_sandbox("{}", "10.0.0.2/24", true).unwrap(), None); + let ipv6 = r#"{"network":{"egress":{"default":"allow","deny":[{"to":[{"cidr":"::/0"}]}]},"ingress":{"default":"deny"}}}"#; + assert_eq!( + host.plan_sandbox(ipv6, "10.0.0.2/24", true) + .unwrap_err() + .code(), + crate::ErrorCode::PolicyValidation + ); + // A plan maps a host directory and attaches a network device at launch. + let directory = tempfile::tempdir().unwrap(); + let policy = serde_json::json!({ + "filesystem": { "readonlyPaths": [directory.path()] }, + "network": { "egress": { "default": "allow" }, "ingress": { "default": "deny" } }, + }) + .to_string(); + let plan = host + .plan_sandbox(&policy, "10.0.0.2/24", false) + .unwrap() + .unwrap(); + let args = host.launch_arguments(&inputs(Some(&plan))).unwrap(); + let args: Vec<_> = args + .iter() + .map(|argument| argument.to_str().unwrap()) + .collect(); + assert!(args.contains(&"--mount") && args.contains(&"--net")); + assert!( + args.iter() + .any(|argument| argument.contains("nvx_overlay_upper=ramfs nvx_map=.,")) + ); } } diff --git a/aci_edge_sandboxes/src/openvmm/config.rs b/aci_edge_sandboxes/src/openvmm/config.rs index 61536dc8..805e20dc 100644 --- a/aci_edge_sandboxes/src/openvmm/config.rs +++ b/aci_edge_sandboxes/src/openvmm/config.rs @@ -316,7 +316,8 @@ impl OpenVmmConfig { } if !valid_guest_network(&self.guest_network) { return invalid(format!( - "guest_network {:?} must be an IPv4 address with a /1 to /30 prefix", + "guest_network {:?} must be an IPv4 address with a /1 to /30 prefix that is not \ + its network's network, broadcast, or gateway (first) address", self.guest_network )); } @@ -383,15 +384,25 @@ fn valid_hostname(hostname: &str) -> bool { && bytes.last() != Some(&b'-') } +/// Applies OpenVMM's rules for a static guest address: a /1 to /30 prefix, and an address other +/// than the network's own, its broadcast address, and the gateway, which is its first address. fn valid_guest_network(value: &str) -> bool { - let Some((address, prefix)) = value.split_once('/') else { + let Some((address, prefix_text)) = value.split_once('/') else { return false; }; - address.parse::().is_ok() - && prefix - .parse::() - .is_ok_and(|prefix| (1..=30).contains(&prefix)) - && !prefix.starts_with('0') + let (Ok(address), Ok(prefix)) = ( + address.parse::(), + prefix_text.parse::(), + ) else { + return false; + }; + if prefix_text.starts_with('0') || !(1..=30).contains(&prefix) { + return false; + } + let mask = u32::MAX << (32 - u32::from(prefix)); + let host = u32::from(address); + let network = host & mask; + ![network, network | !mask, network + 1].contains(&host) } /// Describes why extra kernel parameters are unacceptable, if they are. @@ -495,8 +506,15 @@ mod tests { assert!(!valid_hostname("-nvx")); assert!(!valid_hostname(&"a".repeat(64))); assert!(valid_guest_network("10.0.0.2/24")); + assert!(valid_guest_network("10.0.0.254/24")); assert!(!valid_guest_network("10.0.0.2")); assert!(!valid_guest_network("10.0.0.256/24")); + assert!(!valid_guest_network("10.0.0.2/024")); + // OpenVMM rejects the network, broadcast, and gateway addresses as guest addresses. + for reserved in ["10.0.0.0/24", "10.0.0.255/24", "10.0.0.1/24", "10.0.0.5/30"] { + assert!(!valid_guest_network(reserved), "{reserved}"); + } + assert!(valid_guest_network("10.0.0.6/30")); assert!(kernel_command_line_problem("quiet loglevel=0").is_none()); assert!(kernel_command_line_problem("tsc=reliable").is_some()); assert!(kernel_command_line_problem("hostname=other").is_some()); diff --git a/aci_edge_sandboxes/src/openvmm/filesystem.rs b/aci_edge_sandboxes/src/openvmm/filesystem.rs index 0e8f9359..93681af6 100644 --- a/aci_edge_sandboxes/src/openvmm/filesystem.rs +++ b/aci_edge_sandboxes/src/openvmm/filesystem.rs @@ -331,6 +331,37 @@ pub(crate) fn plan(policy: &FilesystemPolicy) -> Result> { })) } +/// Returns the first mapped path of `policy` that would show workloads the host directory +/// `private`, such as the sandbox state root: a mapped path inside it, or one that contains it +/// without a denied path inside the mapping that hides it. Planning rejects paths that cannot be +/// resolved, so they are skipped here. +#[cfg(feature = "nvxhost")] +pub(crate) fn exposes(private: &Path, policy: &FilesystemPolicy) -> Option { + let private = canonicalize(private).unwrap_or_else(|_| private.to_path_buf()); + let hidden: Vec = policy + .denied_paths + .iter() + .filter_map(|path| canonicalize_lenient(path).ok().map(|(path, _)| path)) + .collect(); + policy + .readonly_paths + .iter() + .chain(&policy.readwrite_paths) + .find(|path| { + let Ok(mapped) = canonicalize(path) else { + return false; + }; + mapped.starts_with(&private) + || private.starts_with(&mapped) + && !hidden.iter().any(|denied| { + private.starts_with(denied) + && denied.starts_with(&mapped) + && *denied != mapped + }) + }) + .cloned() +} + /// Checks that every mapped and denied path inside a read-write mapping still names the object /// it named at provision. /// @@ -875,6 +906,54 @@ mod tests { verify(&mapping).unwrap(); } + #[test] + #[cfg(feature = "nvxhost")] + fn mappings_must_not_expose_a_private_directory() { + let directory = root(); + let base = directory.path(); + // Pretend that work/out holds the sandbox state. + let private = base.join("work/out"); + let policy = |readonly: &[&str], denied: &[&str]| FilesystemPolicy { + readonly_paths: readonly.iter().map(|path| base.join(path)).collect(), + readwrite_paths: Vec::new(), + denied_paths: denied.iter().map(|path| base.join(path)).collect(), + }; + assert_eq!( + exposes(&private, &policy(&["tools", "work/src"], &[])), + None + ); + assert_eq!( + exposes(&private, &policy(&["tools", "work"], &[])), + Some(base.join("work")) + ); + // A denied path inside the mapping hides it; one outside the mapping does not. + assert_eq!(exposes(&private, &policy(&["work"], &["work/out"])), None); + assert_eq!( + exposes(&private, &policy(&["work"], &["."])), + Some(base.join("work")) + ); + assert_eq!( + exposes(&private, &policy(&["work/out"], &["work/out"])), + Some(base.join("work/out")) + ); + fs::create_dir_all(base.join("work/out/data")).unwrap(); + let inside = FilesystemPolicy { + readwrite_paths: vec![base.join("work/out/data")], + ..FilesystemPolicy::default() + }; + assert_eq!(exposes(&private, &inside), Some(base.join("work/out/data"))); + // Another spelling of the private directory is recognized. + let spelled = if cfg!(windows) { + base.join("WORK").join("OUT") + } else { + base.join("work").join(".").join("out") + }; + assert_eq!( + exposes(&spelled, &policy(&["work"], &[])), + Some(base.join("work")) + ); + } + #[test] fn encoding_keeps_tokens_free_of_separators() { assert_eq!(encode("/mnt/c/My Work,1%"), "/mnt/c/My%20Work%2C1%25"); diff --git a/aci_edge_sandboxes/src/openvmm/native.rs b/aci_edge_sandboxes/src/openvmm/native.rs index b01c0c40..e57643e4 100644 --- a/aci_edge_sandboxes/src/openvmm/native.rs +++ b/aci_edge_sandboxes/src/openvmm/native.rs @@ -12,6 +12,7 @@ use prost::Message; use super::artifacts::absolute; use super::config::{OpenVmmConfig, validate_unix_socket_path}; +use super::filesystem; use super::images::{ ImageDigest, ImageId, ImageRecord, ImageStore, RegisteredImage, VerifiedFile, encode_hex, }; @@ -29,13 +30,13 @@ use crate::error::{Error, ErrorCode, Result}; use crate::exec::{Completion, ExecOutcome}; use crate::id::SandboxId; use crate::model::{ - Access, Command, DeprovisionResult, ExecRequest, Metadata, ProvisionRequest, ProvisionResult, - StartResult, StdinMode, StopResult, + Command, DeprovisionResult, ExecRequest, FilesystemPolicy, Metadata, NetworkPolicy, + ProvisionRequest, ProvisionResult, StartResult, StdinMode, StopResult, }; use crate::nvxhost::{HostLibrary, LaunchInputs, Session}; const BACKEND_KEY: &str = NATIVE_BACKEND_KEY; -const RUNTIME_ABI: &str = "microvm-abi-v2-edge-ramfs-v1"; +const RUNTIME_ABI: &str = "microvm-abi-v2-edge-ramfs-v2"; const MAX_EXEC_SECONDS: u64 = 3600; const MAX_OUTPUT_BYTES: usize = 1 << 20; const SHELL: &str = "/bin/sh"; @@ -44,9 +45,9 @@ const CONSOLE_POLL: Duration = Duration::from_millis(250); /// Artifact paths and approved native-library digest for an image-backed guest. /// /// The image is a caller-prepared GPT disk, not a container image reference. This backend -/// creates no disks, filesystem mappings, or network devices. Construct it with -/// [`NvxHostConfig::new`] and its builder methods, which keep callers compatible as options -/// are added. +/// creates no disks; it maps host paths and attaches a network device as each sandbox's policy +/// requests. Construct it with [`NvxHostConfig::new`] and its builder methods, which keep +/// callers compatible as options are added. #[derive(Debug, Clone)] #[non_exhaustive] pub struct NvxHostConfig { @@ -588,7 +589,7 @@ impl NvxHostBackend { Ok(()) } - fn artifact_record(&self) -> NativeArtifactRecord { + fn artifact_record(&self, devices: Option) -> NativeArtifactRecord { let digests = self.runtime.digests(); NativeArtifactRecord { image: self.image.to_string(), @@ -596,6 +597,7 @@ impl NvxHostBackend { kernel_sha256: encode_hex(&digests.kernel), initrd_sha256: encode_hex(&digests.initrd), library_sha256: encode_hex(&self.config.library_sha256), + devices, } } @@ -785,21 +787,24 @@ impl Backend for NvxHostBackend { "microvm.provision.memoryMib must fit a positive signed 32-bit integer", )); } - if request + if let Some(policy) = device_policy(request)? { + self.host + .plan_sandbox(&policy, &self.config.openvmm.guest_network, true)?; + } + // The state root holds every sandbox's record, including the plan that decides what the + // next start exports, so no workload may see it. + let state_root = &self.config.openvmm.state_root; + if let Some(path) = request .filesystem .as_ref() - .is_some_and(|policy| !policy.is_empty()) - || request.network.as_ref().is_some_and(|policy| { - policy.egress.default != Access::Deny - || policy.ingress.default != Access::Deny - || policy.ingress.host_loopback == Some(Access::Allow) - || !policy.egress.allow.is_empty() - || !policy.egress.deny.is_empty() - }) + .and_then(|policy| filesystem::exposes(state_root, policy)) { - return Err(Error::policy_validation( - "the nvxhost backend currently supports no host filesystem or guest network", - )); + return Err(Error::policy_validation(format!( + "the mapped path {} would show workloads the sandbox state in {}; map paths \ + outside the state root, or deny the state root inside the mapped path", + path.display(), + state_root.display() + ))); } Ok(()) } @@ -811,6 +816,19 @@ impl Backend for NvxHostBackend { fn provision(&self, request: &ProvisionRequest) -> Result { self.validate_provision(request)?; self.base.probe()?; + let devices = match device_policy(request)? { + Some(policy) => { + let plan = self + .host + .plan_sandbox(&policy, &self.config.openvmm.guest_network, false)? + .ok_or_else(|| Error::backend_error("nvxhost returned no sandbox plan"))?; + Some(serde_json::from_str(&plan).map_err(|error| { + Error::backend_error("nvxhost returned a malformed sandbox plan") + .with_source(error) + })?) + } + None => None, + }; // Unregistering an image waits for this lock, so the image stays registered until the // sandbox that refers to it is recorded. let _registry = self.images.lock()?; @@ -820,7 +838,7 @@ impl Backend for NvxHostBackend { backend: BACKEND_KEY.to_owned(), network: None, filesystem: None, - native: Some(self.artifact_record()), + native: Some(self.artifact_record(devices)), memory_mib: request .microvm .provision @@ -867,6 +885,15 @@ impl Backend for NvxHostBackend { Error::backend_error("cannot choose a control endpoint").with_source(error) })?; let boot = self.boot_endpoint(sandbox_id, &endpoint)?; + let plan = record + .native + .as_ref() + .and_then(|native| native.devices.as_ref()) + .map(serde_json::to_string) + .transpose() + .map_err(|error| { + Error::backend_error("cannot encode the sandbox plan").with_source(error) + })?; let mut arguments = self.host.launch_arguments(&LaunchInputs { kernel: &self.config.openvmm.kernel, initrd: &self.config.openvmm.initrd, @@ -876,6 +903,7 @@ impl Backend for NvxHostBackend { hypervisor: self.config.openvmm.hypervisor.as_str(), memory_mb: record.memory_mib, guest_debug: self.config.guest_debug, + plan: plan.as_deref(), })?; let report = self.base.store.outcome_path(sandbox_id); remove_if_present(&report)?; @@ -1129,12 +1157,40 @@ fn capabilities() -> Capabilities { capabilities.exec.argv = true; capabilities.exec.max_timeout_ms = Some(MAX_EXEC_SECONDS * 1_000); capabilities.exec.max_output_bytes = Some(MAX_OUTPUT_BYTES as u64); + capabilities.network.egress_allow = true; capabilities.network.egress_deny = true; capabilities.network.ingress_deny = true; capabilities.network.host_loopback_deny = true; + capabilities.network.egress_rules = true; + capabilities.filesystem.readonly_paths = true; + capabilities.filesystem.readwrite_paths = true; + capabilities.filesystem.denied_paths = true; capabilities } +/// The JSON of `request`'s `filesystem` and `network` sections, which nvxhost plans, or `None` +/// when the request has neither. +fn device_policy(request: &ProvisionRequest) -> Result> { + #[derive(serde::Serialize)] + struct Policy<'a> { + #[serde(skip_serializing_if = "Option::is_none")] + filesystem: Option<&'a FilesystemPolicy>, + #[serde(skip_serializing_if = "Option::is_none")] + network: Option<&'a NetworkPolicy>, + } + if request.filesystem.is_none() && request.network.is_none() { + return Ok(None); + } + serde_json::to_string(&Policy { + filesystem: request.filesystem.as_ref(), + network: request.network.as_ref(), + }) + .map(Some) + .map_err(|error| { + Error::policy_validation("filesystem paths must be valid UTF-8").with_source(error) + }) +} + fn finish_session(session: Session, result: Result, deadline: Option) -> Result { match (result, session.close(deadline)) { (Ok(value), Ok(())) => Ok(value), @@ -1500,7 +1556,51 @@ mod tests { assert!(request.encode_to_vec().len() < 100); assert!(capabilities().exec.command_line); assert!(!capabilities().exec.cancel); - assert!(!capabilities().filesystem.readonly_paths); + assert!(!capabilities().exec.cwd); + } + + #[test] + fn host_paths_and_network_follow_the_policy_the_library_plans() { + let capabilities = capabilities(); + assert!(capabilities.filesystem.readonly_paths); + assert!(capabilities.filesystem.readwrite_paths); + assert!(capabilities.filesystem.denied_paths); + assert!(capabilities.network.egress_allow && capabilities.network.egress_rules); + assert!(!capabilities.network.ingress_allow && !capabilities.network.host_loopback_allow); + + assert_eq!(device_policy(&ProvisionRequest::default()).unwrap(), None); + // The JSON is nvxhost's planning input; its own tests parse this exact text. + let request = ProvisionRequest { + filesystem: Some(FilesystemPolicy { + readonly_paths: vec!["/work/src".into()], + readwrite_paths: vec!["/work/out".into()], + denied_paths: vec!["/work/src/secret".into()], + }), + network: Some(NetworkPolicy { + egress: crate::model::EgressPolicy::new(crate::model::Access::Deny) + .with_allow(crate::model::NetworkRule { + to: vec![crate::model::NetworkPeer { + cidr: "192.0.2.0/24".into(), + except: vec!["192.0.2.128/25".into()], + }], + ports: vec![crate::model::NetworkPort { + protocol: crate::model::Protocol::Tcp, + port: Some(443), + end_port: Some(444), + }], + }) + .with_deny(crate::model::NetworkRule::to("192.0.2.7")), + ingress: crate::model::IngressPolicy { + default: crate::model::Access::Deny, + host_loopback: Some(crate::model::Access::Deny), + }, + }), + ..ProvisionRequest::default() + }; + assert_eq!( + device_policy(&request).unwrap().unwrap(), + r#"{"filesystem":{"readonlyPaths":["/work/src"],"readwritePaths":["/work/out"],"deniedPaths":["/work/src/secret"]},"network":{"egress":{"default":"deny","allow":[{"to":[{"cidr":"192.0.2.0/24","except":["192.0.2.128/25"]}],"ports":[{"protocol":"tcp","port":443,"endPort":444}]}],"deny":[{"to":[{"cidr":"192.0.2.7"}]}]},"ingress":{"default":"deny","hostLoopback":"deny"}}}"# + ); } #[test] diff --git a/aci_edge_sandboxes/src/openvmm/platform/linux.rs b/aci_edge_sandboxes/src/openvmm/platform/linux.rs index 522892fb..7093cee0 100644 --- a/aci_edge_sandboxes/src/openvmm/platform/linux.rs +++ b/aci_edge_sandboxes/src/openvmm/platform/linux.rs @@ -19,7 +19,10 @@ pub(crate) const SUPPORTED: bool = true; /// Returns the start time of a live process, `None` if the process no longer exists, or an error /// if its state cannot be determined. /// -/// The value is the `starttime` field of `/proc//stat`. Zombie processes count as exited. +/// The value is the `starttime` field of `/proc//stat`. A process runs until its last thread +/// exits, as on Windows: a zombie whose main thread exited first still runs while its other threads +/// exit, and they keep the process's descriptors, such as an inherited log and its lock, open. +/// Zombies without other threads count as exited. pub(crate) fn process_start_time(pid: u32) -> io::Result> { let stat = match fs::read_to_string(format!("/proc/{pid}/stat")) { Ok(stat) => stat, @@ -31,19 +34,45 @@ pub(crate) fn process_start_time(pid: u32) -> io::Result> { } Err(error) => return Err(error), }; + start_time_from_stat(pid, &stat, || other_threads_remain(pid)) +} + +/// Interprets the `/proc//stat` line of `pid` as [`process_start_time`] does, asking +/// `other_threads_remain` only about a zombie. +pub(crate) fn start_time_from_stat( + pid: u32, + stat: &str, + other_threads_remain: impl FnOnce() -> io::Result, +) -> io::Result> { let malformed = || io::Error::other(format!("/proc/{pid}/stat has an unexpected format")); + // The command name may contain spaces and parentheses, so the fields start after the last + // closing parenthesis. let fields: Vec<&str> = stat[stat.rfind(')').ok_or_else(malformed)? + 1..] .split_whitespace() .collect(); - match fields.first() { - Some(&"Z" | &"X") => Ok(None), - // Field 22 of the stat line; the fields after the command name start at field 3. - Some(_) => fields - .get(19) - .and_then(|value| value.parse().ok()) - .map(Some) - .ok_or_else(malformed), - None => Err(malformed()), + // Field 22 of the stat line; the fields after the command name start at field 3. + let start_time = fields + .get(19) + .and_then(|value| value.parse().ok()) + .ok_or_else(malformed)?; + match fields[0] { + "X" => Ok(None), + "Z" => Ok(other_threads_remain()?.then_some(start_time)), + _ => Ok(Some(start_time)), + } +} + +/// Returns whether a process has threads besides its main thread. +fn other_threads_remain(pid: u32) -> io::Result { + match fs::read_dir(format!("/proc/{pid}/task")) { + Ok(threads) => Ok(threads.count() > 1), + Err(error) + if error.kind() == io::ErrorKind::NotFound + || error.raw_os_error() == Some(libc::ESRCH) => + { + Ok(false) + } + Err(error) => Err(error), } } diff --git a/aci_edge_sandboxes/src/openvmm/process.rs b/aci_edge_sandboxes/src/openvmm/process.rs index 4b9fa707..cea9cee7 100644 --- a/aci_edge_sandboxes/src/openvmm/process.rs +++ b/aci_edge_sandboxes/src/openvmm/process.rs @@ -82,8 +82,8 @@ pub(crate) fn spawn( /// /// On Unix a background thread reaps the child once it exits so it does not linger as a zombie /// while this process lives. If that thread cannot start, the child becomes a zombie after it -/// exits; liveness checks treat zombies as exited, so only the process table entry leaks. The -/// OpenVMM process keeps running either way. +/// exits; liveness checks treat zombies without threads as exited, so only the process table +/// entry leaks. The OpenVMM process keeps running either way. pub(crate) fn detach_child(launched: Launched) { #[cfg(not(windows))] { @@ -220,4 +220,89 @@ mod tests { }; assert!(!child_keeps_inherited_writer(fallback)); } + + #[test] + fn zombies_run_while_other_threads_remain() { + // Field 22 of a stat line, the start time, is 22 here; the command name holds a space + // and a parenthesis. + let stat = |state: &str| { + let fields: Vec = (4..=22).map(|field| field.to_string()).collect(); + format!("7 (open vmm)) {state} {}\n", fields.join(" ")) + }; + let unasked = || -> io::Result { panic!("only zombies count their threads") }; + let start_time = |state: &str, remain: io::Result| { + platform::start_time_from_stat(7, &stat(state), || remain) + }; + for state in ["R", "S", "D", "T"] { + assert_eq!( + platform::start_time_from_stat(7, &stat(state), unasked).unwrap(), + Some(22) + ); + } + assert_eq!( + platform::start_time_from_stat(7, &stat("X"), unasked).unwrap(), + None + ); + assert_eq!(start_time("Z", Ok(true)).unwrap(), Some(22)); + assert_eq!(start_time("Z", Ok(false)).unwrap(), None); + assert!(start_time("Z", Err(io::Error::other("unreadable"))).is_err()); + assert!(platform::start_time_from_stat(7, "7 (vmm) S 1 2", unasked).is_err()); + assert!(platform::start_time_from_stat(7, "garbage", unasked).is_err()); + } + + /// Kills and reaps a child process when dropped, even when a test fails. + struct Reaped(Child); + + impl Drop for Reaped { + fn drop(&mut self) { + let _ = self.0.kill(); + let _ = self.0.wait(); + } + } + + #[test] + fn a_process_runs_until_its_last_thread_exits() { + // The main thread leaves through pthread_exit, which turns it into a zombie, while + // another thread sleeps. Python is the only portable way to arrange that here. + let spawned = Command::new("python3") + .args([ + "-c", + "import ctypes, threading, time\n\ + threading.Thread(target=time.sleep, args=(60,)).start()\n\ + ctypes.CDLL(None).pthread_exit(None)\n", + ]) + .stdin(Stdio::null()) + .spawn(); + let mut child = match spawned { + Ok(child) => Reaped(child), + Err(error) if error.kind() == io::ErrorKind::NotFound => { + eprintln!("skipping: python3 is not installed"); + return; + } + Err(error) => panic!("cannot start python3: {error}"), + }; + let pid = child.0.id(); + let start_time = platform::process_start_time(pid).unwrap().unwrap(); + let deadline = Instant::now() + Duration::from_secs(10); + loop { + let stat = std::fs::read_to_string(format!("/proc/{pid}/stat")).unwrap(); + if stat[stat.rfind(')').unwrap() + 1..] + .trim_start() + .starts_with('Z') + { + break; + } + assert!(Instant::now() < deadline, "the main thread did not exit"); + thread::sleep(POLL_INTERVAL); + } + assert_eq!(platform::process_start_time(pid).unwrap(), Some(start_time)); + assert!(!wait_for_exit( + pid, + start_time, + Instant::now() + Duration::from_millis(100) + )); + child.0.kill().unwrap(); + child.0.wait().unwrap(); + assert!(wait_for_exit(pid, start_time, Instant::now())); + } } diff --git a/aci_edge_sandboxes/src/openvmm/state.rs b/aci_edge_sandboxes/src/openvmm/state.rs index 8a60543e..7e2ecd64 100644 --- a/aci_edge_sandboxes/src/openvmm/state.rs +++ b/aci_edge_sandboxes/src/openvmm/state.rs @@ -81,6 +81,8 @@ pub(crate) struct SandboxRecord { /// /// The image is referenced by its registered content ID (`sha256:`), whose registration /// supplies the path; the digests of the runtime files and host library are lowercase hexadecimal. +/// The host library's plan of the sandbox's host paths and network, if the request had any, is +/// kept as the library wrote it and handed back to the library at every start. #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[serde(rename_all = "camelCase", deny_unknown_fields)] pub(crate) struct NativeArtifactRecord { @@ -89,6 +91,8 @@ pub(crate) struct NativeArtifactRecord { pub(crate) kernel_sha256: String, pub(crate) initrd_sha256: String, pub(crate) library_sha256: String, + #[serde(default, skip_serializing_if = "Option::is_none")] + pub(crate) devices: Option, } /// Identity of the OpenVMM process of a running sandbox. @@ -571,6 +575,7 @@ mod tests { kernel_sha256: "2".repeat(64), initrd_sha256: "3".repeat(64), library_sha256: "4".repeat(64), + devices: Some(serde_json::json!({ "format": 1, "network": { "egress": "deny" } })), }); store.create(&id, &native).unwrap(); let (_guard, loaded) = store.lock_and_load(&id).unwrap(); diff --git a/aci_edge_sandboxes/tests/nvxhost_guest.rs b/aci_edge_sandboxes/tests/nvxhost_guest.rs index 2f99801d..9a665ccc 100644 --- a/aci_edge_sandboxes/tests/nvxhost_guest.rs +++ b/aci_edge_sandboxes/tests/nvxhost_guest.rs @@ -1,25 +1,33 @@ -//! Opt-in WHP lifecycle and robustness proofs with a separately supplied native library and guest. +//! Opt-in lifecycle, robustness, host-path, and network proofs with a separately supplied native +//! library and guest. //! //! The tests need `NVXHOST_TEST_OPENVMM`, `NVXHOST_TEST_KERNEL`, `NVXHOST_TEST_INITRD`, -//! `NVXHOST_TEST_IMAGE`, `NVXHOST_TEST_LIBRARY`, and the approved `NVXHOST_TEST_SHA256`. Run them -//! one at a time so that the check for leftover OpenVMM processes is exact, and optimized, since -//! creating a backend hashes the runtime files and registering an image hashes the image: +//! `NVXHOST_TEST_IMAGE`, `NVXHOST_TEST_LIBRARY`, and the approved `NVXHOST_TEST_SHA256`. +//! `NVXHOST_TEST_HYPERVISOR` selects `whp`, `mshv`, or `kvm`, and defaults to `whp` on Windows. +//! The host-path and network tests also need `python3` in the image. Run the tests one at a time +//! so that the check for leftover OpenVMM processes is exact, and optimized, since creating a +//! backend hashes the runtime files and registering an image hashes the image: //! `cargo test --release --features nvxhost --test nvxhost_guest -- --ignored --test-threads=1`. #![cfg(feature = "nvxhost")] use std::collections::BTreeSet; +use std::io::{ErrorKind, Write}; +use std::net::{IpAddr, Ipv4Addr, TcpListener, UdpSocket}; use std::path::{Path, PathBuf}; use std::process::{Child, Command, Stdio}; use std::sync::Arc; +use std::sync::atomic::{AtomicBool, Ordering}; use std::thread; use std::time::{Duration, Instant}; use aci_edge_sandboxes::openvmm::{ Hypervisor, ImageDigest, ImageId, NvxHostBackend, NvxHostConfig, OpenVmmConfig, + resolve_guest_path, }; use aci_edge_sandboxes::{ - AciEdgeSandbox, Error, ErrorCode, ExecOutcome, ExecOutput, ExecRequest, ProvisionRequest, - Result, SandboxId, StdinMode, StopResult, + Access, AciEdgeSandbox, EgressPolicy, Error, ErrorCode, ExecOutcome, ExecOutput, ExecRequest, + FilesystemPolicy, NetworkPolicy, NetworkRule, Protocol, ProvisionRequest, Result, SandboxId, + StdinMode, StopResult, }; const HELPER_STATE: &str = "NVXHOST_TEST_HELPER_STATE"; @@ -42,6 +50,17 @@ fn approved_digest() -> [u8; 32] { digest } +/// Returns the hypervisor that `NVXHOST_TEST_HYPERVISOR` names, which defaults to WHP on Windows. +fn hypervisor() -> Hypervisor { + match std::env::var("NVXHOST_TEST_HYPERVISOR") { + Ok(name) => name + .parse() + .unwrap_or_else(|error| panic!("NVXHOST_TEST_HYPERVISOR: {error}")), + Err(_) if cfg!(windows) => Hypervisor::Whp, + Err(_) => panic!("NVXHOST_TEST_HYPERVISOR must name the hypervisor, such as mshv"), + } +} + /// Returns the test configuration for `image`, with sandbox state under `state`. fn native_config(state: &Path, image: &Path) -> NvxHostConfig { NvxHostConfig::new( @@ -49,7 +68,7 @@ fn native_config(state: &Path, image: &Path) -> NvxHostConfig { required("NVXHOST_TEST_OPENVMM"), required("NVXHOST_TEST_KERNEL"), required("NVXHOST_TEST_INITRD"), - Hypervisor::Whp, + hypervisor(), state, ), image, @@ -199,10 +218,15 @@ impl Drop for Cleanup<'_> { } fn provision<'a>(client: &'a AciEdgeSandbox, backend: &'a NvxHostBackend) -> Cleanup<'a> { - let id = client - .provision(&ProvisionRequest::new()) - .unwrap() - .sandbox_id; + provision_with(client, backend, &ProvisionRequest::new()) +} + +fn provision_with<'a>( + client: &'a AciEdgeSandbox, + backend: &'a NvxHostBackend, + request: &ProvisionRequest, +) -> Cleanup<'a> { + let id = client.provision(request).unwrap().sandbox_id; Cleanup { client, backend, @@ -211,6 +235,17 @@ fn provision<'a>(client: &'a AciEdgeSandbox, backend: &'a NvxHostBackend) -> Cle } } +/// Provisions a sandbox for `request` and starts it. +fn started<'a>( + client: &'a AciEdgeSandbox, + backend: &'a NvxHostBackend, + request: &ProvisionRequest, +) -> Cleanup<'a> { + let sandbox = provision_with(client, backend, request); + client.start(&sandbox.id).unwrap(); + sandbox +} + fn exec(client: &AciEdgeSandbox, id: &SandboxId, request: &ExecRequest) -> ExecOutput { client .exec(id, request) @@ -228,6 +263,29 @@ fn assert_prints(client: &AciEdgeSandbox, id: &SandboxId, text: &str) { assert_eq!(String::from_utf8_lossy(&output.stdout), text); } +/// Runs `argv` in the guest without a shell, so that paths need no quoting. +fn run(client: &AciEdgeSandbox, id: &SandboxId, argv: &[&str]) -> ExecOutput { + exec(client, id, &ExecRequest::argv(argv.iter().copied())) +} + +/// Runs a Python program with `args` in the guest and returns what it printed. +fn python(client: &AciEdgeSandbox, id: &SandboxId, program: &str, args: &[&str]) -> String { + let mut argv = vec!["/bin/sh", "-c", "exec python3 -c \"$@\"", "sh", program]; + argv.extend_from_slice(args); + let output = exec( + client, + id, + &ExecRequest::argv(argv).with_timeout(Duration::from_secs(120)), + ); + assert_eq!(output.outcome, ExecOutcome::Exited(0), "{output:?}"); + String::from_utf8(output.stdout).unwrap() +} + +/// The guest path of an existing host path, which mappings derive from the resolved path. +fn guest(path: &Path) -> String { + resolve_guest_path(path).unwrap() +} + fn forced(stopped: &StopResult) -> Option { stopped .metadata @@ -305,14 +363,14 @@ fn start_in_helper(state: &Path, id: &SandboxId) -> Child { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_guest_lifecycle_stops_gracefully_and_cleans_up() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn guest_lifecycle_stops_gracefully_and_cleans_up() { let state = tempfile::tempdir().unwrap(); let config = OpenVmmConfig::new( required("NVXHOST_TEST_OPENVMM"), required("NVXHOST_TEST_KERNEL"), required("NVXHOST_TEST_INITRD"), - Hypervisor::Whp, + hypervisor(), state.path(), ); let backend = Arc::new( @@ -422,7 +480,7 @@ fn whp_guest_lifecycle_stops_gracefully_and_cleans_up() { let stop = client.stop(&sandbox_id); let deprovision = client.deprovision(&sandbox_id); panic!( - "WHP lifecycle failed: {error}; recovery stop: {stop:?}; \ + "guest lifecycle failed: {error}; recovery stop: {stop:?}; \ recovery deprovision: {deprovision:?}; OpenVMM log: {log}; \ guest console: {console}; OpenVMM outcome: {outcome}" ); @@ -430,8 +488,8 @@ fn whp_guest_lifecycle_stops_gracefully_and_cleans_up() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_exec_reports_exit_codes_streams_timeouts_and_output_limits() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn exec_reports_exit_codes_streams_timeouts_and_output_limits() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend(state.path()); @@ -508,8 +566,8 @@ fn whp_exec_reports_exit_codes_streams_timeouts_and_output_limits() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_lifecycle_errors_follow_state_and_a_stopped_sandbox_restarts() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn lifecycle_errors_follow_state_and_a_stopped_sandbox_restarts() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend(state.path()); @@ -541,8 +599,8 @@ fn whp_lifecycle_errors_follow_state_and_a_stopped_sandbox_restarts() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_new_backend_reattaches_to_a_running_guest_and_forces_a_stop() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn new_backend_reattaches_to_a_running_guest_and_forces_a_stop() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let id = { @@ -553,8 +611,9 @@ fn whp_new_backend_reattaches_to_a_running_guest_and_forces_a_stop() { sandbox.release() }; + // No graceful shutdown fits in a nanosecond, so the stop has to terminate OpenVMM. let second = backend_with(state.path(), false, |config| { - config.stop_timeout = Duration::from_millis(1); + config.stop_timeout = Duration::from_nanos(1); }); let client = AciEdgeSandbox::from_shared(second.clone()); let sandbox = Cleanup { @@ -579,8 +638,8 @@ fn whp_new_backend_reattaches_to_a_running_guest_and_forces_a_stop() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_a_crashed_guest_is_not_running_and_restarts_with_its_console_captured() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn a_crashed_guest_is_not_running_and_restarts_with_its_console_captured() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend_with(state.path(), true, |_| {}); @@ -615,7 +674,7 @@ fn whp_a_crashed_guest_is_not_running_and_restarts_with_its_console_captured() { } #[test] -#[ignore = "helper process for whp_guest_survives_its_caller_and_a_killed_start_leaves_no_vm"] +#[ignore = "helper process for guest_survives_its_caller_and_a_killed_start_leaves_no_vm"] fn helper_starts_a_sandbox_in_another_process() { let Some(sandbox) = std::env::var_os(HELPER_SANDBOX) else { return; @@ -626,8 +685,8 @@ fn helper_starts_a_sandbox_in_another_process() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_guest_survives_its_caller_and_a_killed_start_leaves_no_vm() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn guest_survives_its_caller_and_a_killed_start_leaves_no_vm() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend_with(state.path(), false, |config| { @@ -667,8 +726,8 @@ fn whp_guest_survives_its_caller_and_a_killed_start_leaves_no_vm() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_parallel_sandboxes_and_repeated_restarts_stay_isolated() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn parallel_sandboxes_and_repeated_restarts_stay_isolated() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend(state.path()); @@ -722,8 +781,8 @@ fn whp_parallel_sandboxes_and_repeated_restarts_stay_isolated() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_concurrent_commands_share_one_guest() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn concurrent_commands_share_one_guest() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let backend = backend(state.path()); @@ -764,8 +823,8 @@ fn whp_concurrent_commands_share_one_guest() { } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_registered_images_start_without_hashing_and_fail_closed_after_a_change() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn registered_images_start_without_hashing_and_fail_closed_after_a_change() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let image = copied_image(state.path()); @@ -832,8 +891,8 @@ fn whp_registered_images_start_without_hashing_and_fail_closed_after_a_change() } #[test] -#[ignore = "requires an approved private DLL, edge initramfs, GPT image, and a WHP host"] -fn whp_trusted_digests_and_content_verification_follow_their_contracts() { +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn trusted_digests_and_content_verification_follow_their_contracts() { let state = tempfile::tempdir().unwrap(); let before = openvmm_processes(); let image = copied_image(state.path()); @@ -905,3 +964,535 @@ fn whp_trusted_digests_and_content_verification_follow_their_contracts() { client.deprovision(&sandbox.release()).unwrap(); assert_no_new_openvmm(&before); } + +#[test] +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn host_paths_follow_the_filesystem_policy() { + let state = tempfile::tempdir().unwrap(); + let host = tempfile::tempdir().unwrap(); + let before = openvmm_processes(); + let work = host.path().join("work"); + for directory in ["src/secret", "out/frozen", "other"] { + std::fs::create_dir_all(work.join(directory)).unwrap(); + } + for (file, content) in [ + ("src/a.txt", "source"), + ("src/secret/key", "hidden"), + ("out/frozen/b.txt", "frozen"), + ("config.json", "{}"), + ("other/c.txt", "other"), + ] { + std::fs::write(work.join(file), content).unwrap(); + } + let (src, out, frozen, config) = ( + work.join("src"), + work.join("out"), + work.join("out").join("frozen"), + work.join("config.json"), + ); + let request = ProvisionRequest::new().with_filesystem(FilesystemPolicy { + // A read-only file inside the read-only src, which OpenVMM exports read-write for out, + // and a read-only directory inside the read-write out, which only the guest protects. + readonly_paths: vec![ + src.clone(), + config.clone(), + src.join("a.txt"), + frozen.clone(), + ], + readwrite_paths: vec![out.clone()], + denied_paths: vec![src.join("secret")], + }); + let backend = backend(state.path()); + let client = AciEdgeSandbox::from_shared(backend.clone()); + // Workloads never see the sandbox state, which holds the plan of every sandbox: a mapping + // inside the state root fails, and one that contains it needs a denied path that hides it. + let exposing = |path: &Path, denied: Vec| { + ProvisionRequest::new().with_filesystem(FilesystemPolicy { + readonly_paths: vec![path.to_path_buf()], + denied_paths: denied, + ..FilesystemPolicy::default() + }) + }; + let parent = state.path().parent().unwrap(); + for request in [ + exposing(state.path(), Vec::new()), + exposing(parent, Vec::new()), + ] { + assert_eq!( + failure(client.provision(&request)), + ErrorCode::PolicyValidation + ); + } + let hidden = client + .provision(&exposing(parent, vec![state.path().to_path_buf()])) + .unwrap(); + client.deprovision(&hidden.sandbox_id).unwrap(); + + let sandbox = started(&client, &backend, &request); + let id = &sandbox.id; + let (guest_src, guest_out, guest_frozen, guest_config) = + (guest(&src), guest(&out), guest(&frozen), guest(&config)); + + let read = run( + &client, + id, + &[ + "/bin/cat", + &format!("{guest_src}/a.txt"), + &guest_config, + &format!("{guest_frozen}/b.txt"), + ], + ); + assert_eq!(read.stdout, b"source{}frozen", "{read:?}"); + let listing = run(&client, id, &["/bin/ls", "-a", &guest_src]); + assert_eq!(listing.stdout, b".\n..\na.txt\n", "{listing:?}"); + let attempts: [[&str; 2]; 7] = [ + ["/bin/cat", &format!("{guest_src}/secret/key")], + ["/bin/touch", &format!("{guest_src}/new")], + ["/bin/touch", &guest_config], + ["/bin/touch", &format!("{guest_frozen}/new")], + ["/bin/rm", &format!("{guest_frozen}/b.txt")], + ["/bin/ls", "/run/nvx/hostfs"], + ["/bin/ls", &guest(&work.join("other"))], + ]; + for denied in attempts { + let output = run(&client, id, &denied); + assert_ne!( + output.outcome, + ExecOutcome::Exited(0), + "{denied:?}: {output:?}" + ); + } + let written = run( + &client, + id, + &[ + "/bin/sh", + "-c", + "echo written > \"$1\"", + "sh", + &format!("{guest_out}/result"), + ], + ); + assert_eq!(written.outcome, ExecOutcome::Exited(0), "{written:?}"); + assert_eq!( + std::fs::read_to_string(out.join("result")).unwrap(), + "written\n" + ); + assert_graceful(client.stop(id)); + client.deprovision(&sandbox.release()).unwrap(); + assert!(!src.join("new").exists() && !frozen.join("new").exists()); + assert_eq!( + std::fs::read_to_string(frozen.join("b.txt")).unwrap(), + "frozen" + ); + assert_no_new_openvmm(&before); +} + +/// Reports the workload's capabilities and the outcome of operations that would lift its mount +/// restrictions, given a read-only mapping inside a read-write one and a file path in the latter. +const CONTAINMENT_PROBE: &str = r#" +import ctypes, errno, json, sys +libc = ctypes.CDLL(None, use_errno=True) +frozen, writable = sys.argv[1], sys.argv[2] +def call(name, *args): + ctypes.set_errno(0) + if getattr(libc, name)(*args) == 0: + return "ok" + return errno.errorcode.get(ctypes.get_errno(), str(ctypes.get_errno())) +def write(path): + try: + with open(path, "w") as file: + file.write("x") + return "ok" + except OSError as error: + return errno.errorcode.get(error.errno, str(error.errno)) +with open("/proc/self/status") as file: + status = dict(line.split(":", 1) for line in file.read().splitlines() if ":" in line) +with open("/proc/sys/user/max_user_namespaces") as file: + user_namespaces = file.read().strip() +names = ("CapInh", "CapPrm", "CapEff", "CapBnd", "CapAmb", "NoNewPrivs") +print(json.dumps({ + "capabilities": [status[name].strip() for name in names], + "userNamespaces": user_namespaces, + "write": write(writable), + "writeFrozen": write(frozen + "/new"), + "remount": call("mount", None, frozen.encode(), None, ctypes.c_ulong(32 | 4096), None), + "tmpfs": call("mount", b"tmpfs", b"/tmp", b"tmpfs", ctypes.c_ulong(0), None), + "export": call("mount", b"microvm", b"/mnt", b"virtiofs", ctypes.c_ulong(0), None), + "unmount": call("umount2", frozen.encode(), 2), + "unshareMount": call("unshare", 0x20000), + "unshareUser": call("unshare", 0x10000000), + "chroot": call("chroot", b"/tmp"), +})) +"#; + +#[test] +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn workloads_cannot_lift_their_mount_restrictions() { + let state = tempfile::tempdir().unwrap(); + let host = tempfile::tempdir().unwrap(); + let before = openvmm_processes(); + let out = host.path().join("out"); + let frozen = out.join("frozen"); + std::fs::create_dir_all(&frozen).unwrap(); + let request = ProvisionRequest::new().with_filesystem(FilesystemPolicy { + readonly_paths: vec![frozen.clone()], + readwrite_paths: vec![out.clone()], + ..FilesystemPolicy::default() + }); + let backend = backend(state.path()); + let client = AciEdgeSandbox::from_shared(backend.clone()); + let sandbox = started(&client, &backend, &request); + let printed = python( + &client, + &sandbox.id, + CONTAINMENT_PROBE, + &[&guest(&frozen), &format!("{}/written", guest(&out))], + ); + let report: serde_json::Value = serde_json::from_str(&printed).unwrap(); + + // The workload stays root, with only the default container capabilities: CHOWN, + // DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, NET_BIND_SERVICE, and AUDIT_WRITE. + let (none, default) = ("0000000000000000", "00000000200004fb"); + assert_eq!( + report["capabilities"], + serde_json::json!([none, default, default, default, none, "1"]), + "{report}" + ); + assert_eq!(report["userNamespaces"], "0", "{report}"); + assert_eq!(report["write"], "ok", "{report}"); + assert_eq!(report["writeFrozen"], "EROFS", "{report}"); + for operation in [ + "remount", + "tmpfs", + "export", + "unmount", + "unshareMount", + "chroot", + ] { + assert_eq!(report[operation], "EPERM", "{operation}: {report}"); + } + assert_ne!(report["unshareUser"], "ok", "{report}"); + assert_graceful(client.stop(&sandbox.id)); + client.deprovision(&sandbox.release()).unwrap(); + assert!(out.join("written").is_file() && !frozen.join("new").exists()); + assert_no_new_openvmm(&before); +} + +#[test] +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn replaced_host_objects_inside_a_writable_mapping_fail_the_next_start() { + let state = tempfile::tempdir().unwrap(); + let host = tempfile::tempdir().unwrap(); + let before = openvmm_processes(); + let out = host.path().join("out"); + let (frozen, secret, moved) = (out.join("frozen"), out.join("secret"), out.join("moved")); + for directory in [&frozen, &secret] { + std::fs::create_dir_all(directory).unwrap(); + } + let request = ProvisionRequest::new().with_filesystem(FilesystemPolicy { + readonly_paths: vec![frozen.clone()], + readwrite_paths: vec![out.clone()], + denied_paths: vec![secret.clone()], + }); + let backend = backend(state.path()); + let client = AciEdgeSandbox::from_shared(backend.clone()); + let sandbox = started(&client, &backend, &request); + let id = &sandbox.id; + assert_prints(&client, id, "planned"); + assert_graceful(client.stop(id)); + + // A decoy at a protected path fails the start until the planned object is back. + for pinned in [&secret, &frozen] { + std::fs::rename(pinned, &moved).unwrap(); + std::fs::create_dir(pinned).unwrap(); + let error = client + .start(id) + .expect_err("a start must fail when a protected object was replaced"); + assert_eq!(error.code(), ErrorCode::BackendError, "{error}"); + assert!(error.message().contains("provision it again"), "{error}"); + std::fs::remove_dir(pinned).unwrap(); + std::fs::rename(&moved, pinned).unwrap(); + client.start(id).unwrap(); + assert_prints(&client, id, "restored"); + assert_graceful(client.stop(id)); + } + client.deprovision(&sandbox.release()).unwrap(); + assert_no_new_openvmm(&before); +} + +#[test] +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn every_mapping_that_fits_the_kernel_command_line_reaches_the_guest() { + let state = tempfile::tempdir().unwrap(); + let host = tempfile::tempdir().unwrap(); + let before = openvmm_processes(); + let directories: Vec = (0..64) + .map(|index| { + let directory = host.path().join(format!("m{index:02}")); + std::fs::create_dir(&directory).unwrap(); + std::fs::write(directory.join("index"), format!("{index} ")).unwrap(); + directory + }) + .collect(); + let request = |count: usize| { + ProvisionRequest::new().with_filesystem(FilesystemPolicy { + readonly_paths: directories[..count].to_vec(), + ..FilesystemPolicy::default() + }) + }; + let backend = backend(state.path()); + let client = AciEdgeSandbox::from_shared(backend.clone()); + + // Provisioning plans the mappings without a guest, so it finds the largest set that fits. + let mut fitting = 0; + for count in 1..=directories.len() { + match client.provision(&request(count)) { + Ok(provisioned) => { + client.deprovision(&provisioned.sandbox_id).unwrap(); + fitting = count; + } + Err(error) => { + assert_eq!(error.code(), ErrorCode::PolicyValidation, "{error}"); + assert!(error.message().contains("kernel command line"), "{error}"); + break; + } + } + } + assert!( + (2..directories.len()).contains(&fitting), + "{fitting} mappings fit the kernel command line" + ); + let sandbox = started(&client, &backend, &request(fitting)); + let files: Vec = directories[..fitting] + .iter() + .map(|directory| format!("{}/index", guest(directory))) + .collect(); + let mut argv = vec!["/bin/cat"]; + argv.extend(files.iter().map(String::as_str)); + let output = run(&client, &sandbox.id, &argv); + let expected: String = (0..fitting).map(|index| format!("{index} ")).collect(); + assert_eq!( + String::from_utf8_lossy(&output.stdout), + expected, + "{output:?}" + ); + eprintln!("{fitting} mappings fit the kernel command line"); + assert_graceful(client.stop(&sandbox.id)); + client.deprovision(&sandbox.release()).unwrap(); + assert_no_new_openvmm(&before); +} + +/// Reports the guest's interfaces, source address, and resolver, and whether it reaches the +/// gateway's DNS service, two host services that greet it, and a host-loopback service through +/// the gateway. +const NETWORK_PROBE: &str = r#" +import json, socket, sys +gateway, host, allowed, other, loopback = sys.argv[1:6] +def connect(address, port, greeted=True): + try: + with socket.create_connection((address, int(port)), timeout=3) as connection: + return "reached:" + connection.recv(16).decode() if greeted else "reached" + except OSError: + return "blocked" +def source(): + try: + with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as probe: + probe.connect((gateway, 53)) + return probe.getsockname()[0] + except OSError: + return "none" +def resolver(): + try: + with open("/etc/resolv.conf") as file: + return file.read() + except OSError: + return None +print(json.dumps({ + "interfaces": sorted(name for _, name in socket.if_nameindex()), + "source": source(), + "dns": connect(gateway, 53, greeted=False), + "allowed": connect(host, allowed), + "other": connect(host, other), + "loopback": connect(gateway, loopback), + "resolver": resolver(), +})) +"#; + +/// Returns the host's IPv4 address on its default route. +fn host_address() -> Ipv4Addr { + let socket = UdpSocket::bind((Ipv4Addr::UNSPECIFIED, 0)).unwrap(); + // Connecting a UDP socket selects a route and a source address without sending anything. + socket + .connect((Ipv4Addr::new(192, 0, 2, 1), 9)) + .expect("the host needs a default IPv4 route"); + match socket.local_addr().unwrap().ip() { + IpAddr::V4(address) if !address.is_loopback() && !address.is_unspecified() => address, + address => panic!("the host's default route uses {address}"), + } +} + +/// Host TCP services that greet every connection with their name until they are dropped. +struct Services { + done: Arc, + threads: Vec>, +} + +impl Services { + fn serve(listeners: Vec<(TcpListener, &'static str)>) -> Self { + let done = Arc::new(AtomicBool::new(false)); + let threads = listeners + .into_iter() + .map(|(listener, greeting)| { + listener.set_nonblocking(true).unwrap(); + let done = done.clone(); + thread::spawn(move || { + while !done.load(Ordering::Relaxed) { + match listener.accept() { + Ok((mut stream, _)) => { + let _ = stream.set_nonblocking(false); + let _ = stream.write_all(greeting.as_bytes()); + } + Err(error) if error.kind() == ErrorKind::WouldBlock => { + thread::sleep(Duration::from_millis(20)); + } + Err(error) => panic!("the {greeting} service failed: {error}"), + } + } + }) + }) + .collect(); + Self { done, threads } + } +} + +impl Drop for Services { + fn drop(&mut self) { + self.done.store(true, Ordering::Relaxed); + for thread in self.threads.drain(..) { + let _ = thread.join(); + } + } +} + +#[test] +#[ignore = "requires an approved private library, edge initramfs, GPT image, and a hypervisor host"] +fn network_policies_are_enforced() { + let state = tempfile::tempdir().unwrap(); + let before = openvmm_processes(); + let host = host_address(); + // The guest must reach the host through its gateway rather than consider it on-link. + let (guest_network, guest_address, gateway) = if host.octets()[..3] == [10, 0, 0] { + ("10.0.1.2/24", "10.0.1.2", "10.0.1.1") + } else { + ("10.0.0.2/24", "10.0.0.2", "10.0.0.1") + }; + let listen = |address: Ipv4Addr| TcpListener::bind((address, 0)).unwrap(); + let (allowed, other, loopback) = (listen(host), listen(host), listen(Ipv4Addr::LOCALHOST)); + let ports = [&allowed, &other, &loopback] + .map(|listener| listener.local_addr().unwrap().port().to_string()); + let _services = Services::serve(vec![ + (allowed, "allowed"), + (other, "other"), + (loopback, "loopback"), + ]); + let backend = backend_with(state.path(), false, |config| { + config.guest_network = guest_network.to_owned(); + }); + let client = AciEdgeSandbox::from_shared(backend.clone()); + let host_rule = || NetworkRule::to(format!("{host}/32")); + let contained = |egress: EgressPolicy| NetworkPolicy { + egress, + ..NetworkPolicy::deny_all() + }; + // Each case lists whether the guest reaches the gateway's DNS service, the allowed host + // service, and the other host service. No case reaches the host's loopback. + let cases = [ + ( + "no device", + NetworkPolicy::deny_all(), + [false, false, false], + ), + ( + "allow", + NetworkPolicy::egress(Access::Allow), + [true, true, true], + ), + ( + "host allow rule", + contained( + EgressPolicy::new(Access::Deny) + .with_allow(host_rule().on_port(Protocol::Tcp, ports[0].parse().unwrap())), + ), + [false, true, false], + ), + ( + "host deny rule", + contained(EgressPolicy::new(Access::Allow).with_deny(host_rule())), + [true, false, false], + ), + ( + "gateway DNS rule", + contained( + EgressPolicy::new(Access::Deny) + .with_allow(NetworkRule::to(gateway).on_port(Protocol::Tcp, 53)), + ), + [true, false, false], + ), + ]; + let host_text = host.to_string(); + let reached = |yes: bool, greeting: &str| match (yes, greeting) { + (false, _) => "blocked".to_owned(), + (true, "") => "reached".to_owned(), + (true, greeting) => format!("reached:{greeting}"), + }; + let gateway_resolver = serde_json::json!(format!("nameserver {gateway}\n")); + for (name, policy, [dns, allowed, other]) in cases { + let nic = name != "no device"; + let sandbox = started( + &client, + &backend, + &ProvisionRequest::new().with_network(policy), + ); + let printed = python( + &client, + &sandbox.id, + NETWORK_PROBE, + &[gateway, &host_text, &ports[0], &ports[1], &ports[2]], + ); + let mut report: serde_json::Value = serde_json::from_str(&printed).unwrap(); + let fields = report.as_object_mut().unwrap(); + let resolver = fields.remove("resolver"); + let interfaces = fields.remove("interfaces").unwrap(); + let (interface_count, source) = if nic { (2, guest_address) } else { (1, "none") }; + assert!( + interfaces.as_array().unwrap().len() == interface_count + && interfaces + .as_array() + .unwrap() + .contains(&serde_json::json!("lo")), + "{name}: interfaces {interfaces}" + ); + assert_eq!( + report, + serde_json::json!({ + "source": source, + "dns": reached(dns, ""), + "allowed": reached(allowed, "allowed"), + "other": reached(other, "other"), + "loopback": "blocked", + }), + "{name}" + ); + // The guest names its gateway as resolver only when it may reach its DNS service. + assert_eq!( + resolver.as_ref() == Some(&gateway_resolver), + dns, + "{name}: resolver {resolver:?}" + ); + assert_graceful(client.stop(&sandbox.id)); + client.deprovision(&sandbox.release()).unwrap(); + } + assert_no_new_openvmm(&before); +}