Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
79 changes: 64 additions & 15 deletions aci_edge_sandboxes/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,8 +117,9 @@ compatible kernel and static edge-agent initramfs, a prepared GPT image, the abs
to `nvxhost.dll` (`libnvxhost.so` on Linux), and the independently approved SHA-256 of that
library. The Rust crate does **not** build, download, or publish the private library.
`NvxHostBackend::new` verifies the file digest and ABI version and requires the additive
`nvx_build_ramfs_launch_arguments` and `nvx_session_connect_verified` exports. A missing,
wrong-version, or wrong-digest library fails rather than falling back.
`nvx_plan_sandbox`, `nvx_build_ramfs_launch_arguments_with_plan`, and
`nvx_session_connect_verified` exports. A missing, wrong-version, or wrong-digest library fails
rather than falling back.

The image is attached read-only in OpenVMM's distro block slot. The edge guest validates
its GPT and p2+ ext4 layers and uses a RAM-backed tmpfs upper/work for its overlay; it
Expand Down Expand Up @@ -165,19 +166,21 @@ and `guestBuildId`; stop reports `forced` and, after a failed graceful shutdown,
`gracefulError`. A capture failure never fails stop or deprovision: both report it as
`consoleError`.

This first backend supports provision/start/exec/stop/deprovision with shell commands or
argv and a fixed caller-provided image. Positive execution timeouts must be whole seconds,
matching the guest RPC's precision; finer-grained timeouts fail validation rather than
silently extending execution. It explicitly rejects host file mappings, network
configuration beyond deny-all, piped stdin, execution cancellation, custom working
directories and environments. Snapshot/restore, image selection, and richer guest
operations are not part of this profile. See
This backend supports provision/start/exec/stop/deprovision with shell commands or argv, a
fixed caller-provided image, and the host paths and network rules described below. Positive
execution timeouts must be whole seconds, matching the guest RPC's precision; finer-grained
timeouts fail validation rather than silently extending execution. It explicitly rejects piped
stdin, execution cancellation, custom working directories and environments. Snapshot/restore,
image selection, and richer guest operations are not part of this profile. See
[`examples/nvxhost_lifecycle.rs`](examples/nvxhost_lifecycle.rs) for a run requiring
`--openvmm`, `--kernel`, `--initrd`, `--image`, `--host-library`, `--host-sha256`,
`--state-root`, `--hypervisor`, and a command after `--`; an optional `--image-sha256`
requires the image's digest when it is registered. The ignored
`tests/nvxhost_guest.rs` exercises an actual WHP guest when the corresponding
`NVXHOST_TEST_*` paths and approved DLL digest are set.
requires the image's digest when it is registered. The repeatable `--readonly`, `--readwrite`,
and `--denied` options map host paths, and `--egress allow|deny` with the repeatable
`--egress-allow` and `--egress-deny` options, each taking `CIDR` or `CIDR:tcp|udp:PORT`, attach a
network. The ignored `tests/nvxhost_guest.rs` exercises an actual guest when the corresponding
`NVXHOST_TEST_*` paths and approved library digest are set: under WHP on Windows, or under the
hypervisor that `NVXHOST_TEST_HYPERVISOR` names, such as `mshv` on Linux.

For example, from `aci_edge_sandboxes` on a Windows WHP host, set the following
paths to compatible, separately built artifacts and a caller-prepared GPT disk
Expand All @@ -199,12 +202,56 @@ cargo run --release --locked --features nvxhost --example nvxhost_lifecycle -- `
--state-root $stateRoot --hypervisor whp -- 'printf READY'
```

To map `C:\work\src` read-only and `C:\work\out` read-write, and let the guest reach only
`192.0.2.10` on TCP port 443, add
`--readonly C:\work\src --readwrite C:\work\out --egress deny --egress-allow 192.0.2.10:tcp:443`
before `--`. The example prints where each mapped path appears in the guest, here
`/mnt/c/work/src` and `/mnt/c/work/out`.

Build the native library and the static edge initramfs separately from their
matching private sources; this example neither fetches nor builds them. Use the
pinned OpenVMM, which includes the scratchless RAM-overlay topology, and a kernel
compatible with that OpenVMM and guest revision. On Linux, supply a matching
`libnvxhost.so` and OpenVMM build and select `--hypervisor mshv`; the Linux edge
lifecycle has not yet been verified end to end.
`libnvxhost.so` and OpenVMM build and select `--hypervisor mshv`.

### Native host paths and network

The native backend accepts the same `filesystem` and `network` policies as the direct
backend: it maps host paths to the same guest paths under the same rules and expands network
rules the same way; see [Host paths](#host-paths) and [Network rules](#network-rules). It also
refuses mappings that would show workloads its state root, described below. Its workloads run
as the guest's root, so it ignores `map_host_identity`. The private library plans the
policies. Provision passes them to `nvx_plan_sandbox`, which resolves the mapped and denied
paths, chooses the export, pins the objects inside read-write mappings, and expands the
network rules; the sandbox record keeps the resulting plan. Start passes the plan back, and
the library checks it again against every planning rule that needs no host access, and checks
the pinned objects: if one was replaced, start fails with `backend_error`, and the sandbox
must be deprovisioned and provisioned again to accept the change.

- The state root holds the record, and so the plan, of every sandbox, which decides what the
next start exports and which egress it allows. Provision therefore refuses, with
`policy_validation`, a mapped path inside the state root, and one that contains it unless a
denied path inside the mapping hides it.
- OpenVMM exports the deepest directory that contains every mapped path through one
virtio-fs device, read-only unless a path is read-write, and hides the denied paths. The
guest mounts the export where workloads cannot reach it and bind-mounts each mapped path
into the RAM overlay at its guest path, read-only where requested. The bind mounts travel on
the guest's 1024-byte kernel command line, which leaves room for roughly a dozen typical
paths; provision rejects a policy whose mounts do not fit with `policy_validation`.
- Workloads run as the guest's root with only the default container capabilities (`CHOWN`,
`DAC_OVERRIDE`, `FOWNER`, `FSETID`, `KILL`, `SETGID`, `SETUID`, `NET_BIND_SERVICE`, and
`AUDIT_WRITE`), with `no_new_privs`, and without user namespaces, so they cannot remount or
unmount the mapped paths or mount the export again. Host-side changes through read-write
mappings, including modes and ownership, happen with the credentials of the account that
runs OpenVMM, so run OpenVMM unprivileged. A host hard link that already joins a file in a
read-write mapping to one in a read-only mapping stays writable through the read-write
path.
- A policy that denies egress without allow rules attaches no network device. Otherwise the
guest gets `OpenVmmConfig::guest_network` (`10.0.0.2/24` by default) behind OpenVMM's NAT
gateway, the network's first address, and names that gateway as its DNS server in
`/etc/resolv.conf` when the policy allows TCP or UDP port 53 to it. Choose a guest network
that contains no address the guest must reach. Ingress and host-loopback access must be
`deny`.

The caller must independently approve and protect the native asset. Checking a caller-supplied
digest does not make a writable path or a self-declared digest trustworthy; use this profile
Expand Down Expand Up @@ -284,7 +331,9 @@ Defaults:
| `stop_timeout`, `exec_response_grace` | 30 s |

Each timeout must be at most 30 days (`OpenVmmConfig::MAX_TIMEOUT`), and all
but `exec_response_grace` must be positive.
but `exec_response_grace` must be positive. `guest_network` must be an address
that OpenVMM accepts: a /1 to /30 prefix, and neither the network's own address,
its broadcast address, nor its first address, which is the gateway.

Choose a `state_root` that only the current user can access.
`OpenVmmConfig::default_state_root` returns `%LOCALAPPDATA%\nvx\sandboxes` on
Expand Down
105 changes: 99 additions & 6 deletions aci_edge_sandboxes/examples/nvxhost_lifecycle.rs
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,12 @@ use std::process::ExitCode;
use std::sync::Arc;

use aci_edge_sandboxes::openvmm::{
Hypervisor, ImageDigest, NvxHostBackend, NvxHostConfig, OpenVmmConfig,
Hypervisor, ImageDigest, NvxHostBackend, NvxHostConfig, OpenVmmConfig, resolve_guest_path,
};
use aci_edge_sandboxes::{
Access, AciEdgeSandbox, EgressPolicy, ExecOutcome, ExecRequest, FilesystemPolicy,
NetworkPolicy, NetworkRule, Protocol, ProvisionRequest,
};
use aci_edge_sandboxes::{AciEdgeSandbox, ExecOutcome, ExecRequest, ProvisionRequest};

fn main() -> ExitCode {
match run() {
Expand All @@ -30,6 +33,10 @@ fn run() -> Result<(), String> {
let mut image_digest = ImageDigest::Compute;
let mut root = None;
let mut hypervisor = None;
let mut filesystem = FilesystemPolicy::default();
let mut egress = None;
let mut allow = Vec::new();
let mut deny = Vec::new();
let mut command = None;
let mut args = std::env::args().skip(1);
while let Some(option) = args.next() {
Expand All @@ -49,6 +56,12 @@ fn run() -> Result<(), String> {
}
"--state-root" => root = Some(value()?),
"--hypervisor" => hypervisor = Some(value()?.parse::<Hypervisor>().map_err(describe)?),
"--readonly" => filesystem.readonly_paths.push(value()?.into()),
"--readwrite" => filesystem.readwrite_paths.push(value()?.into()),
"--denied" => filesystem.denied_paths.push(value()?.into()),
"--egress" => egress = Some(parse_access(&option, &value()?)?),
"--egress-allow" => allow.push(parse_rule(&option, &value()?)?),
"--egress-deny" => deny.push(parse_rule(&option, &value()?)?),
"--" => {
command = Some(args.by_ref().collect::<Vec<_>>().join(" "));
break;
Expand All @@ -69,6 +82,33 @@ fn run() -> Result<(), String> {
let command = command
.filter(|command| !command.is_empty())
.ok_or("pass the guest command after --")?;
let mut request = ProvisionRequest::new();
if !filesystem.is_empty() {
for path in filesystem
.readonly_paths
.iter()
.chain(&filesystem.readwrite_paths)
{
if let Some(guest) = resolve_guest_path(path) {
eprintln!("{} is {guest} in the guest", path.display());
}
}
request = request.with_filesystem(filesystem);
}
match egress {
Some(default) => {
request = request.with_network(NetworkPolicy {
egress: EgressPolicy {
default,
allow,
deny,
},
..NetworkPolicy::deny_all()
});
}
None if allow.is_empty() && deny.is_empty() => {}
None => return Err("--egress-allow and --egress-deny require --egress".to_owned()),
}

let config = OpenVmmConfig::new(openvmm, kernel, initrd, hypervisor, PathBuf::from(root));
let backend = Arc::new(
Expand All @@ -78,10 +118,7 @@ fn run() -> Result<(), String> {
.map_err(describe)?,
);
let client = AciEdgeSandbox::from_shared(backend.clone());
let id = client
.provision(&ProvisionRequest::new())
.map_err(describe)?
.sandbox_id;
let id = client.provision(&request).map_err(describe)?.sandbox_id;
let started = client.start(&id);
let succeeded = started.is_ok();
let executed = started.and_then(|_| {
Expand Down Expand Up @@ -155,6 +192,38 @@ fn parse_digest(option: &str, hex: &str) -> Result<[u8; 32], String> {
Ok(digest)
}

fn parse_access(option: &str, text: &str) -> Result<Access, String> {
match text {
"allow" => Ok(Access::Allow),
"deny" => Ok(Access::Deny),
_ => Err(format!("{option} must be allow or deny")),
}
}

/// Parses `CIDR` or `CIDR:tcp|udp:PORT`, such as `192.0.2.0/24` or `192.0.2.1:tcp:443`.
fn parse_rule(option: &str, text: &str) -> Result<NetworkRule, String> {
let malformed = || format!("{option} must be CIDR or CIDR:tcp|udp:PORT");
let mut parts = text.split(':');
let rule = NetworkRule::to(
parts
.next()
.filter(|cidr| !cidr.is_empty())
.ok_or_else(malformed)?,
);
match (parts.next(), parts.next(), parts.next()) {
(None, ..) => Ok(rule),
(Some(protocol), Some(port), None) => {
let protocol = match protocol {
"tcp" => Protocol::Tcp,
"udp" => Protocol::Udp,
_ => return Err(malformed()),
};
Ok(rule.on_port(protocol, port.parse().map_err(|_| malformed())?))
}
_ => Err(malformed()),
}
}

fn describe(error: aci_edge_sandboxes::Error) -> String {
error.to_string()
}
Expand All @@ -171,4 +240,28 @@ mod tests {
assert!(parse(&"é".repeat(32)).is_err());
assert!(parse(&"gg".repeat(32)).is_err());
}

#[test]
fn egress_rules_name_a_network_and_optionally_one_port() {
let parse = |text: &str| parse_rule("--egress-allow", text);
assert_eq!(
parse("192.0.2.0/24").unwrap(),
NetworkRule::to("192.0.2.0/24")
);
assert_eq!(
parse("192.0.2.1:udp:53").unwrap(),
NetworkRule::to("192.0.2.1").on_port(Protocol::Udp, 53)
);
for malformed in [
"",
":tcp:1",
"192.0.2.1:tcp",
"192.0.2.1:icmp:1",
"192.0.2.1:tcp:x",
] {
assert!(parse(malformed).is_err(), "{malformed}");
}
assert_eq!(parse_access("--egress", "deny").unwrap(), Access::Deny);
assert!(parse_access("--egress", "Deny").is_err());
}
}
Loading