Identity All the Way Down: Zones, SPIRE, and Hybrid Inference on Mandible
How Mandible's platform is built on Edera zones with a SPIRE server and agent on every host, and why I believe hypervisor-rooted workload identity is non-negotiable.
Most platforms treat workload identity as a feature bolted onto the network. We built Mandible the other way around: identity is the foundation, and the network is just how bytes move.
That decision was made possible by one property of the infrastructure underneath us. Every Edera host ships with a SPIRE server and agent built in, rooted in the hypervisor. This post walks through the platform architecture, why that built-in SPIRE matters, and what it let us do when we plugged a local NVIDIA DGX running Nemotron into the cloud.
The platform in one picture
Mandible runs colonies: autonomous agents that coordinate through signals in a shared environment rather than by talking to each other directly. A colony is a workload. The question every platform has to answer is where that workload runs and what it is allowed to touch.
Our hierarchy has four levels, and each is an addressable identity:
account (tenant) billing and data-isolation boundary
└── project collaboration and signal-routing unit
└── zone the attested, hardware-isolated workload unit
└── colony logical workload name (1:1 with a zone today)
Around that hierarchy sit a small number of services, each with an identity of its own:
- cloud-api, the control plane. It owns accounts, projects, deployments, budgets, and model grants, and it is also an OIDC issuer.
- signal-server, the coordination layer colonies use to observe and deposit signals within a project.
- LiteLLM, the model gateway. Every model call from a colony flows through it, so budgets, spend logs, and usage views all hang off one place.
- Edera hosts, bare-metal machines that run zones.
The unit that matters most is the zone.
Zones: hardware isolation as the default
A zone is not a container namespace. Edera Protect runs each zone as its own lightweight virtual machine with its own kernel, launched by a type-1 hypervisor. Two colonies on the same host share silicon and nothing else. A kernel exploit inside one zone lands on that zone's kernel, not the host's and not a neighbor's.
That gives us a clean set of boundaries to reason about:
- The host runs the Edera daemon, the hypervisor, and one SPIRE server.
- Each zone runs a colony, with its own kernel, filesystem, and network interface.
- The control plane decides which zone launches where and what it may access.

The Hosts view: Edera hosts reporting fresh placement telemetry, the zones scheduled on each, and the Edera Protect build they run.
Zones are cheap enough to be disposable. A full VM with its own kernel goes from created to ready in about three seconds, it is destroyed when a deployment ends, and nothing about a zone's lifetime is meant to outlive the credential it was given.

Guest and hypervisor telemetry for one zone: boot time, vCPUs, memory pressure, and network transfer, reported live.
Isolation on its own is only half the story. A perfectly isolated workload that authenticates with a copied API key is still a workload with a copied API key. What makes zones useful as a security primitive is that each one can prove what it is.
SPIRE built into the host
SPIFFE defines a portable identity for workloads. SPIRE is the reference implementation that attests workloads and issues them short-lived credentials called SVIDs. Most teams that adopt SPIRE run it as a separate deployment beside their orchestrator, and attestation quality depends on what that orchestrator can tell SPIRE about a process.
Edera does something different. Every host's daemon configuration enables SPIRE natively, with a server and an agent running on the host and node attestor plugins that are rooted in the hypervisor itself (edera-agent-nodeattestor and edera-server-nodeattestor). When a zone boots, the hypervisor attests the agent inside it. The attestor can pin the exact kernel image the zone booted from by digest, so "this is a real zone" also means "this is a real zone running a kernel we approved."
Identity reaches the zone through Edera's IDM rings, a dedicated channel between hypervisor and guest that exists before the zone's network does. The zone then talks to a local SPIFFE Workload API socket to fetch its SVID.
In our experience this has been the most reliable component in the stack. Attestation completes in under a second on every zone boot, with zero false negatives across the dozens of launches we have done since bringing it into production.

A running zone: attested, with the host it landed on and the digest of the kernel it booted from. That digest is what the attestor pins.
The key insight is that the thing doing the attesting is the thing doing the isolating. There is no gap between "the hypervisor started this VM" and "SPIRE believes this VM is who it says it is." That is a stronger root of trust than any orchestrator-level attestor can offer, because the hypervisor cannot be lied to by the workload it is running.
What Mandible builds on top of it
Edera's SPIRE gives us an attested fact per zone: agent X is running inside a genuine zone, on host Y, with kernel Z. What it does not know is which tenant, project, and colony that zone belongs to. Only the control plane knows that. So we bind the two together.
A single trust domain and identity scheme. Every Mandible identity lives under spiffe://mandible.cloud:
Zone: spiffe://mandible.cloud/account/{account}/project/{project}/zone/{zone}
Service: spiffe://mandible.cloud/service/{name}
Host: spiffe://mandible.cloud/host/{host}
The account segment makes per-tenant policy expressible as a prefix match. Colony name rides as a claim, not a path segment, because the zone is the attested unit and colonies are metadata about what runs in it.
Registration. At launch, cloud-api registers a SPIRE workload entry on the host whose parent is the exact attested agent for that zone and whose SPIFFE ID is the zone's Mandible identity. If that registration fails, the zone does not launch. No verified identity, no zone.
Attestation binds credential renewal. A colony starts with a short-lived bootstrap identity minted by cloud-api. It immediately fetches a JWT-SVID from its local Workload API and presents it to cloud-api. We verify the SVID against that host's SPIRE bundle, confirm the account, project, and zone still exist and are not terminal, record the attestation, and issue a fresh short-lived identity. This runs in require mode on our fleet: a workload that cannot prove itself does not get credentials, tenant secrets, or model-gateway access.
Short-lived means configurable. The default lifetime is one hour, and a fair objection is that an hour is a long time for an autonomous agent. Two things make it shorter than it sounds. The token is only accepted by Mandible services, never by a model provider or a tenant's GitHub, so its blast radius is the platform's own APIs. And the colony renews at half the remaining lifetime, re-proving itself with a fresh SVID each time, so an hour is the ceiling on how long a token could outlive its zone, not the interval between attestations. Operators who want that ceiling lower set MANDIBLE_ZONE_IDENTITY_TTL, down to five minutes; the attestation-freshness window tracks it at twice the lifetime. Short-lived is a property we enforce, and how short is your call.
Terminal states revoke. When a zone moves to destroying, exited, destroyed, or failed, renewal stops and the SPIRE entry is deleted alongside the zone's other credentials. The residual exposure window is bounded by whatever lifetime you configured, not by the lifetime of a static secret.
Trust roots that scale with hosts. Each Edera host runs its own SPIRE server, which means N hosts produce N roots of trust. cloud-api unions every host's bundle and serves the result at a public /.well-known/spiffe-bundle endpoint, so anything that needs to verify a zone's certificate, wherever it runs, can fetch current roots.
A platform CA for what cannot attest. LiteLLM runs on Fargate, which has no SPIRE agent. Some infrastructure will never be an Edera zone. Those services carry certificates from a platform CA we operate, with SPIFFE URI SANs naming what they are. Verifiers check zone credentials against the SPIRE bundle union and service credentials against the platform CA, and an identity is bound to the root class that is allowed to vouch for it. Every host in that union is ours today; binding each zone's credential to the specific host it runs on is the check we are adding before that stops being true. A "measured" identity and an "ours" identity are both first-class, but they are never confused.

The same zone's identity: its full SPIFFE ID under mandible.cloud, a short-lived self-renewing token at the default 60-minute lifetime, credentials that die with the zone, and a hardware-attested badge backed by SPIRE attestation and the enforced root-image policy.
The result is that no long-lived Mandible credential ever crosses into a zone, and every authorization decision downstream can start from a SPIFFE ID rather than an address.
The payoff: a local model joins the platform
Here is where the light bulb went off and we got excited about the architecture.
We have an NVIDIA DGX on a LAN behind NAT serving Nemotron through NVIDIA NIM. Every token it serves is a token not bought from a frontier provider. We wanted colonies to use it as a first-class model, with the same budgets, grants, and metering as any hosted model, without exposing an inference server to the internet on an API key.
Edera does not yet ship for arm64, so the DGX cannot run zones or hold an attested SVID. But the identity model already had an answer for infrastructure that is ours but not measured: the platform CA.
The request path. Policy and metering stay in the cloud. TLS ends on the DGX, and model execution stays on hardware we own.
The pieces are small and each one reuses something the platform already had:
- A gateway on the DGX, modeled on the mTLS terminator we already run in front of LiteLLM, pointed the other direction. It presents a server certificate from our platform CA, requires client certificates, and checks the caller's SPIFFE ID against an explicit allowlist before forwarding to NIM on loopback. The gateway can verify clients against the SPIRE bundle union as well as the platform CA, with each allowlist entry bound to the root class that must vouch for it. As deployed, it trusts the platform CA only: the single allowlisted identity is LiteLLM's, and zones reach the model through LiteLLM rather than directly, so there is no reason to admit bundle-signed certificates at this endpoint. Trust-domain membership is not enough on a public endpoint; the allowlist is required configuration, and an empty one refuses to boot.
- An identified, not attested, client. The credential LiteLLM presents to the DGX is a certificate from our platform CA naming
spiffe://mandible.cloud/service/litellm, minted off-box and stored in Secrets Manager. Unlike a zone's SVID it is not hardware-attested and not self-renewing; it is a service certificate with a fixed lifetime, and rotation is the control. Because it is the only holder of that identity, revoking it is a reissue under a new service name and a one-line allowlist change on the DGX, not a CA rotation. We are saying this plainly because the rest of this post is about short-lived, attested credentials, and this one is neither. - A relay that carries ciphertext. The DGX dials out to a raw TCP relay; a dialer sidecar beside LiteLLM connects to the other end. TLS terminates on the DGX. The relay is an untrusted actor in the availability path, never in the trust path, and can be swapped for a different provider or a plain port-forward without touching the security model.
- LiteLLM as the authorization layer. The DGX is registered as the model group
nemotron. When cloud-api deploys a colony, it mints a scoped, budgeted LiteLLM key listing the models that colony may call. A colony grantednemotronuses the local model; one without the grant cannot, no matter what it knows about the network.
The relay taught us the one lesson worth repeating. Our first configuration used a managed TLS endpoint, and connectivity worked. Then we inspected the peer certificate and found the relay's wildcard, not our gateway's. The relay was terminating TLS and could read plaintext inference traffic. We switched to raw TCP forwarding, and from outside the DGX network the peer certificate became ours, with anonymous clients refused in the handshake before they could send a request.

The inference endpoint card: model group, the logical route, and latency from a real generation through the full path.
To close the loop, staff can see the endpoint's health in the console. The card is driven by a real one-token generation through the full path, cached for five minutes, and it deliberately shows no relay hostnames, secret names, or certificate paths. Connected, degraded, or not configured is all an operator needs.
Why this generalizes
None of the inference work required a new security concept. It required an identity model with three properties, and Edera's built-in SPIRE is what made the first one honest:
- Attested workloads. Zones prove they are real, measured zones before they are given anything, because the hypervisor vouches for them.
- Identified infrastructure. Services that cannot be attested are still named, with certificates from a root we control, and verifiers know which root is allowed to vouch for which identity.
- Authorization above transport. mTLS answers who is connecting. Grants and budgets answer what they may do. Keeping those separate is what let a GPU in a lab become a metered model in a multi-tenant platform.
When Edera ships arm64, the DGX enrolls as a normal host, runs its own SPIRE server, and serves inference from a zone. The gateway and relay collapse away. Nothing about the identity model changes, because the identity model was never about where the hardware sits.
Explore Mandible Cloud or follow the project on GitHub.