Luigit
repositories / bugabinga.net

bugabinga.net

personal infrastructure for bugabinga!

owned by admin

.system/specs/BB-SPEC-5TQ2M8HD-singleton-service-deployment-api/index.md

Raw
Rendered preview

id: BB-SPEC-5TQ2M8HD type: spec title: Toad singleton deployment API

Toad singleton deployment API

Boundary

The deployment service is generic and application-agnostic. It knows enrolled services, immutable revisions, image digests, rollout operations, probation, promotion, and rollback. It must not contain Smith missions, ADW phases, approval semantics, application data models, or per-application policy.

It is a separate klops service, not part of Luci. Luci, operators, and other clients submit deployments; the deployment service owns rollout state and host mutation. Deployment implementation must not enter Luci production code.

systemd remains the process supervisor. The deployment service is the only writer of release activation for enrolled services.

Platform

Toad is a Go service deployed as a repository-managed Quadlet.

It runs under a dedicated rootless toad user with its own Podman socket, systemd user scope, and storage boundary, mirroring the existing ci boundary. Enrolled workloads run in that same toad user scope, so toad never needs root, another user's session bus, or a privilege escalation path. The oli and ci users remain outside toad's control.

Toad runs containerized and reaches the machine only through mounts of its own user's Podman socket, its own user's D-Bus session bus, its Quadlet directory, and its data directory. It uses the Podman HTTP API and the systemd D-Bus interface directly, so the image carries no container or init tooling. Root daemons, polkit rules, sudo rules, and additional public ports are out of scope.

Toad's own unit, image, and enrollment data are owned by OpenTofu, never by toad itself.

Enrollment

Enrollment is operator-managed through OpenTofu and is immutable at request time.

Any number of applications may be enrolled. Durable state, credentials, locks, rollouts, probation policy, and reporting are isolated per service. One service's rollout, probation failure, or rollback must not delay, block, or alter another service.

One enrolled service declares:

  • service identifier and Quadlet unit name;
  • rendered unit template under the toad user scope;
  • fixed volumes, networks, secret references, environment files, and confinement;
  • one readiness endpoint reachable without the client;
  • allowed OCI repository;
  • probation policy: rollout deadline, continuous-ready stability window, and restart budget.

Enrollment carries no credential material. Credentials are issued, listed, and revoked by toad at runtime, so adding or rotating a consumer changes no infrastructure configuration.

A deployment request must not be able to supply Quadlet text, host paths, mounts, commands, environment, capabilities, or probe targets. Changing any enrolled property is an infrastructure change, never a deployment request.

Enrolled services must not use mutable-tag AutoUpdate. Infrastructure apply must not delete, overwrite, or restart enrolled units. Garbage collection must not prune the promoted image, the retained rollback image, or a candidate image under probation.

Caddy, zot, Luci, the secrets service, and the deployment service itself are not enrolled.

Request contract

A deployment request carries exactly:

  • service identifier;
  • image reference pinned by registry manifest digest, fully qualified and platform-resolved;
  • expected generation;
  • idempotency key.

The request is rejected when the expected generation is not the current generation. Generation is monotonic and independent of content, so redeploying an earlier digest produces a new generation and stale requests cannot pass the check.

A repeated idempotency key returns the existing operation and never starts a second rollout. An accepted request is durable before the response returns an operation identifier. Client disconnection, client termination, or Luci restart must not cancel or alter the operation.

At most one rollout per service is active. A conflicting request is rejected with the current generation and operation identifier, not queued.

Rollout

One rollout performs, in order:

  1. persist deployment intent and the candidate revision;
  2. fetch and verify the candidate digest without starting application code;
  3. stop the current instance and verify termination;
  4. activate the candidate revision atomically;
  5. observe probation;
  6. promote, or roll back.

The workload is singleton. Two instances must never run concurrently, so a stop timeout without confirmed termination fails the rollout instead of starting the candidate. Downtime during recreate is accepted and expected.

During probation:

  • systemd may restart the candidate under its ordinary restart policy;
  • every restart or readiness failure resets continuous-ready time;
  • restarts are counted across the complete rollout attempt, not per window;
  • an exhausted restart budget or an exceeded rollout deadline ends probation in failure;
  • a full continuous-ready stability window promotes the candidate.

Readiness is distinct from process liveness. A readiness observation counts only when it is attributable to the expected unit invocation, container, and image digest. An unattributable, cached, or proxied response is not evidence. Monitoring gaps do not accumulate healthy time, and the rollout deadline does not reset when the deployment service restarts.

Promotion is persisted atomically together with generation, current revision, and retained rollback revision.

After promotion, ordinary systemd restart policy applies. Later failures restart and alert; they must not trigger automatic rollback, because they can no longer be attributed to the deployment.

Rollback

Automatic rollback happens only during probation and only to the last promoted healthy revision.

The rollback target, including its image, rendered unit, and referenced configuration, is retained locally until a successor is promoted. A failed candidate must never become the retained rollback target. Registry deletion, registry outage, or garbage collection must not make rollback impossible.

A failed rollback is terminal and visible. The service must not alternate between revisions.

Operator-requested rollback after promotion is an ordinary deployment of an earlier digest with a new generation.

Rollback restores the image and unit specification only. Application data, schema migrations, secret values, published artifacts, and external side effects are not rolled back. Enrolled services must keep schema and data changes backward compatible across the retained rollback window. An incompatible change requires an explicit operator procedure, never automatic reversion.

Durability and recovery

Deployment intent is persisted before every host mutation.

Persisted rollout state distinguishes at least: preparing, activating, probation, promoting, rolling back, promoted, rolled back, and failed.

On start, the deployment service reconciles persisted intent against observed host state. An interrupted operation either continues under its original deadlines or fails explicitly. Recovery must never adopt an unverified candidate as healthy, resume without re-verifying termination guarantees, or replay an already promoted operation.

Single-writer ownership survives retries, concurrent processes, and host restart. A candidate under probation must not perform unreconciled work after boot.

Deployment outcome and current runtime health are separate reported facts.

Interface

The service provides:

  • submit deployment;
  • read operation status: state, generation, candidate and previous digests, timestamps, restart count, and failure reason;
  • read service state: current revision, retained rollback revision, and runtime health;
  • read deployment logs scoped to one operation;
  • request rollback of the active probation.

Toad exposes two transports. A Unix socket, reachable only by the toad identity and host root, authenticates by filesystem permission and carries local administration. A loopback TCP listener authenticates bearer tokens scoped to exactly one service.

Tokens are issued, listed, and revoked through toad, stored only as hashes, and displayed once at issue time. Token administration is available on the Unix socket only, so a leaked service token cannot mint, enumerate, or revoke credentials. Revocation takes effect without restart or redeployment.

Registry push permission does not imply deployment permission. Clients never receive a Podman socket, systemd control, host shell, arbitrary journal access, or another service's scope.

Deployment logs exposed to a public client are sanitized and scoped to the operation.

Surfaces

One binary provides the server and the operator client. The client reaches the same API and offers status, deploy, rollback, log, and token administration. Locally it uses the Unix socket and needs no token at all; remote consumers use an issued service token. The host entry point is a stable wrapper into the running container, so operator commands do not depend on copies of the binary.

A read-only web surface reports current revision, generation, retained rollback revision, live readiness, sanitized enrollment fields, container state, systemd user-unit properties, and recent operations with their Toad-owned logs. It is served on a public hostname behind Caddy basic authentication and offers no mutation, token material, raw Quadlets, environment values, host journal entries, generic host metrics, or host paths. The public Caddy reaches a bundled rootless gateway through a shared Unix socket; the gateway exposes no host TCP listener, accepts only the configured hostname, and forwards over a private Toad network.

Toad exposes no dedicated public SSH endpoint. Relay forwarding, additional public ports, and unauthenticated access to any toad surface are out of scope.

Luci interaction

A Luci job may deploy an enrolled service by:

  1. building and testing;
  2. publishing the image to the registry;
  3. obtaining the published manifest digest;
  4. submitting that digest with a service-scoped credential;
  5. polling until the operation reaches a terminal state;
  6. failing the job when the deployment failed or rolled back.

Luci must expose the published manifest digest to the job before submission. Luci remains repository-agnostic: it learns no application semantics, holds no rollout state, and gains no host control.

Acceptance

  • A valid request deploys the pinned digest, promotes after the stability window, and reports a terminal success.
  • A candidate that never becomes ready fails by rollout deadline and restores the previous revision.
  • A candidate that crashes repeatedly fails by restart budget, not by an indefinitely resetting window.
  • A candidate that becomes ready, then restarts, resets continuous-ready time instead of promoting early.
  • Failures after promotion restart and alert without rolling back.
  • A stale generation is rejected; a repeated idempotency key returns the original operation; a concurrent request is rejected.
  • Redeploying a previously promoted digest yields a new generation and does not accept stale requests.
  • Termination that cannot be confirmed fails the rollout instead of starting a second instance.
  • Killing the deployment service at any mutation boundary, and rebooting the host during probation, leaves a reconcilable state and no silent promotion.
  • A readiness response that cannot be attributed to the expected instance does not count as healthy.
  • Registry outage, registry deletion, and garbage collection do not prevent rollback of a promoted service.
  • A failed rollback is reported as terminal and does not loop.
  • Infrastructure apply, auto-update, and scheduled garbage collection leave enrolled units and retained images intact.
  • A client credential cannot deploy another service, read another service's logs, or reach the host beyond the documented interface.
  • A Luci job in the application repository can publish, deploy, and fail correctly on rollback, without deployment logic inside Luci.

Non-goals

Out of scope for this specification:

  • pods, multi-container revisions, and multiple replicas;
  • rolling, canary, or blue-green strategies;
  • deployment-request-supplied Quadlet, mount, or configuration changes;
  • configuration and secret-value revisioning and rollback;
  • database backup, restore, or migration reversal;
  • Kubernetes targets and multi-host scheduling;
  • enrollment of platform services.
---
id: BB-SPEC-5TQ2M8HD
type: spec
title: Toad singleton deployment API
---

# Toad singleton deployment API

## Boundary

The deployment service is generic and application-agnostic.
It knows enrolled services, immutable revisions, image digests, rollout operations, probation, promotion, and rollback.
It must not contain Smith missions, ADW phases, approval semantics, application data models, or per-application policy.

It is a separate klops service, not part of Luci.
Luci, operators, and other clients submit deployments; the deployment service owns rollout state and host mutation.
Deployment implementation must not enter Luci production code.

systemd remains the process supervisor.
The deployment service is the only writer of release activation for enrolled services.

## Platform

Toad is a Go service deployed as a repository-managed Quadlet.

It runs under a dedicated rootless `toad` user with its own Podman socket, systemd user scope, and storage boundary, mirroring the existing `ci` boundary.
Enrolled workloads run in that same `toad` user scope, so toad never needs root, another user's session bus, or a privilege escalation path.
The `oli` and `ci` users remain outside toad's control.

Toad runs containerized and reaches the machine only through mounts of its own user's Podman socket, its own user's D-Bus session bus, its Quadlet directory, and its data directory.
It uses the Podman HTTP API and the systemd D-Bus interface directly, so the image carries no container or init tooling.
Root daemons, polkit rules, sudo rules, and additional public ports are out of scope.

Toad's own unit, image, and enrollment data are owned by OpenTofu, never by toad itself.

## Enrollment

Enrollment is operator-managed through OpenTofu and is immutable at request time.

Any number of applications may be enrolled.
Durable state, credentials, locks, rollouts, probation policy, and reporting are isolated per service.
One service's rollout, probation failure, or rollback must not delay, block, or alter another service.

One enrolled service declares:

- service identifier and Quadlet unit name;
- rendered unit template under the `toad` user scope;
- fixed volumes, networks, secret references, environment files, and confinement;
- one readiness endpoint reachable without the client;
- allowed OCI repository;
- probation policy: rollout deadline, continuous-ready stability window, and restart budget.

Enrollment carries no credential material.
Credentials are issued, listed, and revoked by toad at runtime, so adding or rotating a consumer changes no infrastructure configuration.

A deployment request must not be able to supply Quadlet text, host paths, mounts, commands, environment, capabilities, or probe targets.
Changing any enrolled property is an infrastructure change, never a deployment request.

Enrolled services must not use mutable-tag `AutoUpdate`.
Infrastructure apply must not delete, overwrite, or restart enrolled units.
Garbage collection must not prune the promoted image, the retained rollback image, or a candidate image under probation.

Caddy, zot, Luci, the secrets service, and the deployment service itself are not enrolled.

## Request contract

A deployment request carries exactly:

- service identifier;
- image reference pinned by registry manifest digest, fully qualified and platform-resolved;
- expected generation;
- idempotency key.

The request is rejected when the expected generation is not the current generation.
Generation is monotonic and independent of content, so redeploying an earlier digest produces a new generation and stale requests cannot pass the check.

A repeated idempotency key returns the existing operation and never starts a second rollout.
An accepted request is durable before the response returns an operation identifier.
Client disconnection, client termination, or Luci restart must not cancel or alter the operation.

At most one rollout per service is active.
A conflicting request is rejected with the current generation and operation identifier, not queued.

## Rollout

One rollout performs, in order:

1. persist deployment intent and the candidate revision;
2. fetch and verify the candidate digest without starting application code;
3. stop the current instance and verify termination;
4. activate the candidate revision atomically;
5. observe probation;
6. promote, or roll back.

The workload is singleton.
Two instances must never run concurrently, so a stop timeout without confirmed termination fails the rollout instead of starting the candidate.
Downtime during recreate is accepted and expected.

During probation:

- systemd may restart the candidate under its ordinary restart policy;
- every restart or readiness failure resets continuous-ready time;
- restarts are counted across the complete rollout attempt, not per window;
- an exhausted restart budget or an exceeded rollout deadline ends probation in failure;
- a full continuous-ready stability window promotes the candidate.

Readiness is distinct from process liveness.
A readiness observation counts only when it is attributable to the expected unit invocation, container, and image digest.
An unattributable, cached, or proxied response is not evidence.
Monitoring gaps do not accumulate healthy time, and the rollout deadline does not reset when the deployment service restarts.

Promotion is persisted atomically together with generation, current revision, and retained rollback revision.

After promotion, ordinary systemd restart policy applies.
Later failures restart and alert; they must not trigger automatic rollback, because they can no longer be attributed to the deployment.

## Rollback

Automatic rollback happens only during probation and only to the last promoted healthy revision.

The rollback target, including its image, rendered unit, and referenced configuration, is retained locally until a successor is promoted.
A failed candidate must never become the retained rollback target.
Registry deletion, registry outage, or garbage collection must not make rollback impossible.

A failed rollback is terminal and visible.
The service must not alternate between revisions.

Operator-requested rollback after promotion is an ordinary deployment of an earlier digest with a new generation.

Rollback restores the image and unit specification only.
Application data, schema migrations, secret values, published artifacts, and external side effects are not rolled back.
Enrolled services must keep schema and data changes backward compatible across the retained rollback window.
An incompatible change requires an explicit operator procedure, never automatic reversion.

## Durability and recovery

Deployment intent is persisted before every host mutation.

Persisted rollout state distinguishes at least: preparing, activating, probation, promoting, rolling back, promoted, rolled back, and failed.

On start, the deployment service reconciles persisted intent against observed host state.
An interrupted operation either continues under its original deadlines or fails explicitly.
Recovery must never adopt an unverified candidate as healthy, resume without re-verifying termination guarantees, or replay an already promoted operation.

Single-writer ownership survives retries, concurrent processes, and host restart.
A candidate under probation must not perform unreconciled work after boot.

Deployment outcome and current runtime health are separate reported facts.

## Interface

The service provides:

- submit deployment;
- read operation status: state, generation, candidate and previous digests, timestamps, restart count, and failure reason;
- read service state: current revision, retained rollback revision, and runtime health;
- read deployment logs scoped to one operation;
- request rollback of the active probation.

Toad exposes two transports.
A Unix socket, reachable only by the toad identity and host root, authenticates by filesystem permission and carries local administration.
A loopback TCP listener authenticates bearer tokens scoped to exactly one service.

Tokens are issued, listed, and revoked through toad, stored only as hashes, and displayed once at issue time.
Token administration is available on the Unix socket only, so a leaked service token cannot mint, enumerate, or revoke credentials.
Revocation takes effect without restart or redeployment.

Registry push permission does not imply deployment permission.
Clients never receive a Podman socket, systemd control, host shell, arbitrary journal access, or another service's scope.

Deployment logs exposed to a public client are sanitized and scoped to the operation.

## Surfaces

One binary provides the server and the operator client.
The client reaches the same API and offers status, deploy, rollback, log, and token administration.
Locally it uses the Unix socket and needs no token at all; remote consumers use an issued service token.
The host entry point is a stable wrapper into the running container, so operator commands do not depend on copies of the binary.

A read-only web surface reports current revision, generation, retained rollback revision, live readiness, sanitized enrollment fields, container state, systemd user-unit properties, and recent operations with their Toad-owned logs.
It is served on a public hostname behind Caddy basic authentication and offers no mutation, token material, raw Quadlets, environment values, host journal entries, generic host metrics, or host paths.
The public Caddy reaches a bundled rootless gateway through a shared Unix socket; the gateway exposes no host TCP listener, accepts only the configured hostname, and forwards over a private Toad network.

Toad exposes no dedicated public SSH endpoint.
Relay forwarding, additional public ports, and unauthenticated access to any toad surface are out of scope.

## Luci interaction

A Luci job may deploy an enrolled service by:

1. building and testing;
2. publishing the image to the registry;
3. obtaining the published manifest digest;
4. submitting that digest with a service-scoped credential;
5. polling until the operation reaches a terminal state;
6. failing the job when the deployment failed or rolled back.

Luci must expose the published manifest digest to the job before submission.
Luci remains repository-agnostic: it learns no application semantics, holds no rollout state, and gains no host control.

## Acceptance

- A valid request deploys the pinned digest, promotes after the stability window, and reports a terminal success.
- A candidate that never becomes ready fails by rollout deadline and restores the previous revision.
- A candidate that crashes repeatedly fails by restart budget, not by an indefinitely resetting window.
- A candidate that becomes ready, then restarts, resets continuous-ready time instead of promoting early.
- Failures after promotion restart and alert without rolling back.
- A stale generation is rejected; a repeated idempotency key returns the original operation; a concurrent request is rejected.
- Redeploying a previously promoted digest yields a new generation and does not accept stale requests.
- Termination that cannot be confirmed fails the rollout instead of starting a second instance.
- Killing the deployment service at any mutation boundary, and rebooting the host during probation, leaves a reconcilable state and no silent promotion.
- A readiness response that cannot be attributed to the expected instance does not count as healthy.
- Registry outage, registry deletion, and garbage collection do not prevent rollback of a promoted service.
- A failed rollback is reported as terminal and does not loop.
- Infrastructure apply, auto-update, and scheduled garbage collection leave enrolled units and retained images intact.
- A client credential cannot deploy another service, read another service's logs, or reach the host beyond the documented interface.
- A Luci job in the application repository can publish, deploy, and fail correctly on rollback, without deployment logic inside Luci.

## Non-goals

Out of scope for this specification:

- pods, multi-container revisions, and multiple replicas;
- rolling, canary, or blue-green strategies;
- deployment-request-supplied Quadlet, mount, or configuration changes;
- configuration and secret-value revisioning and rollback;
- database backup, restore, or migration reversal;
- Kubernetes targets and multi-host scheduling;
- enrollment of platform services.