Pre-release. v0.1 is not out yet, so there is nothing to install and no public source to clone — the quickstart builds from a checkout.
Operate it
The control plane is one service and one database. This page covers what to know before it carries real traffic. Every setting is in configuration, and every command in the command line.
Signing keys
Section titled “Signing keys”Subact ID signs every token and every audit checkpoint with ES256 (ECDSA on P-256). Tool servers verify
tokens against the JWKS at /.well-known/jwks.json.
SubactId.Server keys generate --out /run/secrets/subactid/active.pemThis reads no configuration, so it works before any deployment exists. It writes an owner-readable
PKCS#8 PEM, refuses to overwrite an existing file, and prints the key id and the setting to use.
It never prints the private key. --kid <name> sets the key id instead of the RFC 7638
thumbprint.
A key made another way also works: an unencrypted P-256 private key in PKCS#8
(BEGIN PRIVATE KEY) or SEC 1 (BEGIN EC PRIVATE KEY) form, as a path to a mounted secret or
inline.
Outside the Development environment, the server does not start without a key. In Development
it generates an ephemeral key and logs that it did; tokens signed with it stop verifying after a
restart.
To see the published JWKS and the active key:
SubactId.Server keys{"keys":[{"kty":"EC","crv":"P-256","use":"sig","alg":"ES256","kid":"9Fy-qRMxKnbHaNBF9nwF68GQVxzhgz3miPA15FHHv38","x":"…","y":"…"}]}Active kid: 9Fy-qRMxKnbHaNBF9nwF68GQVxzhgz3miPA15FHHv38Rotating a key
Section titled “Rotating a key”SubactId.Server keys rotate --out /run/secrets/subactid/next.pemThis loads the configured keys as the server would, writes a new key file, and prints the settings for each step. It changes nothing else. Roll the key out in three deploys:
- Publish the new key without signing with it. Add the new key and set
SubactId:Signing:ActiveKidto the current key. With two keys,ActiveKidis required. Deploy, then wait for verifiers to refresh their cached JWKS. - Sign with the new key. Set
ActiveKidto the new key and deploy. Tokens signed by the old key still verify, because it is still published. - Optionally, stop publishing the old key, once the wait
rotateprinted has passed since step 2. The wait is the longestmax_token_ttlof any registered agent, disabled ones included. Waiting longer is always safe.
For a routine rotation, skip step 3. Checkpoints are signed with the active key, and a
checkpoint verifies only while its key is published. Removing a key that was ever active makes
audit-verify report the first checkpoint it signed as a bad signature, and archived months
sealed by it can no longer be verified. If you archive the ledger, keep every key that was ever
active for as long as you keep the exports.
Losing a key
Section titled “Losing a key”A lost key cannot be recovered, and every token it signed becomes unverifiable. Run
keys generate, configure the new key and restart; agents exchange again.
If a key is disclosed, treat every token it signed as compromised:
- Generate a new key and make it active at once.
- Remove the disclosed key without waiting.
- Revoke the affected tasks.
- Keep the last checkpoint
audit-verifyreported before the disclosure. Checkpoints the disclosed key signed can no longer be trusted.
Keys held in a KMS are not in v0.1. The signing key is a file.
Migrations
Section titled “Migrations”SubactId.Server migrateApplied 20 migration(s): - 20260910060829_InitialSchema - 20260910162631_AddAssertionReplays - 20260911081817_AddRevocationExpiry - 20260911124255_AddOutboxNextAttempt - 20260911132757_AddAuditQueryIndexes - 20260912131056_AddAgentJwks - 20260912132211_WidenAuditReason - 20260913223129_AddAuditEventCount - 20260915165924_AddAuditEventChain - 20260917123305_AddSponsorBlocks - 20260917161005_AddSponsorRevocations - 20260917185320_AddSessionsAndSignalReplays - 20260917213039_AddScimUsers - 20260918094512_ClearDeliveredOutboxEntries - 20260918223252_AddDenialIndex - 20260919094500_AddAuditCheckpoints - 20260919120000_PartitionAuditLedger - 20260919160000_AddAuditArchives - 20260919170000_FingerprintAuditQueryIndexes - 20260920185038_AddTaskGrantRenewalsSchema version: 20260920185038_AddTaskGrantRenewalsAudit ledger partitions: created 2, now 2 month(s) ahead.Migrations never run at startup. Run migrate with a MigrationConnectionString for a role that
may change the schema, and let the server connect as a role that may not. Each provider has its
own migrations, because the two schemas differ.
On Postgres, the audit ledger is partitioned by calendar month and has no default partition.
migrate creates the current month’s partition and SubactId:Audit:Partitions:MonthsAhead months
after it, and each server tops this up at start and hourly. A record for a month with no
partition is refused, and so is the action it records. An instance whose current month has no
partition fails /readyz. If the server’s role cannot change the schema, its top-up fails and
logs it, and doctor reports the shortfall: run migrate with the migration credentials.
The embedded database
Section titled “The embedded database”SubactId:Database:Provider=sqlite with a Path runs the same control plane with the same rules. Its
limits are operational:
- One writer at a time. Run one instance.
- Local disk only. Network filesystems are not supported.
- No failover. Losing the host loses anything not backed up.
- Back up with SQLite’s backup API or
VACUUM INTO, not by copying the file under a running server. - Nothing leaves the ledger. SQLite has no partitioning, so
audit-archiverefuses to run.
There is no data migration between providers. Moving to Postgres means a fresh database and registering agents again. Keep the SQLite file if you need its audit history.
The expiry sweeper
Section titled “The expiry sweeper”A background sweeper marks expired tasks terminal and revokes their grants, writing one
task.expired record per task. It runs every SubactId:Tasks:SweepInterval and expires
SubactId:Tasks:SweepBatchSize tasks per transaction, repeating until a batch comes back short. It
also deletes tasks that ended more than SubactId:Tasks:Retention ago (a week by default), with their
grants. The ledger keeps their records. An agent can be deleted only once its tasks are gone.
Expiry is never recorded as a revocation. A kill switch leaves a task already past its expiry to
the sweeper, so that task is not counted in revoked_tasks. A task that renewed more than once
also gets one token.refreshed summary with a count, written just before its terminal record.
Verifying the ledger
Section titled “Verifying the ledger”The ledger is sealed by signed checkpoints. Every SubactId:Audit:Checkpoint:Interval, a pass on each
instance takes the records not yet sealed, builds a Merkle tree over them, signs the root and links
the checkpoint to the previous one. Replicas share the work, and each range is sealed once.
SubactId.Server audit-verifyAudit ledger intact: 272 checkpoint(s) verified sealing 10512 record(s) up to seq 10512, last checkpoint 272:9c1f6a30….It checks each checkpoint’s signature, its link to the previous one, and its root against the
records it covers. It exits 0 when the seal is intact, 3 when it is not, 1 when the ledger
could not be read and 2 for a malformed argument.
Run it on a schedule, and store the last checkpoint somewhere the database cannot reach. Pass
it back on the next run and audit-verify also confirms that checkpoint is still present and
signs the same root. That is how a cut tail is detected:
SubactId.Server audit-verify 272:9c1f6a30…GET /audit/checkpoints publishes the checkpoints so you can keep copies outside the database.
The audit ledger explains what the seal proves.
Retention
Section titled “Retention”No record leaves the ledger on its own. On Postgres, the only way to remove records is to detach a whole month’s partition, and one command does that:
SubactId.Server audit-archive --before 2026-07 --to /var/lib/subactid/archiveFor each month before the cutoff, oldest first, it exports the month, reads the export back and
verifies it against the checkpoints that sealed it, writes an audit.archived record with the
export’s SHA-256, and then detaches and drops the partition. A month that fails stops the run with
nothing detached. With SubactId:Audit:Retention set, --before may be left out; the cutoff is then
the month containing now minus the retention. Retention must be at least P31D.
Archived 2026-06: 339726 record(s) sealed by checkpoints 1-4120, to /var/lib/subactid/archive/audit-2026-06.subactid-archive.gz (sha256 7c1e…).Archived 1 month(s); the online ledger now starts at checkpoint 4121 and holds nothing before 2026-07.Plan around three things:
- Nothing runs it for you. Run it from a timer of your own, for example monthly.
- Store exports away from the database. The export’s digest stays in the online ledger, so a
changed or truncated export is detectable.
audit-verify --archive <export>verifies one export.audit-verifywithout it starts from the checkpoint after the last archived one. Checkpoints are not removed, so acheckpoint_id:rootrecorded earlier still verifies. /auditafter an archive. Every page carriesarchived_before, the earliest instant still online, ornull. A range older than that is an empty page, not an error. A proof for an archived record answers410.
To keep an old month queryable in SQL without archiving it, detach its partition yourself and
drop its query indexes, but keep the table. /audit no longer returns it. There is no command for
this.
Running more than one instance
Section titled “Running more than one instance”Instances are interchangeable, on Postgres. Audit appends do not wait on each other, sealing
claims each range once, and the sink outbox is shared. Two things are per instance: the rate
limiter’s buckets, so n replicas admit up to n times each configured rate, and denial
aggregation, so n instances write up to n summaries per reason per window.
Readiness
Section titled “Readiness”/readyz answers 503 until four checks pass: database, audit-partition (this month has a
ledger partition), signing-key, and upstream-jwks (the identity provider’s discovery document
and keys have been fetched at least once). Point a load balancer at /readyz and a restart policy
at /healthz. doctor names a failing check.
Readiness does not go red when the identity provider stops answering after the instance is ready. The instance still holds the provider’s keys, so it can still validate subject tokens and issue tokens; taking every replica out of rotation would turn a provider outage into a control plane outage. Instead:
- The keys are re-fetched every five minutes.
- The first failed fetch logs a warning, and recovery logs an information line. Both contain
identity providerand name no person or token. - Once the last successful fetch is more than fifteen minutes old,
/readyzstill answers200, with the bodyDegraded.
During the outage, a refresh that needs a fresh sponsor check is refused with
temporarily_unavailable and audited as sponsor_status_unavailable. If the provider rotates its
keys during the outage, tokens signed with the new key are refused. Alert on the log warning, not
on readiness.
Kubernetes
Section titled “Kubernetes”The control plane repository has a Helm chart. It installs a Deployment and Service, a migrate
Job as a pre-install and pre-upgrade hook, and optionally an Ingress, a separate admin Ingress
and a NetworkPolicy. helm test runs doctor and checks /readyz.
- It supports Postgres only.
- It creates no Secret and no key. It reads Secrets you create, by name.
- It requires four values:
issuer,upstream.issuer,database.existingSecretandsigning.existingSecret. - It refuses to render
/admin,/auditor/on the public Ingress. Use the admin Ingress on an internal-only controller. - With the Ingress on, set
rateLimit.trustedProxiesto the ingress controller’s network. The chart refuses to render the Ingress without it, unless rate limiting is off.
Release builds sign the chart and the image with cosign. v0.1 is not released yet, so neither is published.
Sizing
Section titled “Sizing”Measured on one 4-vCPU machine running the control plane, Postgres, Keycloak and the load generator together. Treat the numbers as a floor and re-measure on your own hardware.
| Measure | Value |
|---|---|
| Idle, settled | 16 millicores |
| A token exchange | 5.9 CPU-ms, 21 database statements, 1 ledger row |
| An introspection | 2.7 CPU-ms, no ledger row |
| Cost of one exchange a second | 4.4 millicores of Subact ID, 2.9 of Postgres, 1.1 of the identity provider |
| Ten exchanges a second | 200m CPU and 512 MiB, database included |
| Exchange ceiling | about 350 a second, clean to 300 |
- The ledger is never pruned except by
audit-archive. A row is about 500 bytes. A task writes 1 to 5 rows however often it renews, so budget 0.5 to 2.5 KB per exchange. - Memory does not grow with rate. Overload shedding (
SubactId:Overload) refuses work past the ceiling with503instead of queueing it. - Size every replica for the whole rate, since a survivor carries all the traffic when one fails.
- Keep the connection pool at or below the default of 100. A much larger pool lowered throughput on the test machine.
- Size the rate-limit buckets your traffic spends. A fleet behind one egress address is one source. Each high-risk tool call spends one introspection permit, and a realm-wide sign-out one signal permit per session. Raise the introspection limit with the agents’ limit.
Before real traffic
Section titled “Before real traffic”SubactId:Admin:ApiKeyset, at least 32 characters, not the quickstart’s.- The signing key mounted from a secret store, not set inline.
ASPNETCORE_ENVIRONMENTnotDevelopment, so a missing key stops the server.migraterun with a separate role.high_risk_audiencesset for every audience you would not want reached after a revocation.SubactId.Server doctorrun from where the control plane runs, with no failures.SubactId:RateLimit:PermitsPerMinuteand the introspection limit raised to your fleet, andSubactId:RateLimit:TrustedProxiesset if anything fronts the service. The limiter andSubactId:Overloadleft on.SubactId:Audit:Aggregation:Enabledleft on.audit-verifyon a schedule, its last checkpoint stored away from the database.SubactId:Audit:Retentiondecided, andaudit-archiveon a timer if it is set, with the exports stored away from the database.- Log shipping that keeps
decision: "deny"records.
© 2026 Nikola Živković PR Agencija za programerske usluge Novi Sad. Subact ID is its product.