fix: ACME ownership self-heal + daily timer, apply-all force, firewall baseline re-stamp

acme:
- acme.sh chmods its tree to owner-only (700/600) every run, which
  broke the two-user model: a tree left owner-only by one user made
  every acme.sh call of the other exit 2
- normalize_acme_home() reopens group access (sudo chmod g+rwX,
  files only — setgid dirs trip RestrictSUIDSGID); _run_acme_preflight
  is the choke point before every daemon acme.sh call + startup
- acme service now runs as the daemon user; --log persists the raw CA
  transcript; SYS_LOG=6 journals manual issue/renew runs
- timer daily-only: two runs/day landed inside ZeroSSL's 24h
  validation backoff (Retry-After: 86400) — a permanent renewal lockout
- _collect_acme no longer raises on cert-list failure; reports
  status.error (AcmeState.status) so the certs page can surface it

firewall: re-stamp the applied baseline on live zone mutations
(interfaces/services/rich-rules/masquerade/forward-ports) so cancel-all
reverts to post-mutation state, not a stale install-era snapshot;
set_masquerade syncs the declarative config for existing zones;
add_forward_port records toaddr only with toport

status: apply-all accepts {"force": true} (forwarded to the firewall
apply only); ApplyConfirm force checkbox; applyResultToasts() — the
errors map wins over the 200; ActionButton checks errors before the
success toast; dashboard uses ApplyConfirm

system_import: drift re-imports carry the existing apply-meta; first
import stamps the adopted content as applied (it is the running state)
— no phantom pending changes

nginx: get_config only re-saves when migration actually changed the
config (no more owner/mtime churn on every read)

install: repair mis-owned top-level system dirs (tmpfiles
unsafe-path-transition), warn with a full-repair command for deeper
mis-ownership

daemon/server: loop.get_exception_handler() (aiohttp API fix)

tests: 888 pytest + 24 node passing; ruff clean
This commit is contained in:
2026-09-01 02:35:04 +00:00
parent ac52918df5
commit 75b86fd60d
30 changed files with 738 additions and 53 deletions
+17 -4
View File
@@ -1976,6 +1976,16 @@ POST /api/status/apply-all
Apply pending changes for all subsystems in dependency order.
**Request Body (optional):**
```json
{ "force": true }
```
`force` is forwarded to the firewall apply only — it overrides the
management-lockout and interface-coverage guards. Other subsystems
ignore it.
**Response (`data`):**
| Field | Type | Description |
@@ -1983,10 +1993,13 @@ Apply pending changes for all subsystems in dependency order.
| `applied` | `[string, ...]` | List of subsystems that were applied |
| `errors` | `object` | Map of subsystem label → error message |
The firewall apply runs with `force=false`, so if a firewall interface
would be left without zone coverage (the coverage guard), a
`ConflictError` surfaces in `errors` under `"Firewall"` while the other
subsystems proceed — the desired no-silent-apply behavior.
The endpoint returns `200` even when some subsystems failed — per-subsystem
failures are reported in `errors`, so clients must check `errors` (not just
the HTTP status) before reporting success. Without `force`, the firewall
apply refuses if an interface would be left without zone coverage (the
coverage guard) or both https/ssh would be stripped from the default zone
(lockout guard); the `ConflictError` surfaces in `errors` under
`"Firewall"` while the other subsystems proceed.
---