You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(distributed): make backend upgrade actually re-install on workers (#9708)
* fix(distributed): make backend upgrade actually re-install on workers
UpgradeBackend dispatched a vanilla backend.install NATS event to every
node hosting the backend. The worker's installBackend short-circuits on
"already running for this (model, replica) slot" and returns the
existing address — so the gallery install path was skipped, no artifact
was re-downloaded, no metadata was written. The frontend's drift
detection then re-flagged the same backends every cycle (installedDigest
stays empty → mismatch → "Backend upgrade available (new build)") while
"Backend upgraded successfully" landed in the logs at the same time.
The user-visible symptom: clicking "Upgrade All" silently does nothing
and the same N backends sit on the upgrade list forever.
Two coupled fixes, one PR:
1. Force flag on backend.install. Add `Force bool` to
BackendInstallRequest and thread it through NodeCommandSender ->
RemoteUnloaderAdapter. UpgradeBackend (and the reconciler's pending-op
drain when retrying an upgrade) sets force=true; routine load events
and admin install endpoints keep force=false. On the worker, force=true
stops every live process that uses this backend (resolveProcessKeys
for peer replicas, plus the exact request processKey), skips the
findBackend short-circuit, and passes force=true into
gallery.InstallBackendFromGallery so the on-disk artifact is
overwritten. After the gallery install completes, startBackend brings
up a fresh process at the same processKey on a new port.
2. Liveness check on the fast path. installBackend's "already running"
branch read getAddr without verifying the process was alive, so a
gRPC backend that died without the supervisor noticing left a stale
(key, addr) entry. The reconciler then dialed that address, got
ECONNREFUSED, marked the replica failed, retried install — and the
supervisor said "already running addr=…" again. Loop forever, exactly
what we observed on a node whose llama-cpp process had died but whose
supervisor record persisted. Verify s.isRunning(processKey) before
trusting getAddr; if the entry is stale, stopBackendExact cleans up
and we fall through to a real install.
Backwards-compatible: the new Force field is omitempty, older workers
ignore it (their default behavior matches force=false). The signature
change on NodeCommandSender.InstallBackend is internal-only.
Verified: unit tests in core/services/nodes pass (108s suite). The
pre-existing core/backend build break (proto regen pending for
word-level timestamps) blocks core/cli and core/http/endpoints/localai
package tests but is unrelated to this change.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-7 [Claude Code]
* test(e2e/distributed): pass force=false to adapter.InstallBackend
NodeCommandSender.InstallBackend gained a final force bool in the
upgrade-force commit; the e2e distributed lifecycle tests still called
the old 8-arg signature and broke compilation. These tests exercise the
routine install path (single replica, default behavior), so force=false
preserves their existing semantics.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-7 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
0 commit comments