Restoring the PowerShell helper put it in update.go, which is shared —
so a Linux build would have tried to run `powershell` to relaunch
itself. It compiled and vetted cleanly for linux/amd64, which is exactly
why it needed catching before somebody built it: the fault only shows on
a real update, on a machine that has no PowerShell.
scheduleRelaunch now lives in the platform files. Windows keeps the
helper. Linux starts the new binary directly, which is right there and
not a compromise: nothing holds an executable open while it runs, so the
swap has already succeeded, and there is no mutex to race — the
single-instance guard is an flock the dying process releases as it
exits, and the new one waits for our pid first.
The two guards were looking at the old location and had to follow: the
Wait-Process/Start-Process check moves into relaunch_windows_test.go
where it belongs, and TestEveryRelaunchPassesItsPid now scans
updateswap_linux.go too — the direct spawn moved there, and without it
the test would have gone quiet again.
Checked from Windows, as BUILDING-LINUX.md says is done at every
release: GOOS=linux go build ./... and go vet ./... both clean.
Operators are still losing the relaunch, so I went and read the code
from before 0430aab — the commit that removed the helper — instead of
theorising again. It was:
Wait-Process -Id <pid>; Start-Sleep -Milliseconds 400;
Start-Process -FilePath <exe> -ArgumentList '--post-update'
Two things in there that the direct launch never had:
1. Start-Process made the new OpsLog a child of PowerShell, which then
exited — so it was detached. The direct launch makes it a child of
the instance that is dying, in the same process group and console.
DETACHED_PROCESS | CREATE_NEW_PROCESS_GROUP puts it back on its own,
so nothing aimed at the old process can reach the new one.
2. It slept 400 ms AFTER the old process was gone, before starting
anything. A process's handles are released by the kernel as it dies,
so the mutex is free the moment the wait returns — but the things
around it are not on that clock: the WebView2 user-data lock, the log
file, an antivirus that woke up when the exe was replaced. This is
not a theory about which of those it was; it is the pause being put
back where it was.
Not restored: the PowerShell itself. Defender removed 0.27.14 from a
station as Trojan:Script/Wacatac.H!ml, and "Script/" was that helper —
an unsigned binary replacing itself and spawning a windowless script to
start another executable is, byte for byte, a dropper. Bringing it back
trades this fault for one that deletes the program.
Also: the relaunch now logs the child's pid, which the new instance's
startup.log already records on the other side. Without the pair there is
no telling "the new instance never started" from "it started and gave up
waiting", and those have different causes.