Embedded Linux ·
An A/B update does not roll back a bad image that hangs
An A/B update scheme doesn't roll back a bad image. It rolls back a bad image that reboots. Those are different guarantees, and the gap between them is where a lot of field-bricked devices come from.
How bootcount rollback works
The usual U-Boot mechanism is bootcount. While upgrade_available is set, U-Boot increments bootcount on every boot. Once it exceeds bootlimit, U-Boot runs altbootcmd instead of bootcmd, and that is where you switch slots.
Userspace is responsible for zeroing bootcount and clearing upgrade_available once it decides the new image is healthy.
Where it breaks: an image that hangs
Now take a kernel that hangs in a driver probe, or an init that deadlocks before your health check ever runs. Nothing reboots. bootcount stays at 1, altbootcmd never fires, and the device sits on the bad slot indefinitely.
The counter only helps if something forces the next boot.
The fix: a hardware watchdog that is already armed
That something is a hardware watchdog that is already armed when the failure happens. Three things to check:
- Arm it in the bootloader, not in userspace, so a hang in early boot is covered.
- If the watchdog is running at boot, the kernel's watchdog core can keep feeding it until userspace opens
/dev/watchdog. Setwatchdog.open_timeout(CONFIG_WATCHDOG_OPEN_TIMEOUT) to a finite value. At 0 it feeds the watchdog forever, which quietly defeats the scheme. - Tie the keepalive to the application. systemd's
RuntimeWatchdogSeconly shows that PID 1 is being scheduled. A per-serviceWatchdogSecwithsd_notifyWATCHDOG=1is what ties a reset to your actual process.
Check that the counter survives the reset
Also confirm that the bootcount storage survives a watchdog reset. If it does not, the count starts again from zero on every watchdog-forced boot and the limit is never reached.
Test the path you are relying on
Does your rollback path have a test that deliberately hangs the kernel and checks that the device ends up on the old slot? That is the test that shows whether rollback works for the failure that matters.