Safe Mode, crash reports, and a watched main loop

Every build now records at boot which version runs and, after a crash
restart, which one crashed (even across a Rollback). The core dump
summary (task, PC, reason, backtrace) is printed and raised as a
Notification; `crash` shows it later. After 3 crash restarts in a row
the firmware starts in Safe Mode: clock, Wi-Fi, Update Service and Debug
Console only (SafeMode, 2 host tests). A normal restart or a minute up
resets the count.

The main loop is now on the task watchdog (enableLoopWDT): Arduino only
watched core 0's idle task, so a stuck loop hung the device for good.
The Update Service restarts into an installed update by itself if the
main loop hasn't after 90 s.

Debug Builds: `coredump get` and `reset` are answered by the console's
own task; rdbg.py crash decodes the backtrace and rdbg.py coredump runs
esp-coredump, against ELFs archived by version and digest in .pio/elves.

The StorageService mutex is now made in the constructor: Safe Mode never
starts that Service, and `info` crashed on the null mutex, 29 times in a
row before the fix was pushed into Safe Mode over Wi-Fi.

Verified on the device: crash report and full core dump decoded over
Wi-Fi; Safe Mode at exactly 3 crashes, left by `reboot`; a hung loop
caught by the watchdog in 5 s; `reset` from the console task. ADR 0005.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
This commit is contained in:
2026-10-04 03:07:56 +02:00
co-authored by Claude Opus 5.5
parent 0fb7f4e9d5
commit 14ff13f634
16 changed files with 452 additions and 21 deletions
@@ -0,0 +1,13 @@
# Safe Mode, crash reports and a watched main loop, in every build
Rollback protects against new firmware that fails Probation. It does nothing for firmware that was confirmed and crashes later: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. Without a cable, such a device would restart forever. Three measures, in release and Debug Builds alike, keep it reachable:
- **Safe Mode.** The firmware counts starts that follow a crash (panic or watchdog) in NVS, first thing at boot. After 3 in a row, it starts only the clock, Wi-Fi, the Update Service and, in a Debug Build, the Debug Console: no Apps, no IRC, no SD card, and a screen that says so with the address to push an update to. Any normal restart (a `reboot`, an update), or a minute of uptime, resets the count.
- **Crash reports.** The same boot record keeps which version was running, so after a crash the firmware knows which one crashed, even when a Rollback has switched slots since. ESP-IDF already writes a core dump to its flash partition on a panic; after the restart the firmware prints its summary (task, PC, reason, backtrace) and raises a Notification. The `crash` command shows it again later. In a Debug Build, `scripts/rdbg.py crash` decodes the backtrace and `scripts/rdbg.py coredump` fetches the whole dump for `esp-coredump`, against the ELF of that exact build (`.pio/elves/`, named by version and ELF digest).
- **The main loop is watched.** Arduino-ESP32 subscribes only core 0's idle task to the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with the screen frozen and the Debug Console unable to run commands. `enableLoopWDT()` makes a loop stuck for 5 s a panic, with a core dump, counted towards Safe Mode. And an installed update no longer depends on the main loop: the Update Service restarts into it by itself after 90 s.
## Consequences
- Nothing in the main loop may block for 5 s. Network and card work already run on their own tasks.
- Safe Mode can't help when Wi-Fi or the Update Service itself is what crashes; that still needs USB.
- Three crashes within a minute of each restart are needed to reach Safe Mode, so a crash loop costs about half a minute before the device becomes reachable.