Public Access
There is no Debug Build any more (ADR 0010, issue #68, Q188 to Q195). The console and the test commands are compiled into every firmware. It listens only while Settings > Debug Console is on, which isn't the default; off, neither its task nor its 4 KB ring exists. The token is made by the device and shown on that page; a client proves it knows it by answering a challenge with an HMAC, so it never crosses the network, and five wrong answers close the console for a minute. DBG in the Status Bar while it listens. Over USB serial only: debug on, debug token <value>, debug token new. scripts/flash.sh --debug uses them to set a device up with the developer's token. scripts/rdbg.py takes the token from -t, $RORO_DEBUG_TOKEN or the file, answers the challenge, and fetches a release's ELF to decode a crash. Gone: the cardputer-adv-debug environment, RORO_DEBUG, the +debug version, scripts/debug_flags.py, update install ... force, and the rule that a Debug Build doesn't install releases. Old clients and old firmwares don't talk to each other. Against the builds it replaces: 30 KB more flash and 88 bytes more static RAM than the release, 4 KB less RAM than the Debug Build. 468 host tests. Checked on the device: off by default, login, the pause after wrong tokens, Safe Mode with the console, the setting surviving an update, debug off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
14 lines
2.2 KiB
Markdown
14 lines
2.2 KiB
Markdown
# Safe Mode, crash reports and a watched main loop, in every build
|
|
|
|
Rollback protects against new firmware that fails Probation. It does nothing for firmware that was confirmed and crashes later: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. Without a cable, such a device would restart forever. Three measures, in every build, keep it reachable:
|
|
|
|
- **Safe Mode.** The firmware counts starts that follow a crash (panic or watchdog) in NVS, first thing at boot. After 3 in a row, it starts only the clock, Wi-Fi, the Update Service and, if it's switched on, the Debug Console: no Apps, no IRC, no SD card, and a screen that says so with the address to push an update to. Any normal restart (a `reboot`, an update), or a minute of uptime, resets the count.
|
|
- **Crash reports.** The same boot record keeps which version was running, so after a crash the firmware knows which one crashed, even when a Rollback has switched slots since. ESP-IDF already writes a core dump to its flash partition on a panic; after the restart the firmware prints its summary (task, PC, reason, backtrace) and raises a Notification. The `crash` command shows it again later. With the Debug Console, `scripts/rdbg.py crash` decodes the backtrace and `scripts/rdbg.py coredump` fetches the whole dump for `esp-coredump`, against the ELF of that exact build (`.pio/elves/`, named by version and ELF digest).
|
|
- **The main loop is watched.** Arduino-ESP32 subscribes only core 0's idle task to the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with the screen frozen and the Debug Console unable to run commands. `enableLoopWDT()` makes a loop stuck for 5 s a panic, with a core dump, counted towards Safe Mode. And an installed update no longer depends on the main loop: the Update Service restarts into it by itself after 90 s.
|
|
|
|
## Consequences
|
|
|
|
- Nothing in the main loop may block for 5 s. Network and card work already run on their own tasks.
|
|
- Safe Mode can't help when Wi-Fi or the Update Service itself is what crashes; that still needs USB.
|
|
- Three crashes within a minute of each restart are needed to reach Safe Mode, so a crash loop costs about half a minute before the device becomes reachable.
|