Files
twislaandClaude Opus 5.5 c68741cc46
CI / build (pull_request) Successful in 7m20s
Site / build (pull_request) Successful in 9s
One firmware: the Debug Console in every build, off until switched on, with the device's own token
There is no Debug Build any more (ADR 0010, issue #68, Q188 to Q195). The
console and the test commands are compiled into every firmware. It listens
only while Settings > Debug Console is on, which isn't the default; off,
neither its task nor its 4 KB ring exists. The token is made by the device
and shown on that page; a client proves it knows it by answering a challenge
with an HMAC, so it never crosses the network, and five wrong answers close
the console for a minute. DBG in the Status Bar while it listens.

Over USB serial only: debug on, debug token <value>, debug token new.
scripts/flash.sh --debug uses them to set a device up with the developer's
token. scripts/rdbg.py takes the token from -t, $RORO_DEBUG_TOKEN or the
file, answers the challenge, and fetches a release's ELF to decode a crash.

Gone: the cardputer-adv-debug environment, RORO_DEBUG, the +debug version,
scripts/debug_flags.py, update install ... force, and the rule that a Debug
Build doesn't install releases. Old clients and old firmwares don't talk to
each other.

Against the builds it replaces: 30 KB more flash and 88 bytes more static
RAM than the release, 4 KB less RAM than the Debug Build. 468 host tests.
Checked on the device: off by default, login, the pause after wrong tokens,
Safe Mode with the console, the setting surviving an update, debug off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
2026-10-06 22:59:35 +02:00

5.2 KiB

+++ title = "Crashes, core dumps and Safe Mode" description = "What happens when the firmware crashes: the report, the core dump, the build it is decoded against, the main-loop watchdog and Safe Mode." weight = 5 [extra] tag = "Console" +++

Rollback (see How an update works) protects against new firmware that fails Probation. These measures cover firmware that was confirmed and then crashes: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. They are in every firmware; the Debug Console is what makes them convenient. The decision is ADR 0005.

What a crash leaves behind

  • A core dump in a flash partition: ESP-IDF writes it on a panic.
  • A crash record in NVS, written first thing at the next boot: which version was running (so after a Rollback has switched slots, the firmware still knows which one crashed), and how many starts in a row followed a crash.
  • A summary on the console after the restart (task, program counter, reason, backtrace), and a Notification on screen.

The crash command prints it again whenever you like:

> crash
crash: last one in v0.9.0-1-g4ab873e-dirty (panic)
crash: task loopTask, pc 0x4037e179, cause 0, address 0x00000000
crash: reason: abort() was called at PC 0x421209b3 on core 1
crash: backtrace 0x4037e179 0x4037e141 0x4038582d 0x421209b3 ...
crash: elf sha256 3c6a185e5

coredump erase forgets the dump.

Decode it: rdbg.py crash and rdbg.py coredump

An address is no use without the exact build that crashed. Every build archives its ELF in .pio/elves/, named by version and the first 16 hex digits of its SHA-256 (scripts/version.py); the core dump names the crashed firmware by the same digest. So a crash can be decoded after later builds, including a build you have since replaced:

scripts/rdbg.py crash              # the crash report, with the backtrace turned into functions and source lines
scripts/rdbg.py coredump           # fetch the whole core dump, then decode it
scripts/rdbg.py coredump my.bin    # ...to a file you name
  • crash runs the device's crash, finds the ELF by digest (or by version), and passes the backtrace to scripts/decode_backtrace.sh, which runs addr2line from the build container: one line for each frame, with function and file:line.
  • coredump fetches the raw dump over the console (it works when the main loop is stuck) and runs scripts/decode_coredump.sh: esp-coredump and GDB, giving every task's backtrace, the registers, and the crashed task's stack.
  • By hand: scripts/decode_backtrace.sh <version|digest> <address>..., and scripts/decode_coredump.sh <core.bin> <version|digest>. Both say which ELF they used.
  • A release's ELF is fetched for you. When the crash names a released version (v0.12.0) and no local ELF matches, rdbg.py downloads roro9stack-<version>.elf.gz from the release on Gitea into .pio/elves/, so a crash on a firmware you did not build can be decoded. Decoding itself still runs in the build container.
  • If nothing matches, the scripts say so: an unreleased build made on another machine, or .pio/ was cleaned.

The main loop is watched

Arduino-ESP32 puts only core 0's idle task on the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with a frozen screen and a console that could not run commands. Now a loop stuck for 5 seconds is a panic, with a core dump, counted towards Safe Mode. The rule that follows: nothing in the main loop may block for 5 seconds. Network and card work already run on their own tasks.

An installed update no longer depends on the main loop either: the Update Service restarts into it by itself after 90 seconds.

Safe Mode

The count of starts that follow a crash (a panic or the watchdog) is kept in NVS. After three in a row, the firmware starts Safe Mode instead of everything else:

  • only the clock, Wi-Fi, the Update Service and, if it is switched on, the Debug Console start: no Apps, no IRC, no SD card;
  • the screen says so, with the address to push an update to;
  • only a few commands run (the command reference lists them); anything else answers not available in Safe Mode.

Any normal restart (reboot, an update) or a minute of uptime resets the count. So Safe Mode is reached by crashing three times quickly, and a crash loop costs about half a minute before the device becomes reachable. To leave it: push a fix (scripts/flash.sh --ota), or reboot. Over USB, debug on works in Safe Mode too, if the console was off.

Its limit: if Wi-Fi or the Update Service is what crashes, Safe Mode cannot help, and it takes USB.

Crash on purpose

crash abort      # abort(): a panic with a core dump
crash wdt        # hang the main loop until the task watchdog fires

Use them to see a real report and decode it, to check that the counter reaches Safe Mode, and to prove that a build you are about to rely on will survive its own crashes. The device restarts by itself after either, and the report is there to read when the console comes back.