Files
roro9stack/site/content/dev/debug/crashes.md
T
twislaandClaude Sonnet 5.5 3b4100dc6a
CI / build (pull_request) Successful in 8m38s
Site / build (pull_request) Successful in 9s
Site: the developer docs (phase 4), with the Debug Builds and the Debug Console first
/dev/ has Debug Builds and the Debug Console (builds and the token, the
console and its protocol, files and screenshots, driving the UI, crashes and
Safe Mode, the command reference), Build, test and release (including how an
update works), the architecture decisions and the milestone plans.

Generated from the repository by site/tools/gen_dev_docs.py: the ADRs, the
milestones, the README's sections, and the command reference, read from the
firmware's own `help` text. The pages are committed (Zola cannot read outside
its folder); the Site workflow checks they are current, and now also runs
when src/main.cpp changes. M0, M1 and CONTEXT.md are not published.
README: the gnss commands that the table lacked.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
2026-10-06 21:25:33 +02:00

4.8 KiB

+++ title = "Crashes, core dumps and Safe Mode" description = "What happens when the firmware crashes: the report, the core dump, the build it is decoded against, the main-loop watchdog and Safe Mode." weight = 5 [extra] tag = "Console" +++

Rollback (see How an update works) protects against new firmware that fails Probation. These measures cover firmware that was confirmed and then crashes: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. They are in every build, release and Debug alike; the Debug Console is what makes them convenient. The decision is ADR 0005.

What a crash leaves behind

  • A core dump in a flash partition: ESP-IDF writes it on a panic.
  • A crash record in NVS, written first thing at the next boot: which version was running (so after a Rollback has switched slots, the firmware still knows which one crashed), and how many starts in a row followed a crash.
  • A summary on the console after the restart (task, program counter, reason, backtrace), and a Notification on screen.

The crash command prints it again whenever you like:

> crash
crash: last one in v0.9.0-1-g4ab873e-dirty+debug (panic)
crash: task loopTask, pc 0x4037e179, cause 0, address 0x00000000
crash: reason: abort() was called at PC 0x421209b3 on core 1
crash: backtrace 0x4037e179 0x4037e141 0x4038582d 0x421209b3 ...
crash: elf sha256 3c6a185e5

coredump erase forgets the dump.

Decode it: rdbg.py crash and rdbg.py coredump

An address is no use without the exact build that crashed. Every build archives its ELF in .pio/elves/, named by version and the first 16 hex digits of its SHA-256 (scripts/version.py); the core dump names the crashed firmware by the same digest. So a crash can be decoded after later builds, including a build you have since replaced:

scripts/rdbg.py crash              # the crash report, with the backtrace turned into functions and source lines
scripts/rdbg.py coredump           # fetch the whole core dump, then decode it
scripts/rdbg.py coredump my.bin    # ...to a file you name
  • crash runs the device's crash, finds the ELF by digest (or by version), and passes the backtrace to scripts/decode_backtrace.sh, which runs addr2line from the build container: one line for each frame, with function and file:line.
  • coredump fetches the raw dump over the console (it works when the main loop is stuck) and runs scripts/decode_coredump.sh: esp-coredump and GDB, giving every task's backtrace, the registers, and the crashed task's stack.
  • By hand: scripts/decode_backtrace.sh <version|digest> <address>..., and scripts/decode_coredump.sh <core.bin> <version|digest>. Both say which ELF they used.
  • If no archived ELF matches, they say so: the build was made on another machine, or .pio/ was cleaned.

The main loop is watched

Arduino-ESP32 puts only core 0's idle task on the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with a frozen screen and a console that could not run commands. Now a loop stuck for 5 seconds is a panic, with a core dump, counted towards Safe Mode. The rule that follows: nothing in the main loop may block for 5 seconds. Network and card work already run on their own tasks.

An installed update no longer depends on the main loop either: the Update Service restarts into it by itself after 90 seconds.

Safe Mode

The count of starts that follow a crash (a panic or the watchdog) is kept in NVS. After three in a row, the firmware starts Safe Mode instead of everything else:

  • only the clock, Wi-Fi, the Update Service and, in a Debug Build, the Debug Console start: no Apps, no IRC, no SD card;
  • the screen says so, with the address to push an update to;
  • only a few commands run (the command reference lists them); anything else answers not available in Safe Mode.

Any normal restart (reboot, an update) or a minute of uptime resets the count. So Safe Mode is reached by crashing three times quickly, and a crash loop costs about half a minute before the device becomes reachable. To leave it: push a fix (scripts/flash.sh --debug --ota), or reboot.

Its limit: if Wi-Fi or the Update Service is what crashes, Safe Mode cannot help, and it takes USB.

Crash on purpose

On a Debug Build:

crash abort      # abort(): a panic with a core dump
crash wdt        # hang the main loop until the task watchdog fires

Use them to see a real report and decode it, to check that the counter reaches Safe Mode, and to prove that a Debug Build you are about to rely on will survive its own crashes. The device restarts by itself after either, and the report is there to read when the console comes back.