Files
roro9stack/site/content/dev/debug/crashes.md
T
twislaandClaude Opus 5.5 c68741cc46
CI / build (pull_request) Successful in 7m20s
Site / build (pull_request) Successful in 9s
One firmware: the Debug Console in every build, off until switched on, with the device's own token
There is no Debug Build any more (ADR 0010, issue #68, Q188 to Q195). The
console and the test commands are compiled into every firmware. It listens
only while Settings > Debug Console is on, which isn't the default; off,
neither its task nor its 4 KB ring exists. The token is made by the device
and shown on that page; a client proves it knows it by answering a challenge
with an HMAC, so it never crosses the network, and five wrong answers close
the console for a minute. DBG in the Status Bar while it listens.

Over USB serial only: debug on, debug token <value>, debug token new.
scripts/flash.sh --debug uses them to set a device up with the developer's
token. scripts/rdbg.py takes the token from -t, $RORO_DEBUG_TOKEN or the
file, answers the challenge, and fetches a release's ELF to decode a crash.

Gone: the cardputer-adv-debug environment, RORO_DEBUG, the +debug version,
scripts/debug_flags.py, update install ... force, and the rule that a Debug
Build doesn't install releases. Old clients and old firmwares don't talk to
each other.

Against the builds it replaces: 30 KB more flash and 88 bytes more static
RAM than the release, 4 KB less RAM than the Debug Build. 468 host tests.
Checked on the device: off by default, login, the pause after wrong tokens,
Safe Mode with the console, the setting surviving an update, debug off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
2026-10-06 22:59:35 +02:00

72 lines
5.2 KiB
Markdown

+++
title = "Crashes, core dumps and Safe Mode"
description = "What happens when the firmware crashes: the report, the core dump, the build it is decoded against, the main-loop watchdog and Safe Mode."
weight = 5
[extra]
tag = "Console"
+++
Rollback (see [How an update works](/dev/build/how-an-update-works/)) protects against **new** firmware that fails Probation. These measures cover firmware that was **confirmed and then crashes**: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. They are in every firmware; the Debug Console is what makes them convenient. The decision is [ADR 0005](/dev/decisions/0005-safe-mode-crash-reports-watchdog/).
## What a crash leaves behind
- **A core dump** in a flash partition: ESP-IDF writes it on a panic.
- **A crash record** in NVS, written first thing at the next boot: which **version** was running (so after a Rollback has switched slots, the firmware still knows which one crashed), and how many starts in a row followed a crash.
- **A summary on the console** after the restart (task, program counter, reason, backtrace), and a **Notification** on screen.
The `crash` command prints it again whenever you like:
```
> crash
crash: last one in v0.9.0-1-g4ab873e-dirty (panic)
crash: task loopTask, pc 0x4037e179, cause 0, address 0x00000000
crash: reason: abort() was called at PC 0x421209b3 on core 1
crash: backtrace 0x4037e179 0x4037e141 0x4038582d 0x421209b3 ...
crash: elf sha256 3c6a185e5
```
`coredump erase` forgets the dump.
## Decode it: `rdbg.py crash` and `rdbg.py coredump`
An address is no use without the **exact build** that crashed. Every build archives its ELF in **`.pio/elves/`**, named by version and the first 16 hex digits of its SHA-256 (`scripts/version.py`); the core dump names the crashed firmware by the same digest. So a crash can be decoded **after later builds**, including a build you have since replaced:
```sh
scripts/rdbg.py crash # the crash report, with the backtrace turned into functions and source lines
scripts/rdbg.py coredump # fetch the whole core dump, then decode it
scripts/rdbg.py coredump my.bin # ...to a file you name
```
- **`crash`** runs the device's `crash`, finds the ELF by digest (or by version), and passes the backtrace to `scripts/decode_backtrace.sh`, which runs `addr2line` from the build container: one line for each frame, with function and file:line.
- **`coredump`** fetches the raw dump over the console (it works when the main loop is stuck) and runs `scripts/decode_coredump.sh`: `esp-coredump` and GDB, giving **every task's backtrace, the registers, and the crashed task's stack**.
- By hand: `scripts/decode_backtrace.sh <version|digest> <address>...`, and `scripts/decode_coredump.sh <core.bin> <version|digest>`. Both say which ELF they used.
- **A release's ELF is fetched for you.** When the crash names a released version (`v0.12.0`) and no local ELF matches, `rdbg.py` downloads `roro9stack-<version>.elf.gz` from the release on Gitea into `.pio/elves/`, so a crash on a firmware you did not build can be decoded. Decoding itself still runs in the build container.
- If nothing matches, the scripts say so: an unreleased build made on another machine, or `.pio/` was cleaned.
## The main loop is watched
Arduino-ESP32 puts only core 0's idle task on the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with a frozen screen and a console that could not run commands. Now **a loop stuck for 5 seconds is a panic**, with a core dump, counted towards Safe Mode. The rule that follows: **nothing in the main loop may block for 5 seconds.** Network and card work already run on their own tasks.
An installed update no longer depends on the main loop either: the Update Service restarts into it by itself after 90 seconds.
## Safe Mode
The count of starts that follow a crash (a panic or the watchdog) is kept in NVS. After **three in a row**, the firmware starts **Safe Mode** instead of everything else:
- only the **clock, Wi-Fi, the Update Service and, if it is switched on, the Debug Console** start: no Apps, no IRC, no SD card;
- the screen says so, with the **address to push an update to**;
- only a few commands run (the [command reference](/dev/debug/commands/) lists them); anything else answers `not available in Safe Mode`.
Any **normal restart** (`reboot`, an update) or **a minute of uptime** resets the count. So Safe Mode is reached by crashing three times *quickly*, and a crash loop costs about half a minute before the device becomes reachable. To leave it: push a fix (`scripts/flash.sh --ota`), or `reboot`. Over USB, `debug on` works in Safe Mode too, if the console was off.
**Its limit:** if Wi-Fi or the Update Service is what crashes, Safe Mode cannot help, and it takes USB.
## Crash on purpose
```
crash abort # abort(): a panic with a core dump
crash wdt # hang the main loop until the task watchdog fires
```
Use them to see a real report and decode it, to check that **the counter reaches Safe Mode**, and to prove that a build you are about to rely on will survive its own crashes. The device restarts by itself after either, and the report is there to read when the console comes back.