Public Access
/dev/ has Debug Builds and the Debug Console (builds and the token, the console and its protocol, files and screenshots, driving the UI, crashes and Safe Mode, the command reference), Build, test and release (including how an update works), the architecture decisions and the milestone plans. Generated from the repository by site/tools/gen_dev_docs.py: the ADRs, the milestones, the README's sections, and the command reference, read from the firmware's own `help` text. The pages are committed (Zola cannot read outside its folder); the Site workflow checks they are current, and now also runs when src/main.cpp changes. M0, M1 and CONTEXT.md are not published. README: the gnss commands that the table lacked. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
73 lines
4.8 KiB
Markdown
73 lines
4.8 KiB
Markdown
+++
|
|
title = "Crashes, core dumps and Safe Mode"
|
|
description = "What happens when the firmware crashes: the report, the core dump, the build it is decoded against, the main-loop watchdog and Safe Mode."
|
|
weight = 5
|
|
[extra]
|
|
tag = "Console"
|
|
+++
|
|
|
|
Rollback (see [How an update works](/dev/build/how-an-update-works/)) protects against **new** firmware that fails Probation. These measures cover firmware that was **confirmed and then crashes**: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. They are in **every build**, release and Debug alike; the Debug Console is what makes them convenient. The decision is [ADR 0005](/dev/decisions/0005-safe-mode-crash-reports-watchdog/).
|
|
|
|
## What a crash leaves behind
|
|
|
|
- **A core dump** in a flash partition: ESP-IDF writes it on a panic.
|
|
- **A crash record** in NVS, written first thing at the next boot: which **version** was running (so after a Rollback has switched slots, the firmware still knows which one crashed), and how many starts in a row followed a crash.
|
|
- **A summary on the console** after the restart (task, program counter, reason, backtrace), and a **Notification** on screen.
|
|
|
|
The `crash` command prints it again whenever you like:
|
|
|
|
```
|
|
> crash
|
|
crash: last one in v0.9.0-1-g4ab873e-dirty+debug (panic)
|
|
crash: task loopTask, pc 0x4037e179, cause 0, address 0x00000000
|
|
crash: reason: abort() was called at PC 0x421209b3 on core 1
|
|
crash: backtrace 0x4037e179 0x4037e141 0x4038582d 0x421209b3 ...
|
|
crash: elf sha256 3c6a185e5
|
|
```
|
|
|
|
`coredump erase` forgets the dump.
|
|
|
|
## Decode it: `rdbg.py crash` and `rdbg.py coredump`
|
|
|
|
An address is no use without the **exact build** that crashed. Every build archives its ELF in **`.pio/elves/`**, named by version and the first 16 hex digits of its SHA-256 (`scripts/version.py`); the core dump names the crashed firmware by the same digest. So a crash can be decoded **after later builds**, including a build you have since replaced:
|
|
|
|
```sh
|
|
scripts/rdbg.py crash # the crash report, with the backtrace turned into functions and source lines
|
|
scripts/rdbg.py coredump # fetch the whole core dump, then decode it
|
|
scripts/rdbg.py coredump my.bin # ...to a file you name
|
|
```
|
|
|
|
- **`crash`** runs the device's `crash`, finds the ELF by digest (or by version), and passes the backtrace to `scripts/decode_backtrace.sh`, which runs `addr2line` from the build container: one line for each frame, with function and file:line.
|
|
- **`coredump`** fetches the raw dump over the console (it works when the main loop is stuck) and runs `scripts/decode_coredump.sh`: `esp-coredump` and GDB, giving **every task's backtrace, the registers, and the crashed task's stack**.
|
|
- By hand: `scripts/decode_backtrace.sh <version|digest> <address>...`, and `scripts/decode_coredump.sh <core.bin> <version|digest>`. Both say which ELF they used.
|
|
- If no archived ELF matches, they say so: the build was made on another machine, or `.pio/` was cleaned.
|
|
|
|
## The main loop is watched
|
|
|
|
Arduino-ESP32 puts only core 0's idle task on the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with a frozen screen and a console that could not run commands. Now **a loop stuck for 5 seconds is a panic**, with a core dump, counted towards Safe Mode. The rule that follows: **nothing in the main loop may block for 5 seconds.** Network and card work already run on their own tasks.
|
|
|
|
An installed update no longer depends on the main loop either: the Update Service restarts into it by itself after 90 seconds.
|
|
|
|
## Safe Mode
|
|
|
|
The count of starts that follow a crash (a panic or the watchdog) is kept in NVS. After **three in a row**, the firmware starts **Safe Mode** instead of everything else:
|
|
|
|
- only the **clock, Wi-Fi, the Update Service and, in a Debug Build, the Debug Console** start: no Apps, no IRC, no SD card;
|
|
- the screen says so, with the **address to push an update to**;
|
|
- only a few commands run (the [command reference](/dev/debug/commands/) lists them); anything else answers `not available in Safe Mode`.
|
|
|
|
Any **normal restart** (`reboot`, an update) or **a minute of uptime** resets the count. So Safe Mode is reached by crashing three times *quickly*, and a crash loop costs about half a minute before the device becomes reachable. To leave it: push a fix (`scripts/flash.sh --debug --ota`), or `reboot`.
|
|
|
|
**Its limit:** if Wi-Fi or the Update Service is what crashes, Safe Mode cannot help, and it takes USB.
|
|
|
|
## Crash on purpose
|
|
|
|
On a Debug Build:
|
|
|
|
```
|
|
crash abort # abort(): a panic with a core dump
|
|
crash wdt # hang the main loop until the task watchdog fires
|
|
```
|
|
|
|
Use them to see a real report and decode it, to check that **the counter reaches Safe Mode**, and to prove that a Debug Build you are about to rely on will survive its own crashes. The device restarts by itself after either, and the report is there to read when the console comes back.
|