2 Commits
Author SHA1 Message Date
twislaandClaude Opus 5.5 c68741cc46 One firmware: the Debug Console in every build, off until switched on, with the device's own token
Site / build (pull_request) Successful in 9s
CI / build (pull_request) Successful in 7m20s
There is no Debug Build any more (ADR 0010, issue #68, Q188 to Q195). The
console and the test commands are compiled into every firmware. It listens
only while Settings > Debug Console is on, which isn't the default; off,
neither its task nor its 4 KB ring exists. The token is made by the device
and shown on that page; a client proves it knows it by answering a challenge
with an HMAC, so it never crosses the network, and five wrong answers close
the console for a minute. DBG in the Status Bar while it listens.

Over USB serial only: debug on, debug token <value>, debug token new.
scripts/flash.sh --debug uses them to set a device up with the developer's
token. scripts/rdbg.py takes the token from -t, $RORO_DEBUG_TOKEN or the
file, answers the challenge, and fetches a release's ELF to decode a crash.

Gone: the cardputer-adv-debug environment, RORO_DEBUG, the +debug version,
scripts/debug_flags.py, update install ... force, and the rule that a Debug
Build doesn't install releases. Old clients and old firmwares don't talk to
each other.

Against the builds it replaces: 30 KB more flash and 88 bytes more static
RAM than the release, 4 KB less RAM than the Debug Build. 468 host tests.
Checked on the device: off by default, login, the pause after wrong tokens,
Safe Mode with the console, the setting surviving an update, debug off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
2026-10-06 22:59:35 +02:00
twislaandClaude Opus 5.5 14ff13f634 Safe Mode, crash reports, and a watched main loop
Every build now records at boot which version runs and, after a crash
restart, which one crashed (even across a Rollback). The core dump
summary (task, PC, reason, backtrace) is printed and raised as a
Notification; `crash` shows it later. After 3 crash restarts in a row
the firmware starts in Safe Mode: clock, Wi-Fi, Update Service and Debug
Console only (SafeMode, 2 host tests). A normal restart or a minute up
resets the count.

The main loop is now on the task watchdog (enableLoopWDT): Arduino only
watched core 0's idle task, so a stuck loop hung the device for good.
The Update Service restarts into an installed update by itself if the
main loop hasn't after 90 s.

Debug Builds: `coredump get` and `reset` are answered by the console's
own task; rdbg.py crash decodes the backtrace and rdbg.py coredump runs
esp-coredump, against ELFs archived by version and digest in .pio/elves.

The StorageService mutex is now made in the constructor: Safe Mode never
starts that Service, and `info` crashed on the null mutex, 29 times in a
row before the fix was pushed into Safe Mode over Wi-Fi.

Verified on the device: crash report and full core dump decoded over
Wi-Fi; Safe Mode at exactly 3 crashes, left by `reboot`; a hung loop
caught by the watchdog in 5 s; `reset` from the console task. ADR 0005.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EhqxQ49eCju4CzKYNjZzwT
2026-10-04 03:07:56 +02:00