p4: an unreachable C6 must not take the device down with it

beta_86 flipped the T-Display P4's default C6 backend to esp-hosted. Current
V1 units ship the C6 with esp-hosted firmware and are fine; OLDER units still
carry the factory ESP-AT firmware, and on those the SDIO stream never frames.
That turned into a board that will not boot at all, reported by three people
so far (#600 jalevdev, myshoeisonfire on Discord, and Kaj's own unit), each
with the same signature:

  H_SDIO_DRV: len[25970]>max[1524] OR offset[25697] != exp[12], Drop

Two separate things made a co-processor mismatch fatal, and both had to go.

1. The UI was behind the C6. hostedInitBLE() ran unconditionally at setup, a
   hosted handshake, with ui_task.begin() ~80 lines later. The LEGACY backend
   always started its worker AFTER the UI, which is exactly why a dead C6 used
   to cost Wi-Fi and nothing else; the hosted path was wired in ahead of it.
   Both backends now start at the end of setup, after the UI, with a note
   saying nothing C6-dependent may move above it.

2. esp-hosted resets the host. CONFIG_ESP_HOSTED_HOST_RESTART_NO_COMMUNICATION_
   WITH_SLAVE=n covers one path; the one that actually fires is
   H_TRANSPORT_RESTART_ON_FAILURE in sdio_drv.c, about 10 s after the connect
   attempt, and it is on by default. Now off. The failure still arrives as
   ESP_HOSTED_EVENT_TRANSPORT_FAILURE, which is what an application should act
   on instead of the transport rebooting the board underneath it.

Measured on an ESP-AT unit, 45 s captures, same hardware:

  as shipped       5 resets, host restarts, UI never appears
  reorder only     2 resets, host restarts, UI appears then dies
  reorder + config 0 resets, 0 restarts, UI stays up

The C6 is still unreachable on that unit, as it must be - it is the wrong
firmware for it. It now costs Wi-Fi and BLE and nothing else.

Also here, both needed before this can ship:

* The ESP-AT build gets its own artifact name (tdisplay-p4-at, and the matching
  -lcd/-lr2021 combinations). Both backends used to answer "tdisplay-p4", so an
  ESP-AT unit's self-update would have fetched the hosted bin and put itself
  straight back into the loop, recoverable only over USB. That is the T-Deck
  Max bug again (a Max self-updating to Pro firmware) and would have been the
  fourth time this table shipped wrong.

* The keyboard expansion probe prints a one-shot bus scan. "It does not work"
  has three causes that are indistinguishable from outside: not clipped on,
  clipped on but unpowered, or present and failing bring-up. An unpowered board
  does not NACK cleanly either - its pull-ups are dead, the lines sit low
  through its ESD diodes, and every transaction returns ESP_ERR_INVALID_STATE,
  which reads like a driver fault. The scan says which it is in one line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kaj Schittecat
2026-10-05 10:03:19 +02:00
co-authored by Claude Opus 5
parent 66e1f37d2c
commit 7ce1b3f48c
4 changed files with 110 additions and 42 deletions
+17 -4
View File
@@ -32231,17 +32231,30 @@ static int wifiScanWatchdogSafe(uint32_t cap_ms, uint16_t per_chan_ms = 300) {
#elif defined(HAS_TDECK_PRO)
#define WADA_BOARD_ID "tdeck-pro-v1-1"
#elif defined(HAS_TDISPLAY_P4)
// The C6 backend is part of the artifact identity, not just a build flavour. An
// older unit runs factory ESP-AT and MUST be built with WADA_P4_LEGACY_AT=1; the
// esp-hosted image boot-loops it (the SDIO stream never frames, and the transport
// resets the host every ~10 s). Both backends used to answer "tdisplay-p4" here,
// so an ESP-AT unit's self-update would fetch the hosted bin and put itself back
// into that loop with no way out but USB. That is precisely the T-Deck Max bug
// (a Max self-updating to Pro firmware), and this would have been the fourth time
// this table shipped wrong. Give the ESP-AT build its own name.
#if TDP4_C6_HOSTED
#define TDP4_C6_ID ""
#else
#define TDP4_C6_ID "-at"
#endif
#if defined(HAS_TDP4_LCD)
#if defined(USE_LR2021)
#define WADA_BOARD_ID "tdisplay-p4-lcd-lr2021"
#define WADA_BOARD_ID "tdisplay-p4-lcd-lr2021" TDP4_C6_ID
#else
#define WADA_BOARD_ID "tdisplay-p4-lcd"
#define WADA_BOARD_ID "tdisplay-p4-lcd" TDP4_C6_ID
#endif
#else
#if defined(USE_LR2021)
#define WADA_BOARD_ID "tdisplay-p4-lr2021"
#define WADA_BOARD_ID "tdisplay-p4-lr2021" TDP4_C6_ID
#else
#define WADA_BOARD_ID "tdisplay-p4"
#define WADA_BOARD_ID "tdisplay-p4" TDP4_C6_ID
#endif
#endif
#elif defined(HAS_TANMATSU)
+58 -38
View File
@@ -53,13 +53,23 @@
#endif
// --- C6 / esp-hosted bring-up gate -----------------------------------------------------------------
// The ESP32-C6 (Wi-Fi/BLE co-processor over SDIO) doesn't come up yet: the SDIO bus inits, but the
// C6's esp-hosted slave never signals ready, so esp_hosted_connect_to_slave() times out (~12 s) and
// esp-hosted resets the P4 — a boot loop that never reaches ui_task.begin(), leaving the AMOLED dark.
// Until the C6 reset/power/slave-firmware path is proven (needs a reset callback via the XL9535 — see
// TDISPLAY_P4_PORT.md), gate ALL C6-dependent bring-up OFF so the display + touch + LoRa come up.
// LoRa is a raw SX1262/LR2021 on P4 GPIOs, independent of the C6, so the mesh still works over USB/LoRa.
// Flip to 1 once the C6 link is solid to restore Wi-Fi + BLE.
// TWO HARDWARE GENERATIONS, and the image has to match the one in front of it. Current T-Display P4
// V1 units ship the C6 with esp-hosted firmware (2.12.3) and are what TDP4_C6_HOSTED=1 (the default)
// is for. OLDER units still carry the factory ESP-AT firmware and need WADA_P4_LEGACY_AT=1. Flash a
// hosted image onto an ESP-AT unit and the SDIO stream never frames — the log fills with
// "H_SDIO_DRV: len[...]>max[1524] ... Drop" and Wi-Fi/BLE simply never arrive.
//
// That mismatch must cost Wi-Fi and BLE and NOTHING ELSE. All C6-dependent bring-up therefore runs
// at the END of setup, after ui_task.begin() — see the block there. LoRa is a raw SX1262/LR2021 on
// P4 GPIOs, independent of the C6, so the mesh still works over USB/LoRa on a mismatched unit.
//
// esp-hosted RESETS THE HOST on a transport failure, which is what turns a mismatch into a boot
// loop rather than a missing feature. There are two reset paths and they need separate switches:
// CONFIG_ESP_HOSTED_HOST_RESTART_NO_COMMUNICATION_WITH_SLAVE=n covers one, and
// CONFIG_ESP_HOSTED_TRANSPORT_RESTART_ON_FAILURE=n covers the one that actually fires here, about
// 10 s after the connect attempt begins (sdio_drv.c, H_TRANSPORT_RESTART_ON_FAILURE). Both are off
// in sdkconfigs/tdisplay_p4. Measured on an ESP-AT unit before the second one was set: five resets
// in a 45 s capture, each rst:0x9 HP_SYS_LP_WDT_RESET. Do not re-enable either.
#define TDP4_C6_READY TDP4_C6_HOSTED
// BLE over the C6 uses arduino-esp32's hostedInitBLE(), which only exists in arduino-esp32 >=3.3.10.
// We pin 3.3.0 (so esp_hosted can be the C6-compatible 2.0.17), which predates hostedInitBLE — so BLE
@@ -147,10 +157,10 @@ static void wadameshSetup() {
esp_netif_init();
esp_event_loop_create_default();
// 4. C6: runs the FACTORY ESP-AT firmware (never esp-hosted, never reflashed — Kaj's product
// decision 2026-07-15). Wi-Fi comes up via the c6_at AT-over-SDIO worker spawned at the end of
// setup; BLE companion is unavailable on this AT build (advertising commands stubbed) so the
// phone pairs over Wi-Fi (TCP:5000) or USB.
// 4. C6: which firmware it runs depends on the unit's age — see the generation note at the top.
// Either way nothing happens HERE: both backends are started at the end of setup, after the
// UI. On the ESP-AT build the c6_at AT-over-SDIO worker carries Wi-Fi and the BLE companion is
// unavailable (advertising commands stubbed), so the phone pairs over Wi-Fi (TCP:5000) or USB.
#if TDP4_C6_HOSTED
printf("[BOOT] C6 = esp-hosted (lazy Wi-Fi/BLE initialization)\n");
#else
@@ -298,39 +308,13 @@ static void wadameshSetup() {
the_mesh.begin(disp != NULL);
#if TDP4_C6_HOSTED
// Cold-boot clock sync over SAVED Wi-Fi (#383, BootTimeSync.h) -- the same opt-in
// Settings > Clock switch the T-Deck/M9 have. Same place in the sequence as the S3
// main.cpp: before the TCP/WS/BLE transports exist, so its temporary Wi-Fi session has
// nothing live to disturb. Returns Skipped at no cost unless this was a true power-on,
// the user turned it on, and the PCF8563 did not already give a trustworthy time.
{
wifiConfigBegin();
uint32_t synced_epoch = 0;
const BootTimeSyncResult r = bootTimeSyncRun(rtc_clock.timeIsCurrent(), synced_epoch);
if (r == BootTimeSyncResult::Ok) {
rtc_clock.setCurrentTime(synced_epoch); // floor + system clock + RTC chip, like an NTP sync
printf("[BOOT] cold-boot time sync ok: %lu\n", (unsigned long)synced_epoch);
} else if (r != BootTimeSyncResult::Skipped) {
printf("[BOOT] cold-boot time sync: %s\n", bootTimeSyncResultName(r));
}
}
wifiConfigBegin(); // NVS only, no radio — safe this early, and the getters assume it
#endif
serial_interface.begin(Serial, TCP_PORT, WS_PORT);
serial_interface.setBroadcastResponses(true);
the_mesh.startInterface(serial_interface);
#if defined(BLE_PIN_CODE) && TDP4_C6_READY && TDP4_BLE_READY
if (hostedInitBLE()) {
char* nm = the_mesh.getNodePrefs()->node_name;
serial_interface.prepareBle("wadamesh-", nm, the_mesh.getBLEPin());
if (wifiConfigGetBleEnabled())
serial_interface.beginBle("wadamesh-", nm, the_mesh.getBLEPin());
} else {
printf("[BOOT] hostedInitBLE FAILED\n");
}
#endif
// GPS UART resilience (same fix as the S3 boards): the core opens Serial1 with Arduino's
// 256-byte RX ring; one long LVGL frame overflows it and corrupts NMEA, stretching TTFF from
// ~1 min to many. Must precede sensors.begin() (a no-op once the UART runs). Unconditional
@@ -414,9 +398,45 @@ static void wadameshSetup() {
printf("[BOOT] setup done (helper)\n");
return;
#endif
// ---- C6 bring-up: AFTER the UI, for BOTH backends --------------------------------------------
// Nothing below here may run before ui_task.begin(). The C6 is a separate chip that can be
// unreachable — an older unit still carries the factory ESP-AT firmware while this build speaks
// esp-hosted, and then the SDIO stream never frames (H_SDIO_DRV drop storm). The legacy backend
// always started its worker here, after the UI, so a dead C6 cost you Wi-Fi and nothing else.
// The hosted backend was wired in ahead of the UI instead, so the same dead C6 held the AMOLED
// dark through a handshake that could not succeed, and the board read as bricked. Both now start
// here. A co-processor must never be able to take the display, touch and LoRa down with it.
#if !TDP4_C6_HOSTED
// Legacy factory ESP-AT units use the asynchronous AT worker.
c6at_worker_start();
#else
#if defined(BLE_PIN_CODE) && TDP4_C6_READY && TDP4_BLE_READY
if (hostedInitBLE()) {
char* nm = the_mesh.getNodePrefs()->node_name;
serial_interface.prepareBle("wadamesh-", nm, the_mesh.getBLEPin());
if (wifiConfigGetBleEnabled())
serial_interface.beginBle("wadamesh-", nm, the_mesh.getBLEPin());
} else {
printf("[BOOT] hostedInitBLE FAILED\n");
}
#endif
// Cold-boot clock sync over SAVED Wi-Fi (#383, BootTimeSync.h) -- the same opt-in
// Settings > Clock switch the T-Deck/M9 have. Returns Skipped at no cost unless this was a
// true power-on, the user turned it on, and the PCF8563 did not already give a trustworthy
// time. It used to sit before the transports so its temporary Wi-Fi session had nothing live
// to disturb; from here the transports exist but none of them is associated yet, and the UI
// being up first is worth more than that ordering nicety. The clock simply corrects itself a
// moment after the screen appears instead of the screen waiting on the network.
{
uint32_t synced_epoch = 0;
const BootTimeSyncResult r = bootTimeSyncRun(rtc_clock.timeIsCurrent(), synced_epoch);
if (r == BootTimeSyncResult::Ok) {
rtc_clock.setCurrentTime(synced_epoch); // floor + system clock + RTC chip, like an NTP sync
printf("[BOOT] cold-boot time sync ok: %lu\n", (unsigned long)synced_epoch);
} else if (r != BootTimeSyncResult::Skipped) {
printf("[BOOT] cold-boot time sync: %s\n", bootTimeSyncResultName(r));
}
}
#endif
}
+12
View File
@@ -44,6 +44,18 @@ CONFIG_BT_CONTROLLER_DISABLED=y
# C6 SDIO/power path isn't proven yet; without this the host reboots ~10 s after a failed C6 init,
# preventing the UI (and LoRa) from ever coming up. Keep the device alive; Wi-Fi/BLE just stay down.
CONFIG_ESP_HOSTED_HOST_RESTART_NO_COMMUNICATION_WITH_SLAVE=n
# ...and the OTHER reset path, which is the one that actually bites. The line
# above only covers "no communication with slave"; a transport failure after the
# connect attempt times out goes through H_TRANSPORT_RESTART_ON_FAILURE in
# sdio_drv.c and calls _h_restart_host() regardless. On a unit whose C6 still
# runs factory ESP-AT the SDIO stream never frames, so that fires every ~10 s
# and the board boot-loops: measured on Kaj's own P4, 2026-10-05, five resets in
# a 45 s capture, each one rst:0x9 HP_SYS_LP_WDT_RESET right after
# "Host is resetting itself, to avoid any sdio race condition".
# A co-processor that cannot be reached must cost Wi-Fi and BLE, not the device.
# The failure still arrives as ESP_HOSTED_EVENT_TRANSPORT_FAILURE either way,
# which is what the app should surface instead.
CONFIG_ESP_HOSTED_TRANSPORT_RESTART_ON_FAILURE=n
# Panic coredump -> the dedicated 64K 'coredump' partition (see tdisplay_p4_16M.csv). Unlike the
# Tanmatsu (whose table has no coredump partition, so UITask stubs the API), the P4 has one, so
+23
View File
@@ -34,6 +34,7 @@
namespace {
constexpr uint32_t kHotplugMs = 1500; // how often the expansion is looked for
static bool s_scanned = false; // one-shot bus scan (see p4KeyboardPoll)
constexpr uint32_t kIdlePollMs = 120; // backstop drain when INT says nothing is waiting
constexpr uint32_t kModTimeoutMs = 1500; // a one-shot modifier lapses after this
constexpr uint32_t kRepeatDelayMs = 450;
@@ -312,6 +313,28 @@ void p4KeyboardPoll() {
s_next_probe_ms = now + kHotplugMs;
uint8_t cfg = 0;
const bool there = regRead(w, P4KB_EXP_ADDR, XL_CFG0, cfg);
// One-shot bus scan on the first probe. "It does not work" on this
// expansion has three very different causes and they are impossible to
// tell apart from the outside: the board is not clipped on, it is
// clipped on but unpowered (it carries its own cells, and the firmware
// has no enable for it), or it answers and the bring-up fails. An
// unpowered board does NOT give a clean NACK either: with its pull-ups
// dead the lines sit low through its ESD diodes and every transaction
// comes back ESP_ERR_INVALID_STATE, which reads like a driver fault.
// So print what is actually on the bus, once, and let the log say which
// of the three it is. Same one-shot scan the touch bring-up does.
if (!s_scanned) {
s_scanned = true;
char sc[96]; int n = 0;
n += snprintf(sc, sizeof sc, "[P4KB] i2c46/45 scan:");
for (uint8_t a = 0x08; a < 0x78 && n < (int)sizeof sc - 6; a++) {
w->beginTransmission(a);
if (w->endTransmission() == 0) n += snprintf(sc + n, sizeof sc - n, " %02X", a);
}
if (n < (int)sizeof sc - 24 && !there)
snprintf(sc + n, sizeof sc - n, " (expansion absent)");
printf("%s\n", sc);
}
if (there && !s_present) {
if (!attach(w)) printf("[P4KB] keyboard expansion found, bring-up failed\n");
} else if (!there && s_present) {