A realtime `lf read` streams samples straight to the host at the LF sample rate,
around 125 kB/s. If the host stopped draining for even a moment,
async_usb_write_requestWrite() returned false, and ReadLF_realtime() treated that
as fatal: it returned PM3_EIO through a goto that also skipped
async_usb_write_stop(). The IN endpoint was left busy, every later usb_write()
then returned PM3_EIO, and the device went silent until it was physically
replugged. A USB bus reset does not clear it -- EP0 keeps working, descriptors
read fine, only bulk IN is dead. That is what a long `lf read` looked like from
the client: a transfer that stopped early, and then a device that would not
answer the next command.
A busy endpoint is back pressure, not an error. The sampling loop now waits for
the host, up to 100ms, and only gives up if it never comes back. Every exit path
closes the async write. The spins inside the USB helpers are bounded and drop the
stuck packet instead of turning into an infinite loop, which is the other half of
why the device never recovered. Measured against a reader deliberately held at
41 kB/s: before, the stream died at 212238 of 300000 bytes and the unit needed a
replug; after, 300000 of 300000 and it still answers.
The last packet of a stream was being lost as well. The host's CDC read buffer is
two max sized packets, and a bulk IN transfer only completes on a short packet or
a full buffer, so a stream ending on a full 64 byte packet was left sitting in a
half filled buffer. It showed as exact parity on the packet count:
513 packets requested -> 512 delivered 514 -> 514
515 packets requested -> 514 delivered 516 -> 516
async_usb_write_stop() now always sends a closing packet, the leftover partial
bytes when there are any and a zero length packet otherwise, the same way
usb_write() already did for its own transfers.
On the client side a short transfer is reported instead of being presented as a
complete read, the device is told to stop streaming on that path too, and the in
place byte counter is repeated with its final value, since the loop only samples
it every 10ms and the last line printed was stale.
Lastly `lf read` and `lf sniff` cap the sample count to the graph buffer size.
Anything past it was streamed and then discarded by getSamplesFromBufEx(), so
`-s 1500000` spent about two extra seconds collecting 220000 samples that were
thrown away, and printed "Received 1500000 / 1500000 bytes" followed by "Got
1280000 samples" with nothing to explain the gap.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
AT91F_USB_SendStall spun on STALLSENT with no exit. That bit only
arrives once the host polls the endpoint and takes the handshake, so a
host that vanishes in between wedges the device. It was the last
unbounded wait in the file without an escape.
Exit on RXSETUP (host abandoned the request and sent a new SETUP), on
ENDBUSRES, or on a spin count, reusing the 0x1FFF that usb_read already
uses. usb_check() is not usable here: it calls back into
AT91F_CDC_Enumerate, which is what calls this.
Do not test RXSUSP. Nothing writes it back to UDP_ICR, so it latches on
the first bus idle and stays set, which makes the guard fire on every
stall and stops the STALL from ever being delivered.
The second wait is bounded too, since leaving the first one early can
let the host set STALLSENT after the clear.
Also ISOERROR -> STALLSENT. Same bit 3, but only one of those names is
true on a control endpoint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
uint32_t plus three uint8_t pads to 8, so sizeof(line) was 8 where CDC
defines 7. Hosts asking for more than 7 got a byte of uninitialised
static padding. The commented out SET_LINE_CODING loop already hardcodes
i < 7, so the wire format was understood; the struct just did not match.
No observable change for cdc_acm or usbser.sys, which both ask for
exactly 7.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
AT91F_USB_SendData only took length, and every caller passed
MIN(sizeof(x), wLength), so the function could not tell it had
truncated. A control IN transfer ends on wLength bytes or on a short
packet; returning fewer than wLength in whole 8 byte packets is neither,
and the host polls until it times out.
Take wLength too, do the MIN inside, and send a ZLP when the reply is
shorter than asked and lands on a packet boundary. The length > 0 guard
avoids a second termination, since the send loop already emits an empty
packet for an empty payload.
This is what the serial number descriptor's '(size % 8) == 0 OS bug
workaround' padding was dodging. Not an OS bug.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
StrMS_OSDescriptor advertises MSFT100 and vendor code 0x1C, but the
handler for that request was commented out, so it fell through to the
standard switch and stalled. Windows caches that under
usbflags\<VID><PID><bcdDevice> and stops asking, and bcdDevice never
changes, so no reflash could recover it.
Enable both feature descriptors and answer 0x1C in AT91F_CDC_Enumerate.
The compatible ID is left empty so usbser.sys keeps binding; WINUSB
there would take interface 0 away from CDC and kill the COM port.
DeviceInterfaceGUID was GUID_DEVCLASS_PORTS, a setup class GUID in an
interface GUID field - replaced with a fixed project GUID.
AT32 is unchanged; both descriptors are behind #ifndef PM5.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PLATFORM=PM3ICOPYX has not compiled at all -- it dies on obj/start.o,
before anything else is built:
common_arm/gpio/gpio_hw_at91.h:77: error: 'GPIO_FPGA_ON' undeclared
(first use in this function); did you mean 'GPIO_FPGA_DONE'?
PA26 has two different jobs depending on the FPGA, per
config_gpio_proxmark3.h:
#if defined XC3
#define GPIO_FPGA_SWITCH AT91C_PIO_PA26
#else
#define GPIO_FPGA_ON AT91C_PIO_PA26
#endif
ICopyX has no FPGA power rail to switch; that pin selects the FPGA
instead. But Gpio_FPGA_ON_High(), Gpio_FPGA_ON_Low() and
gpio_fpga_on_setup() all referenced GPIO_FPGA_ON unconditionally, so the
symbol simply does not exist on XC3.
'This board has no FPGA power switch' is already a supported case: the
AT32 port stubs both accessors with a bare '// Unsupported' comment, and
PM5 depends on that, while Gpio_FPGA_SWITCH_High()/Low() two functions
further down in this same header are guarded with #ifdef for the mirror
case. The AT91 header just never grew the XC3 arm. All three now carry
the same guard.
fpga_loader.c calls Gpio_FPGA_ON_High() from shared code; on ICopyX that
compiles to nothing, which is correct -- the rail is always on.
Verified: fullimage builds for PM3ICOPYX (356832 bytes, with
PLATFORM_EXTRAS=FLASH), and for PM3RDV4, PM3GENERIC, PM3ULTIMATE and
PM5. For the boards that do have the pin the guard costs nothing -- the
.text section is byte-identical before and after, 289464 bytes on
PM3RDV4 and 243584 on PM3GENERIC; only the embedded build timestamp and
source hash move in the full ELF.
Co-Authored-By: Claude Opus 5 (1M context)
The 1 MB block branch exists for the ARM .data section: start.c's
uncompress_data_section() reads one 4-byte length and does one
LZ4_decompress_safe(), so .data has to arrive as a single block. It was
selected by 'num_infiles == 1', which is not what tells the two callers apart.
A build that skips LF, FeliCa and ISO15693 leaves FPGA_BITSTREAMS holding just
fpga_pm3_hf.bit, so the bitstream took that same branch and was packed as one
42 kB block. get_from_fpga_combined_stream() decompresses into a
FPGA_RING_BUFFER_BYTES buffer, 16 kB since 83c3f81b1:
[#] inflate returned: -13247
[#] reset_fpga_stream failed
Before 83c3f81b1 the copy was clamped with MIN(FPGA_RING_BUFFER_BYTES, ...)
whatever buffer_size said, so the blocks came out at 30 kB and the 30 kB ring
buffer still took them. That is why the commit looks like the cause - it only
removed the clamp that was covering for the wrong branch.
Add -s for the single block case and let the FPGA path always chop at
FPGA_RING_BUFFER_BYTES, however many bitstreams went in:
4 bitstreams 169344 in -> 106933 out, byte identical to before
1 bitstream 42172 in -> 28718 out, 3 blocks 13265/13206/2247,
was 1 block of 27627
.data (-s) 14944 in -> 8786 out, byte identical to the
obj/fullimage.data.bin.z in tree
Also hand the ring buffer back when reset_fpga_stream() fails. The early
return left it allocated for the rest of the session, which is the reporter's
[#] BigBuf_size............. 48116
[#] Available memory........ 31732
48116 - 31732 is 16384, exactly FPGA_RING_BUFFER_BYTES.
No CAPABILITIES_VERSION bump: fpga_all.bit.z is objcopy'd into the same
fullimage as the decompressor that reads it, so nothing here is client facing.
Reported and correctly diagnosed by @ewangsoft.
Fixes#3599
Co-Authored-By: Claude Opus 5 (1M context)
CMD_SPIFFS_DOWNLOAD read the whole file into BigBuf first. BigBuf_calloc()
takes a uint16_t while the size is a uint32_t, so a 64K file allocated zero
bytes and anything larger wrapped to a short buffer the file was then read
straight past the end of. rdv40_spiffs_read_stream() opens the file once and
hands out one frame at a time, so the size never reaches the allocator, and the
SPIFFS path now honours start_index. Verified with an 80K round trip: upload,
download, byte identical.
The same BigBuf_calloc(filesize) wrap was in copy_in_spiffs(); it now copies in
SPIFFS_WRITE_CHUNK_SIZE chunks.
Separately, every flash failure was invisible. SPIFFS_CHECK_RES() only treats a
negative result as an error and the HAL hooks returned 128/129/130, so SPIFFS
saw success. The erase path was worst: Flash_Erase4k() only reports that the
command was sent, both Flash_CheckBusy() results were discarded, and
'return (SPIFFS_OK == erased)' yields 1 on failure. A flash dump from a device
that hit this showed one block written without being erased - its object index
header held '0x14000 & 0x10000', and the filename ANDed away to empty. Erases
are now read back and retried, and the HAL returns real SPIFFS error codes.
Flash_Write() returned len unconditionally even after a rejected page, which
also defeated the 'res == payload->len' check in CMD_FLASHMEM_WRITE.
Feedback, so none of this is silent again: CMD_SPIFFS_WRITE answers with the
SPIFFS result, the client aborts on the first refusal and names the byte it
stopped at, dl_it() checks the terminator status, and both transfer directions
print inline progress.
Co-Authored-By: Claude Opus 5 (1M context)
Send a synchronous zero-length packet after CDC responses whose length is
an exact multiple of the endpoint packet size, preventing oversized host
reads from remaining pending.
Simulation now completes the full exchange with a genuine Paxton reader in
password mode, and crypto mode read/write passes Proxmark-to-Proxmark.
Firmware:
- SOF was one bit period short. The lead-in that compensated for the lost
head half bit was removed and nothing replaced it, so readers rejected
every answer with a second START_AUTH. Default is now 6.
- The edge-detect threshold was latched before being measured, so the value
chosen depended on whether the Proxmark was in a field when sim started.
It is now measured on field entry and re-armed when the reader leaves.
- The percentile walk latched on run-scoped variables, so one attempt made
outside a field poisoned every later one.
- Field loss was detected from TIMESTAMP, which is free-running MCU time and
never stalls. Detect it from receive silence instead.
- Frames of a length the protocol does not have no longer reach the state
machine; our own modulation tail was resetting the session and breaking
every write.
- A dropped edge merges two or three reader bit periods into one gap. Those
bits were discarded; they are now recovered by decomposition, which is what
made crypto mode work (AUTH decode 15% -> 100%).
- Threshold selection is limited to 20 and 32 and settles in under 25 ms.
Client:
- lf hitag info printed a hardcoded 0x06 and reported 'Password mode' for
every tag. It now reads page 3, takes -k (4 bytes password, 6 bytes
crypto), and says so when the config cannot be read.
- lf hitag restore: writes a dump back in dependency order - user pages,
then key material, then config last - validates the config byte, and
prints the credential the tag will require afterwards.
- lf hitag crack2 now reports why it failed instead of a bare 'fail'.
- trace list: bit count moved to its own column, relative mode shows a
Frame Delay Time row rather than renaming Start/End, --frame and -r
rejected together.
FeliCa reading was broken on every card tested: 'hf felica reader' returned PM3_ETIMEOUT while the tag was answering correctly. The cause was in the FPGA demodulator, not the ARM.
fpga/hi_flite.v
---------------
Adaptive hysteresis thresholds. The envelope tracker clamped curmin to <= 70 and curmax to >= 180, so curminthres/curmaxthres were pinned near 91/160 no
matter where the signal actually sat. Measured on a RDV4 with the field on, the peak detector idles near 112 and a tag swings it by about +/-35, ie entirely
inside that window - so nothing ever crossed a threshold and every frame demodulated as a constant. The band is now derived from the tracked envelope,
3/16 of its span, floored at +/- 8 to stay clear of the 4..6 counts of carrier ripple.
Matched-filter bit detector. The slicer counted comparator trips (+1 above curmaxthres, -1 below curminthres, repeat the last crossing direction inside
the dead band), so every bit depended on where the band happened to sit. A mispositioned band railed the output to a constant and, since only the stable
branch can recompute thresholds or desync, it stayed that way for the rest of the session. It also discarded amplitude, gaining nothing from 32x
oversampling. Each half-bit is now integrated in the ADC domain and the larger half wins. Thresholds still drive bit phase and the desync, they no longer
decide bit values, so a clipped or mispositioned envelope can no longer rail the output.
Polarity lock guard. try_sync arms part way through a half-bit, so the first decision after arming is meaningless and could latch 'zero' inverted, decoding
the whole frame with the wrong polarity and losing the sync word. Skip the first two decisions; the preamble is 48 bits.
curbit re-timing. The bit decision was made in the bit-phase domain, which is aligned to the tag's edges, but sampled by the SSC in the carrier domain. The
ARM could latch a bit mid-transition at a phase that varied per frame. Both run at 64 carrier periods per bit, so re-timing curbit half an SSP bit away from the
sampling edge is a re-time, not a resample.
Envelope watchdog. FPGA registers persist across PM3 commands - only a bitstream reload clears them - so the tracker could enter a state it never left and the
first command after the client started would work while every one after it failed. Force a re-centre when the demodulator has not reached a known-good idle
for 19.3 ms, held off at the start of each frame so it cannot fire mid-reply.
state is marked (* fsm_extract = 'no' *). The project synthesises with -fsm_style bram; once XST recognised this register as a state machine it placed
the state ROM in a block RAM, and the xc2s30's six were already spoken for. MAP then failed to fit with nothing but a generic 'design is too large' error, no
BITGEN, and no new bitstream.
armsrc/felica.c
---------------
- felica_select_card() returning 4 (response too short for IDm+PMm) fell through to PM3_SUCCESS, so 'hf felica reader' reported an all-zero IDm as a good read.
- After a poll timeout the code still read FelicaFrame; with a stale POLLING_RES and len == 0, check_crc() was handed (len - 2) as a size_t, ie 65534.
- WaitForFelicaReply() could only time out from STATE_UNSYNCD/TRYING_SYNC and would spin forever if a frame never completed.
- felica_sniff() decremented and broke before LogTrace, so '-s 10' logged nine frames and '-s 0' logged none. CRC-failed noise no longer pollutes the trace.
- felica_sendraw() sent no reply at all for some flag combinations, leaving the client blocked until its own timeout.
- Polling used time slot 0 only, so several cards in the field collided forever. Retries now widen the TSN window.
- BuildFliteRdblk() warned about a bad block count and built the frame anyway.
Signal probe
------------
'hf felica raw -p' streams the per-window envelope min and max instead of demodulated bits, so reading distance and coupling can be measured rather than
guessed. This is what told 'tag out of range' apart from 'demodulator not locking', which are otherwise identical from the ARM's point of view.
Measured on a RDV4, both cards previously unreadable:
FeliCa Standard RC-S830 (CJRC 0003): reader 4/4, info 4/4, 39 nodes discovered, dump complete in 2.0 s, 37/40 single polls.
FeliCa Standard RC-S962 (Octopus 8008): reader 10/10, 23 nodes discovered, dump complete in 1.5 s, 40/60 single polls. This one drives the envelope onto the bottom ADC rail; the matched filter reads it anyway.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`hw fpga config` drives the FPGA over JTAG, but on PM5 the FPGA has to be
clocked/powered first - previously that only happened as a side effect of a
reader command (e.g. `hf 14a read`), so `hw fpga config` failed if run on its
own. Bring the FPGA up in the HAL at the start of FpgaStartConfig()
(FpgaSetup24MHzClk) so the command works standalone; no doc workaround needed.
Adds an opt-in `--rgb` flag to the continuous `hf tune` / `lf tune` commands that
mirrors the antenna tuning level on the PM5 antenna RGB LED: blue = low, green =
mid, red = high, tracking the on-screen bar so you can find coupling (e.g. an
implant) by feel without watching the screen.
The colour is computed client-side from the same per-iteration voltage / running
peak the bar uses (so it matches the bar and auto-scales), and pushed to the
device via a new dumb CMD_PM5_RGB_SET {r,g,b}. That command is handled (#ifdef
PM5) by a dedicated AT32 RGB HAL module, common_arm/rgb/{rgb_apis.h,
rgb_hw_at32.c} (RgbLedSet(), I2C controller @ 0x48), wired into the armsrc
Makefile/CMake as SRC_RGB for PM5 only - so no other platform is affected and no
hardware code lands in shared files.
The AT32 battery-domain helper at32_bpr_write_dt1() enabled the ERTC (ertcen)
but never selected an ERTC clock source (ertcsel stayed NOCLK). Writing the
ERTC write-protection register (ERTC->wp) requires the ERTC to be clocked, so
with no source the write stalled the APB once the CPU ran at the full PLL clock
(288MHz, 144MHz APB); it only survived on the slow HICK clock.
The bootrom reaches this via check_goto_flash_mode() -> system_bpr_chk_clear()
while reading/clearing the "enter flash mode" magic the OS sets on a software
reset. That runs after ConfigSystemClocks() (the button branch needs
SpinDelayUs()/the timer clock, so the clock must be up first), i.e. at 288MHz,
so the bootrom hung and the device never re-enumerated into the bootloader
until a physical USB replug.
Fix: select the ERTC clock (HEXT/20, as the tick HAL does) before touching the
ERTC registers. HEXT is running by the time this executes, so the wp write no
longer stalls. The fix is contained in the AT32 HAL (sys_hw_at32.c, which only
builds for AT32/PM5); the shared bootrom flow and its ConfigSystemClocks()-first
ordering (needed by AT91/RDV4 SpinDelayUs) are unchanged.