PLATFORM=PM3ULTIMATE has not built since c0ecb1c62 (2026-04-14), which
fixed a real heap overflow by changing one character:
- if (total_size > num_infiles * FPGA_CONFIG_SIZE) {
+ if (total_size >= num_infiles * FPGA_CONFIG_SIZE) {
The old '>' checked after the fact, so a round could start with the
buffer already full and write num_infiles * FPGA_INTERLEAVE_SIZE bytes
past the end. But '>=' answers the wrong question: a buffer that is
exactly full is not an error, it is the expected end state for a
platform whose bitstreams are sized to FPGA_CONFIG_SIZE.
Two things had to change.
all_feof() tested the EOF flag, and C only sets that once a read has
already run off the end -- consuming the last byte of a file leaves it
clear. A file whose length is an exact multiple of FPGA_INTERLEAVE_SIZE
therefore still looked unfinished after its final whole chunk, and the
interleave loop ran one more round of pure zero padding. It now peeks a
byte with fgetc/ungetc instead, so a file read to its last byte counts
as finished right away.
PM3ULTIMATE is the only platform this reaches:
fpga_pm3_ult_felica.bit 69984 = 243 * 288 exactly
fpga_pm3_ult_hf_15.bit 69983
fpga_pm3_ult_hf.bit 69980
fpga_pm3_ult_lf.bit 69980
243 rounds * 4 files * 288 = 279936, which is exactly
4 * FPGA_CONFIG_SIZE for 2s50vq144. Round 244 was padding only, and '>='
killed it there. Stock PM3's largest bitstream is 42172, not a chunk
multiple, so its last round trips feof naturally and it never gets that
far -- which is why this went unnoticed, nothing in CI builds ULTIMATE.
The guard now asks whether the next round fits, which is the condition
it was always meant to express and is strictly stronger than the
original '>':
if (total_size + (num_infiles * FPGA_INTERLEAVE_SIZE) > num_infiles * FPGA_CONFIG_SIZE)
Verified: PM3ULTIMATE compresses 279936 bytes to 40195 and all four
bitstreams decompress back byte-exact. Output is byte-identical to
before for stock 2s30vq100 (169344 -> 106628), for icopyx XC3
(72864 -> 27292) and for the '-s' .data section path. fullimage builds
for PM3ULTIMATE, PM3RDV4, PM3GENERIC and PM5.
Note FPGA_CONFIG_SIZE is now exactly the size of the largest ULTIMATE
bitstream. Regenerate that one byte larger and the guard fires again,
correctly; bump the constant by one interleave step rather than touching
the guard.
Co-Authored-By: Claude Opus 5 (1M context)
zlib_decompress() walks its output in whole FPGA_INTERLEAVE_SIZE chunks:
for (long k = 0; k < *outsize / (FPGA_INTERLEAVE_SIZE * num_outfiles); k++)
so a stream that is not a whole number of chunks loses its trailing partial one.
With two or more inputs the read loop zero-pads each stream past EOF and the
total lands on a boundary, but the padding was guarded by 'num_infiles > 1', so
the single input case was left ragged. 42172 bytes of fpga_pm3_hf.bit is 146.43
chunks, and -d handed back 39788 - a clean looking prefix, short by 2384 bytes,
22 of them real bitstream data.
Gate the padding on single_block instead. It must not be 'always pad': -s is the
.data section, and start.c's uncompress_data_section() sizes the decompression
with __data_end__ - __data_start__. Rounding .data from 14944 up to 14976 makes
LZ4_decompress_safe() return an error, and that path is the LED panic loop, so
the firmware would never reach AppMain().
1 bitstream 42172 in -> 42336 packed, -d round trip byte identical,
archive 28730 -> 28731
4 bitstreams archive byte identical to 43fe6c3eb, round trip exact
-s .data byte identical to 43fe6c3eb on the same input, unpadded
The 164 padding bytes never reach the FPGA: DownloadFPGA() shifts out only
bitstream_length bytes, taken from the .bit 'e' section header.
Also simulated the ARM decoder over the new archive - LZ4_decompress_safe_continue()
block by block into a FPGA_RING_BUFFER_BYTES buffer - 3 blocks of 16384/16384/9568,
none over the ring buffer.
Co-Authored-By: Claude Opus 5 (1M context)
The 1 MB block branch exists for the ARM .data section: start.c's
uncompress_data_section() reads one 4-byte length and does one
LZ4_decompress_safe(), so .data has to arrive as a single block. It was
selected by 'num_infiles == 1', which is not what tells the two callers apart.
A build that skips LF, FeliCa and ISO15693 leaves FPGA_BITSTREAMS holding just
fpga_pm3_hf.bit, so the bitstream took that same branch and was packed as one
42 kB block. get_from_fpga_combined_stream() decompresses into a
FPGA_RING_BUFFER_BYTES buffer, 16 kB since 83c3f81b1:
[#] inflate returned: -13247
[#] reset_fpga_stream failed
Before 83c3f81b1 the copy was clamped with MIN(FPGA_RING_BUFFER_BYTES, ...)
whatever buffer_size said, so the blocks came out at 30 kB and the 30 kB ring
buffer still took them. That is why the commit looks like the cause - it only
removed the clamp that was covering for the wrong branch.
Add -s for the single block case and let the FPGA path always chop at
FPGA_RING_BUFFER_BYTES, however many bitstreams went in:
4 bitstreams 169344 in -> 106933 out, byte identical to before
1 bitstream 42172 in -> 28718 out, 3 blocks 13265/13206/2247,
was 1 block of 27627
.data (-s) 14944 in -> 8786 out, byte identical to the
obj/fullimage.data.bin.z in tree
Also hand the ring buffer back when reset_fpga_stream() fails. The early
return left it allocated for the rest of the session, which is the reporter's
[#] BigBuf_size............. 48116
[#] Available memory........ 31732
48116 - 31732 is 16384, exactly FPGA_RING_BUFFER_BYTES.
No CAPABILITIES_VERSION bump: fpga_all.bit.z is objcopy'd into the same
fullimage as the decompressor that reads it, so nothing here is client facing.
Reported and correctly diagnosed by @ewangsoft.
Fixes#3599
Co-Authored-By: Claude Opus 5 (1M context)
Now, you can enable at least two of your favorite technologies (such as LF and HF 14443A) attached a standalone mode and still have spare ROM space for other functionalities on a Proxmark3 Easy with a 256KiB ROM.
fpga_compress.c:176:32: warning: cast from 'char *' to 'int *' increases required alignment from 1 to 4 [-Wcast-align]
const int cmp_bytes = *(int*)(compressed_fpga_stream.next_in);
^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~