Dumping Flash
- Dumping Flash
Why this matters
Everything up to this point has been preparation. Identifying chips, finding pins, decoding buses — all of it exists to get you to the moment where you have the device’s firmware on your own disk.
Once you do, the assessment changes character entirely. You stop poking at a black box and start reading its source of truth: the code, the filesystem, the credentials, the keys, the URLs it contacts, the update mechanism it trusts. Everything in Software Reverse Engineering and most of what is in Attacks starts here.
Three routes, in order of preference
Try them in this order. People reach for the soldering iron far too early.
1. Download it
Vendor support sites, OTA update endpoints, and GPL source releases. Free, non-destructive, and legal for anything publicly posted.
A surprising fraction of assessments end here. Vendors publish firmware because customers need to recover bricked devices, and the same file that fixes a customer’s router is the file you wanted. GPL obligations frequently force publication of the kernel and toolchain even when the vendor would rather not.
If the device fetches updates itself, intercepting that traffic is the same thing without needing the vendor’s website — see HW #3 Part B.
2. Read it off the running system
If UART or JTAG gave you a shell or debug access, take the image from the device itself:
# cat /proc/mtd
dev: size erasesize name
mtd0: 00020000 00010000 "bootloader"
mtd1: 00040000 00010000 "kernel"
mtd2: 00080000 00010000 "rootfs"
mtd3: 00010000 00010000 "config"
# dd if=/dev/mtd2 of=/tmp/rootfs.bin bs=4096
# base64 /tmp/rootfs.bin # then reassemble at the other end
Non-destructive, and it gives you the deployed image including device-specific configuration — which a vendor download does not. If you have a shell, this is the best artifact you can get.
Getting bulk data off a device with only a serial console is the awkward part.
base64 over the console works and is slow. If the device has networking,
nc, tftp, or an HTTP POST are all faster.
3. Read the flash chip directly
The technique that always works, and where this course spends its bench time.
Internal versus external flash
Before reaching for a clip, establish where the firmware actually lives. The answer determines everything that follows.
| Internal flash | External SPI flash | |
|---|---|---|
| Where | Inside the MCU package | A separate 8-pin part beside the SoC |
| Typical size | 32 KB – 2 MB | 2 MB – 32 MB |
| How to read | Over SWD/JTAG | SPI clip or chip-off |
| Blocked by | Read protection fuses | Bus contention |
| Common on | Microcontroller-class devices | Linux-class devices |
A device with a large Linux filesystem almost certainly has external flash. A sensor or a small controller almost certainly does not. If you found an 8-pin part next to the SoC with a JEDEC ID, it is external. If the SoC is a microcontroller and there is no such part, it is internal.
Reading internal flash
$ st-flash read dump.bin 0x08000000 0x40000
or, more generally:
$ openocd -f interface/stlink.cfg -f target/stm32f3x.cfg \
-c init -c "reset halt" \
-c "flash read_bank 0 dump.bin 0 0x40000" \
-c reset -c shutdown
Both the length and the OpenOCD target config have to match the part you
actually have — 0x40000 is 256 KB, the size of an STM32F303VC. Get the length
from the device rather than from a tutorial: st-info --probe reports it.
st-flash happens to clamp an over-long read to the real flash size without
complaining, so an inherited 0x100000 gives you a correct dump and a wrong
idea of how big it is.
No clip, no desoldering, no contention. If this fails while the probe still identifies the part, you are looking at read protection rather than a wiring problem — see the worked example in JTAG.
Reading external flash in circuit
A SOIC-8 test clip onto the flash chip, wired to a programmer. The clip is the fiddly part: it must seat squarely on all eight legs, and on a densely populated board there may not be room.
flashrom is the standard tool and supports a wide range of programmers:
$ flashrom --programmer ch341a_spi
flashrom v1.8.0 on Linux 6.12.0 (x86_64)
Found Winbond flash chip "W25Q128.V" (16384 kB, SPI) on ch341a_spi.
$ flashrom --programmer ch341a_spi --read dump1.bin
Reading flash... done.
Common programmer options:
| Programmer | --programmer value |
Notes |
|---|---|---|
| CH341A | ch341a_spi |
Cheap and ubiquitous. Many boards ship at 5 V — see the warning below |
| Bus Pirate | buspirate_spi:dev=/dev/ttyUSB0 |
Slower, safer, more configurable |
| FT2232-based (incl. Tigard) | ft2232_spi:type=..., port=... |
Fast and reliable |
| Raspberry Pi SPI | linux_spi:dev=/dev/spidev0.0 |
Needs the SPI kernel driver |
⚠️ The classic CH341A modules drive 5 V on the SPI lines while powering the chip at 3.3 V. That is out of specification for a 3.3 V flash part and will eventually damage something. Modified modules and 3.3 V-level versions exist; check yours with a meter before trusting it on a board you care about.
Contention
If the SoC is powered and running, it is also driving the bus. Remedies, with their risks:
| Method | Risk |
|---|---|
| Hold the SoC in reset (assert NRST) | Some SoCs do not tri-state on reset; a watchdog may release it mid-read |
| Power the flash from the programmer, board unpowered | Back-powering through the SoC’s ESD diodes partially powers it, and it may drive the bus anyway |
| Desolder the chip | Thermal damage, pad lift, rework needed afterwards |
A concrete decision rule: if two consecutive dumps produce different SHA-256 values, in-circuit reading is not working. Try one remedy. If it still differs, desolder. That is a falsifiable criterion, unlike “when it gets difficult”.
Chip-off
Hot air at around 300 °C, flux, and patience. Lift the part, read it in a socket or on a breakout, and reball or re-solder afterwards if the device needs to keep working.
Slower, destructive-ish, and it always works. It also removes every question about contention and back-powering at once.
For BGA packages this stops being a bench technique and becomes a rework-station one — and at that point, reconsider routes 1 and 2.
Verification: the step people skip
Dump twice and compare. This is not optional and it costs nothing:
$ flashrom --programmer ch341a_spi --read dump1.bin
$ flashrom --programmer ch341a_spi --read dump2.bin
$ sha256sum dump1.bin dump2.bin
3e6ccc79596b6d6c464bb16d90d768562338bf502f2f2d419a04adcc4fb3b5ec dump1.bin
3e6ccc79596b6d6c464bb16d90d768562338bf502f2f2d419a04adcc4fb3b5ec dump2.bin
Matching hashes mean the read is repeatable. Differing hashes mean it is not, and everything you conclude from that image afterwards is unreliable. Common causes: a marginal clip connection, a running SoC contending for the bus, or the clock rate set too high.
Then sanity-check the content:
$ ls -l dump1.bin
-rw-r--r-- 1 student student 16777216 Sep 3 14:02 dump1.bin
$ xxd dump1.bin | head -2
00000000: 27051956 5b91d1d3 66d2a4c1 00600000 '..V[...f....`..
00000010: 001d6f42 00600000 8b5a1c74 05050200 ..oB.`...Z.t....
The size should match the capacity from the JEDEC ID. All 0xFF means an
erased chip or a failed read. All 0x00 usually means the programmer never
talked to the part at all.
Making sense of the image
Mapping it
$ binwalk dump.bin
DECIMAL HEXADECIMAL DESCRIPTION
--------------------------------------------------------------------------------
0 0x0 uImage firmware image, header size: 64 bytes,
data size: 1929538 bytes, compression: lzma,
CPU: ARM, OS: Linux, image type: OS Kernel Image
64 0x40 LZMA compressed data, properties: 0x5D
1929602 0x1D7A82 SquashFS file system, little endian, version: 4.0,
compression: xz, image size: 8534016 bytes
--------------------------------------------------------------------------------
Analyzed 1 file for 85 file signatures (187 magic patterns) in 22.0 milliseconds
🛠️ binwalk 3 is a rewrite, and its output moved. Version 3 (Rust) replaced the Python version 2 that most tutorials were written against. The columns are wider, field names now carry colons (
version: 4.0, notversion 4.0), and it prints the signature-count footer above. If a guide’s output does not look like yours, checkbinwalk --versionbefore you doubt your dump.
Account for every byte, including what binwalk does not identify.
Unidentified regions are findings, not gaps. Two byte values deserve
distinguishing:
0xFF— flash that was never written. Erased.0x00— something wrote zeros deliberately. Often filesystem padding to a block boundary, occasionally something more interesting.
Extraction
$ binwalk -e dump.bin
$ ls extractions/dump.bin.extracted/1D7A82/squashfs-root/
bin dev etc lib mnt proc sbin sys tmp usr var www
Note the shape of that path. binwalk 3 writes to
extractions/<file>.extracted/<hex offset>/, one subdirectory per signature it
extracted, named for where in the image it was found. binwalk 2 used
_dump.bin.extracted/ with a flat layout, which is what older write-ups show.
⚠️
binwalk -eruns external extractors on untrusted data. That is a code-execution surface pointed at a file you pulled off a device you do not trust. Extract inside a container or a throwaway VM, never on your daily driver.
If binwalk -e cannot handle a filesystem, carve and extract manually:
$ dd if=dump.bin bs=1 skip=$((0x1D7A82)) count=8534016 of=rootfs.sqfs
$ unsquashfs -d root rootfs.sqfs
Take the offset and the length from the binwalk map above — 0x1D7A82 and
8534016 — not from a tutorial. bs=1 is what makes an arbitrary byte offset
expressible, and it is also why this is slow; once it works, skip/count in
larger blocks will do the same job in a fraction of the time when the offset
happens to be block-aligned.
Entropy, and what it can and cannot tell you
$ binwalk -E dump.bin
| Entropy | Meaning |
|---|---|
| ~0.0 | Erased flash or padding |
| 0.3–0.7 | Code and data |
| ~0.8 | Uncompressed filesystem: executables plus strings and metadata |
| 0.99+ | Compressed or encrypted — the number alone cannot distinguish them |
That last row is where the common advice is wrong. A good compressor’s output is statistically indistinguishable from random data; that is close to the definition of a good compressor. LZMA-compressed firmware and AES-encrypted firmware both measure ~0.9998.
So entropy tells you where the interesting regions are. Three other things tell you what they are:
- Is there a header? Compressed data in firmware is nearly always wrapped in a container — uImage, gzip, LZMA magic. Encrypted regions typically begin with ciphertext immediately, because a header would leak structure.
- Does a decompressor accept it? The definitive test. If
unlzmaorgunzipconsumes it and produces sensible output, it was compressed. - What do the edges look like? Compressed regions usually show a dip at the header and ragged edges where padding starts; encrypted regions start and stop abruptly.
A maximum-entropy region with no header that nothing will decompress is encrypted. That does not end the assessment — it relocates it. The key must be somewhere the device can reach at boot: the bootloader, an OTP region, a secure element, or derived from a device identifier.
What to look for once it is open
| Target | Where |
|---|---|
| Accounts and hashes | /etc/passwd, /etc/shadow |
| Hardcoded credentials and keys | /etc/, config files, CGI scripts |
| Endpoints the device contacts | Config files, and strings in binaries |
| Services started at boot | /etc/init.d/, /etc/inittab |
| Third-party component versions | /usr/share/, /etc/os-release, binary strings |
| The web interface | /www/, /var/www/, cgi-bin |
For any version string you find, check a vulnerability database. Embedded devices ship years-old components as a matter of routine, and the gap between the shipped version and the current one is frequently the entire finding.
Key takeaways
- Try downloading, then reading off a running system, then reading the chip. People reach for hardware far too early.
- Establish whether the firmware is in internal or external flash before choosing a technique; a large Linux filesystem means external.
- Dump twice and compare hashes. A read you cannot reproduce is a read you cannot trust.
0xFFis erased flash;0x00is something that wrote zeros. They mean different things.- Entropy locates interesting regions but cannot distinguish compressed from encrypted — headers and decompressibility can.
binwalk -eexecutes external tools on hostile data. Contain it.- If two dumps differ, stop being clever and desolder.
References
- flashrom — https://www.flashrom.org/
- flashrom manual page — https://flashrom.org/classic_cli_manpage.html
- binwalk (v3, Rust rewrite) — https://github.com/ReFirmLabs/binwalk
- OWASP Firmware Security Testing Methodology — https://github.com/scriptingxss/owasp-fstm
- OWASP IoTGoat, deliberately vulnerable firmware — https://github.com/OWASP/IoTGoat
- SquashFS documentation — https://docs.kernel.org/filesystems/squashfs.html
- Linux MTD subsystem — http://www.linux-mtd.infradead.org/
- stlink tools — https://github.com/stlink-org/stlink
- NIST National Vulnerability Database — https://nvd.nist.gov/
Related course pages: Schedule · SPI · JTAG and SWD · Software Reverse Engineering · Attacks · Tools of the Trade
🛠️ Maintenance note:
binwalkv3 is a Rust rewrite whose CLI differs from the v2.x Python tool that most online tutorials still describe — re-verify everybinwalkinvocation here against the installed version each term.flashromprogrammer names occasionally change between releases, and the CH341A 5 V warning applies to the classic black modules that remain the cheapest option and therefore the most common in student hands.