Writing · Connected products

Decoding a format you did not design

The instrument's own software gave me raw data, slowly and awkwardly. The useful system was locked inside a proprietary binary format. Decoding it was the easy part. Trusting the decoder was the real work.

This summer I built a noise-monitoring system as unpaid community work. A Class 1 sound level meter records to an SD card: for each measurement, a summary file and a file of one-second readings, in a proprietary binary format. The system's output goes to a local authority, so the system has one rule above all others: a plausible but wrong number is worse than a crash. A crash gets noticed. A confident wrong figure in a report gets relied on.

I have written before about five things I was sure of on that project and had to retract. This piece is about the step before those stories: why I decoded the meter's files at all, and the checks that make a decoder of someone else's format safe to rely on.

Why decode the files at all

The manufacturer's export software did one job: it produced raw data. Getting it out was slow, the result was hard to analyse, and everything useful still had to be built on top: statistics, reports, a record of which measurement belongs to which assessment, and two small computers at different sites that copy their data to each other so that either can fail without losing anything.

The meter's own files held everything. They were simply in a format nobody had documented for me. Reading them directly turned a limited tool into a useful system.

Many hardware teams are in the same position, with a logger, a sensor or a legacy controller whose vendor tool does less than they need. A parser that runs is easy. A parser you can put your name to takes discipline, because a format you did not design will lie to you in ways that look perfectly reasonable.

The trap: wrong output that looks right

The first decoder ran. Values came out, plotted sensibly and passed the tests. Then a batch of fifty-four files decoded to a logging interval of 1.25 seconds instead of one. It was consistent across every file, which felt like confirmation. It was the bug: those files used eight-byte records and the parser was reading them as ten-byte ones, and ten divided by eight is 1.25. The full story is in the earlier essay. It was caught before any of those files reached a real result, which is the only reason it is a good story rather than a bad one. The lesson for anyone decoding a format is the one that shaped everything after it: when a decoder is wrong, it is usually wrong consistently, and consistency is exactly what makes wrong output believable.

So the question changed from "does it decode?" to "how would I know if it decoded wrongly?" These are the checks that answer it.

1. Confirm the structure more than one way

Never let one clue decide how the bytes are laid out. The record size is now confirmed three independent ways, and each can be checked on its own:

  • The length divides. All 54 files in the affected group divide exactly by eight; only eight of them also divide by ten.
  • The count matches the clock. At eight bytes, one test file holds 895 records for a measurement 894 seconds long: one per second, plus a final partial one. At ten bytes it would be 0.8 of that.
  • The values agree with the meter's own summary. The meter stores headline figures (the average level, the maximum, the peak) in a separate summary file. Decoded at the right size, the one-second readings reproduce them exactly.

Two of the channels had averages that agreed to two decimal places, so nothing above could tell them apart; a swap would have passed every check. A fourth test, against the meter's percentile figures, separates them in every file. Look for the case your checks cannot distinguish, because that is where a silent error will live.

The function that decides record size compares two independent signals and refuses to guess when they disagree. That sentence was in my design notes before it was true of the code, which returned as soon as the first signal gave an answer. Read the code, not the notes.

2. One decoder, not several

Over time the project grew three copies of some of the decoding, one of which had drifted. I found out the hard way. A fix to the decoder changed nothing, because the comparison against the manufacturer's figures exercised the export path, while a different copy of the same decoding fed the database. The divergent copy had also been discarding legitimate high readings whole on import that the other path kept.

A duplicated decoder fails silently, because each copy passes its own checks. All raw decoding now lives in one module, and everything that reads a file goes through it.

3. Check against something that did not write the code

The manufacturer's export, limited as it was, became the reference. For the same recording, my system's figures had to match the vendor's: the one-second profile agrees to about 0.03 dB, which is the resolution of the data and the rounding in the vendor's spreadsheet. Tests written by the same person, or the same agent, as the code tend to share its assumptions. A reference that knows nothing about your implementation does not. It began as one recording and grew to seven, from three different years. That comparison exposed seven defects in my own export that the test suite had passed.

What the reference found

Three of its findings are worth telling, because each was a plausible number that was wrong.

The safety check that destroyed data. The decoder clamped every reading below 20 dB up to 20.0, as a guard against nonsense values. The vendor's export showed 19.6, 19.8 and 19.9 where mine showed 20.0. Across 100 runs, 78,888 one-second readings were genuinely below 20 dB, the lowest 17.43. The guard protected against nothing and quietly rewrote real measurements. A defensive limit is a claim about the data. Check it against the data.

The marker that passes every check. The meter writes a value meaning "not recorded" that decodes to exactly −20 dB, twenty decibels below the threshold of hearing. Only the export code knew what it meant. Everywhere else it was just a number, and it passes every "is this value missing?" check on its way into a report. Used as the background level in a standard noise assessment, it would have produced a difference of +81 dB. It appears in 27% of the archived runs, and it is not corruption: the meter is correctly declining to calculate a statistic for runs too short to support it. The fix was to turn it into a real missing value in one place, at the decode boundary, where its meaning is unambiguous.

The error blamed on the hardware. 222 of 4,500 one-second values were exactly 0.1 dB low. My notes put that down to the meter's rounding convention. It was my system rounding twice: storing at one decimal place, then rounding again on export. Three separate checks confirmed it, including the error rate: 5.0% predicted for double rounding, 4.93% observed. A plausible explanation that blames the other party deserves the same suspicion as any other claim.

4. Test on real data, lots of it

Synthetic files test what you thought of. The suite runs against 527 real recordings from the meter's own cards, each a summary file and a profile file, and the 1.25-second oddity came from them: 54 of the 527 used the shorter record. Six files in the archive do not parse at all, and the suite has to handle those too. A hand-made fixture would have contained none of this.

5. Refuse to guess

When the decoder cannot be sure, it says so. If the two signals for record size disagree, the measurement is skipped, not decoded on a best guess. As the comment in the code puts it, a skipped run is recoverable where a silently misdecoded one is not.

The same rule applies to display. An older piece of code once worked out a missing end time by arithmetic when the meter's own value looked implausible. It produced confident, wrong times. It was deleted. The display now shows the meter's value or nothing. No fallbacks, no estimates dressed as data. A blank invites a question. A guess does not.

6. Test every boundary the data crosses

The decoder was right and the data was still at risk. The two sites copy their measurements to each other, and one side sent a field called end while the other read end_time. Every synced measurement would have lost its end time, and every test passed, because none of them exercised the hand-over between the two. A format you decode correctly can still be mangled by the format you invent to move it.

7. Write the claims down, ranked by damage

The project keeps a validation brief. It lists every significant claim the system makes, ranks them by how much harm a wrong one would do, says how confident I am in each and where it could still be wrong, and gives a reviewer the commands to try to break it. That last part matters most. A claim with no way to break it is an opinion.

What I have not published

The format itself. This piece describes the method, not the layout of anyone's files, and I have not named the manufacturer or the meter. The method is the useful part, and it applies to any data you did not design.

The question to ask of your own parser

The suite now runs 673 checks, and the number is not the point. The point is that each check above fails if the decoder is wrong in a believable way. If your system reads a format you did not design, ask one question of it: if this parser were consistently wrong, what would tell me?

© 2026 Catherine Ives-Yim. All rights reserved.

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.