Blog

Second Pass on the Phish Analyzer: Hashing, Thread Hijacking, and a Bug I Introduced Myself

← Back to Blog

After the rebuild I wrote about last time, the analyzer had a sound scoring model and a set of checks that finally told you something true. This pass was less about fixing what was wrong and more about asking what a real detection engineer would do with a report like this once they had it in hand, and it turned up one bug I genuinely didn't expect to find.

Hashing without opening

The biggest documented gap after the rebuild was that nothing about attachments beyond filename and declared type ever got looked at. That was a deliberate boundary, not an oversight, because the moment you start interpreting a file's actual contents, you've taken on the job of a hostile-input parser, which is a different and much bigger commitment than everything else in this tool.

A cryptographic hash sits on the safe side of that line, and I hadn't noticed until I looked closely. Computing an MD5, SHA-1, or SHA-256 over a file's bytes never asks what those bytes mean. It doesn't matter if the file is a PDF, an executable, or garbage, the hash function treats it as an undifferentiated stream either way. There's no format to get wrong, no structure to misparse, nothing for a malformed file to exploit. That's categorically different from decoding an image or parsing a document, which is why the QR-code gap I flagged last time is still there and this isn't a step toward closing it.

Once I confirmed that distinction held up, the rest was straightforward. For a raw .eml file or pasted email text, the library already parsing the message has the attachment bytes sitting in memory as an ordinary part of doing its job. Nothing new gets read, nothing new gets opened, I just stopped throwing the bytes away before computing a digest over them.

The one place I drew a hard line was Outlook's .msg format. That parser was written from the start to never call the library function that extracts attachment content, specifically so .msg attachment bytes are never in memory at all. Extending hashing to that path would have meant reversing a real design decision, not just adding a feature, so I didn't. A .msg attachment still shows up with a filename and a type. It just doesn't get a hash, and the tool doesn't pretend otherwise.

Why the hash actually matters

It would have been easy to add hashing and stop there, treating it as a nice-to-have detail on the attachment list. What makes it worth the effort is what it does to the Sigma rule the tool generates.

Everything else that rule matches on is inherently soft. A sender domain might be a compromised host you don't want to block outright. Subject-line keywords are the most false-positive-prone thing in the whole document, and the rule says so in its own comments. A hash is different. If a file's SHA-256 matches, it is that file, not something that resembles it. So the rule now includes an attachment_hash selection whenever one is available, and I marked it explicitly as the one part of the generated rule that's trustworthy as-is rather than a starting point to review.

A pattern I hadn't named yet: thread hijacking

Once I started thinking about what a real analyst would want from this tool, one gap stood out that had nothing to do with attachments at all.

A reply carries specific technical headers, In-Reply-To and References, that a mail client generates from an actual prior message. A phishing kit sending a cold message from scratch has no reason to fake them, and mostly doesn't bother. So a message that presents itself as a reply, with those headers genuinely present, while failing real authentication, is telling you something more specific than ordinary phishing would. Either the account sending it is actually compromised and mailing from infrastructure it doesn't normally use, or the thread headers were lifted from a real message to make an unrelated one look like a continuation of it. Either way, that's the shape of business email compromise, not a generic phishing blast.

I was careful about what counts as a real authentication failure here. SPF softfail or an absent DMARC policy are extremely common on completely ordinary mail, so gating this on "not a perfect pass" would have made it fire constantly and taught people to ignore it, which is exactly the failure mode I wrote about avoiding last time. It only fires on an actual failure: SPF fail, DKIM fail, DMARC fail, or DMARC enforcement being applied. A reply-looking subject with none of the real headers behind it doesn't count either, since anyone can type "Re:" by hand.

The bug I found while writing the test for it

Testing the hashing feature meant constructing a .msg-style message and checking that its attachment metadata came through correctly. It didn't.

I traced it back to how a .msg file gets turned into something the rest of the pipeline can read. Since the analyzer's other checks all expect plain email text, .msg files get reconstructed into that shape first, either by reusing the real transport headers if the file still has them, or by synthesizing a minimal From: and Subject: block if it doesn't. Both of those reconstructions start with header lines that look exactly like the start of a real email, which is the point, that's what lets every other check run against .msg mail without needing to know .msg exists.

The problem is that the code path that decides "does this text look like a real email" runs before attachment handling, and virtually always says yes for reconstructed .msg text. When it does, the parser hands the whole thing to the library that reads real MIME messages, which correctly finds zero actual attachment parts, because the reconstructed text is headers and body only, nothing else. And the attachment metadata that .msg parsing had already extracted separately, filenames, types, sizes, just wasn't consulted at all in that branch. Every .msg attachment had been silently disappearing before it ever reached the analyzer, since the feature was first built.

This has nothing to do with hashing specifically. It would have caused the same problem for the double-extension check, the macro-document check, any of it, for anyone who uploaded a .msg file with an attachment. I only found it because I happened to write a test that exercised that exact path while checking something unrelated.

The fix was to make explicitly-supplied attachment metadata take priority whenever a caller provides it, regardless of which parsing branch the message text falls into. Only the .msg code path ever supplies that second argument, so the change is a no-op for ordinary .eml uploads and pasted text, and it stops discarding real data for .msg uploads specifically.

A narrower check I almost didn't add

I also added one more attachment check: a filename with an ordinary extension, a photo or a document, whose declared type in the message is a Windows executable format. That combination has no innocent explanation.

I nearly left it out. The most common failure mode for a check like this is declaring the wrong content type entirely, since plenty of mail clients tag an attachment as application/octet-stream whenever they don't recognize it, and treating that as suspicious would flag a large share of completely normal mail. The version I built deliberately excludes that generic type and only looks for a small, specific list of executable-flavored types against a small, specific list of benign extensions. It's a narrow check on purpose, precisely because the last round of work was so much about the cost of a check that isn't.

What it looks like put together

A test message with a spoofed sender, a failing DMARC check, reply headers pointing at a real prior message, and an attachment declaring itself as invoice.pdf.exe now produces a report that leads with what to do about it, groups the DMARC failure and the thread-hijack pattern separately so it's clear both fired independently, lists the attachment's hash with a one-click lookup, and hands over both a Sigma rule with that hash wired in and the whole result as JSON for anything that wants to consume it programmatically. None of that existed before this pass, and most of it exists because I kept asking what the report was actually for, rather than what else it could check.