I set out to review this site for anything worth cleaning up. What I got was three confirmed bugs in the phishing analyzer that had been live for weeks, a scoring model that was wrong in a way I had not noticed, and eventually a fairly complete rebuild of how the tool thinks.
Writing it down partly because the bugs are instructive, and partly because the reason they survived is more interesting than the bugs themselves.
Three bugs, all live
A real UPS tracking link was scored as malicious. The typosquat check compares a hostname against a list of known brands using edit distance. The problem was where the "this is actually the brand" check sat: inside the comparison loop rather than before it. usps.com comes earlier in my brand list than ups.com does, and the two are one character apart. So a genuine UPS link matched usps.com at distance 1 and got flagged as a typosquat before the loop ever reached ups.com and its own whitelist entry. Any real package tracking email came back malicious.
Any unsubscribe link was scored as malicious. There is a classic trick where http://paypal.com@evil.tk sends you to evil.tk, because everything before the @ is treated as a username. My check for it stripped the scheme off the URL and then looked for an @ anywhere in what was left. Query strings routinely carry the recipient's email address, which contains an @. So every unsubscribe?email=you@example.com link in ordinary bulk mail tripped the check and scored 35 points.
Links that existed only in the HTML part were never looked at. The line responsible was extractUrls(textBody || htmlBody). Read that carefully: if there is a text part, the HTML part is never reached. Almost all real mail is multipart with both. And a benign-looking plain text part sitting next to an HTML part carrying the actual payload is not an edge case, it is the standard shape of a phishing email. The tool scored those zero.
There was a fourth, outside the analyzer: visiting /blog/constructor returned a 500 rather than a 404. The lookup table for post slugs was a plain object, so asking it for constructor walked up the prototype chain and returned a function, which is truthy, which sailed straight past the "not found" check and then threw in the template. Same for __proto__, toString, and friends.
Why they lasted
None of these were subtle. The reason they survived is that nothing in lib/ had a single test.
There was one test touching the analyzer. It posted an email and asserted that the string verdict-badge appeared in the response. That assertion passes if every heuristic in the codebase returns an empty array. It was testing that the page rendered, and I had been reading it as testing that the analysis worked.
That is the part worth taking away. The test was not missing. It was there, it was green, and it was measuring the wrong thing, which is worse than having no test at all because it buys you confidence you have not earned.
The scoring model was wrong
Fixing the bugs was the easy part. The harder problem was that the scoring did not hold up once I actually looked at it.
Every finding carried a hand-picked number and the numbers were added together. SPF failure, 25. DKIM failure, 25. DMARC failure, 25. A message that failed all three scored 75 and landed in high risk.
Except those are not three independent findings. DMARC failing is SPF and DKIM failing to align. It is one fact about the message, counted three times. Any unauthenticated email hit high risk on that alone, before it had done a single phishy thing. Meanwhile a message doing four different suspicious things across four unrelated areas could score lower than that.
Two other things were missing. Nothing could ever lower a score, so there was no way to say "this looks fine, and here is why." And a bare number out of 100 implies a precision that hand-picked weights do not have, with nothing anywhere saying how much of the message the tool had actually been able to see.
So the model got rebuilt. Findings now carry a category and a severity instead of a number. Inside a category the strongest finding counts fully and the rest count for a quarter, which is roughly how corroborating evidence works: the second and third signs of the same thing are worth something, but not the same as the first. That unauthenticated message now scores 35 instead of 75.
Categories are the independent axes, and none of them can reach the top band alone. A high risk verdict now requires evidence in more than one place.
Evidence can also argue in the other direction. An aligned authentication pass, or a message with no links in it at all, pulls the score down. That stops the moment anything serious fires, though, because plenty of real phishing is sent from perfectly configured infrastructure that the attacker owns outright. Valid SPF is not a character reference.
Confidence is reported separately now, along with a plain list of what could not be checked.
Then it needed to actually be useful
With the scoring sound, the remaining problem was that the tool stopped at a number. "High Risk, 87 out of 100" is not an answer to the question anyone is actually asking, which is either "am I about to get owned" or "what do I put in the ticket."
The biggest analytical gap was the delivery path. The Received headers were being parsed and then thrown away. That chain is the one part of a message an attacker cannot fully forge. They can prepend anything they like, but every hop after the point they inject at is written by servers they do not control, and those servers record both the name the sending machine announced and the name their own reverse lookup returned. When a connection announces itself as paypal.com and resolves to a cheap VPS on a throwaway domain, that is a spoofing indicator sitting in plain text, obtainable without making a single network request. That matters here because the tool deliberately makes no external calls, so a check that needs no lookup is worth a lot.
The chain also gives you forged timestamps for free. Time should move forward as a message travels. A server can rewrite the headers below it but cannot reach forward and change the clocks of servers further along, so time running backwards means a hop was fabricated.
I also stopped reporting "uses punycode" and started decoding it. Telling someone a domain uses punycode is technically true and completely useless. Showing them that xn--pypal-4ve.com reads as pаypal.com, with a Cyrillic character sitting in the middle of a word they thought was Latin, is the actual finding. That turned out to need two separate checks rather than one, which I will come back to.
Then the output. It now leads with what to do, ordered by urgency, and the actions at the top are the ones aimed at someone who has already clicked: change that password, turn on MFA, disconnect the machine. That window is minutes long, and someone in it is exactly the kind of person who ends up pasting an email into a tool like this.
Below that, collapsed, is a section for whoever is writing the thing up rather than deciding what is safe to open. Delivery path hop by hop. Indicators defanged so pasting them into a chat does not produce a live clickable link. ATT&CK technique mappings. And a Sigma rule generated from what was actually observed, which is the piece that turns one analyzed email into a detection for the next copy of the campaign.
That rule is labelled a starting point in its own comments. Email telemetry field names differ across every platform, and shipping something that looks ready to deploy but quietly matches nothing would be worse than shipping no rule at all.
Two bugs I wrote on the way
Both of these were caught by tests I wrote against the new code, which felt like some kind of vindication given how this whole thing started.
The punycode decoder. I wrote the decoder by hand rather than pulling in a dependency, since Node's own module has been deprecated for years and the algorithm is small and fully specified. The reference implementation is in C and checks whether a character is a valid digit with a single subtraction, relying on unsigned arithmetic to wrap negative values around into a large number that fails the range test. JavaScript arithmetic is signed. Any character below 0 produced a negative number, which passed the check, and decoded as a garbage digit instead of being rejected. Malformed input came back as plausible-looking nonsense rather than an error.
A YAML injection in my own output. The generated Sigma rule puts the email's subject line in the title field. I wrote it as a bare YAML value, and a subject containing a colon followed by a space parses as a nested key. An attacker-controlled subject could break the rule or smuggle a key into it. Quoting the value fixes it. Worth remembering that generated output is still output, and the input to it is still hostile.
The one that changed my mind
The homograph check is the piece I got most wrong before getting right.
A quick note on reading this section: the whole attack is that these domains look identical to the real thing, so on screen they will look identical here too. That is not a rendering mistake, it is the point. I have led with the punycode form, which is what the domain actually is underneath, because that is the only way to tell them apart in writing.
I started with a mixed-script test: flag any hostname whose label draws on more than one alphabet. That catches xn--pypal-4ve.com, which renders as PayPal's domain with a Cyrillic а sitting where the first a should be. Latin word, one foreign character swapped in, two alphabets in one label.
It completely misses xn--80ak6aa92e.com, which is the most famous example of this attack there is. That one renders as Apple's domain and is not mixed at all: every single character in it is Cyrillic. One alphabet, so nothing for a mixed-script test to catch, and my check called it clean.
So I added a second test for a word written entirely in another script but built only from characters that have Latin lookalikes. That caught it, and then immediately flagged xn--80adxhks.xn--p1ai, which is москва.рф, the Moscow city domain. Ordinary Russian, and also made entirely of Cyrillic characters that happen to resemble Latin ones. Structurally the two are identical. There is no property of the label itself that separates a real Russian word from an imitation of an English one.
What separates them is where they live. A Cyrillic word under a Cyrillic TLD, .рф, is a domain doing exactly what internationalized domains exist for. A Cyrillic word under .com is imitating something. Qualifying the check on the TLD fixes it without giving up the detection.
I only found that because I wrote a test asserting that legitimate international domains stay clean. Every version of that check passed the tests for the attacks. It was the test for the non-attacks that showed me the check was wrong.
A pin that quietly came undone
One last thing, found while updating the README rather than the code.
I wrote a post recently about migrating this site to TypeScript, and one of the things I called out was pinning @types/node to the version line that actually runs in production. If the type definitions describe a newer Node than the server has, the compiler will happily accept code using APIs that are not there, and it fails at runtime instead.
A grouped dependency update bumped it two majors ahead of the deployed runtime. No alarm, because nothing was broken yet: it only breaks the first time someone writes code against an API that exists in the types and not on the server. I put it back and told Dependabot to leave that one alone.
Pinning a version is a decision. Automation does not know it was a decision unless you tell it.
Where it ended up
Tests went from 18 to 124. Most of the new ones assert things that were already broken when I wrote them, which is the useful kind.
The thing I keep coming back to is the false positives. A tool that cries wolf on real UPS emails and ordinary unsubscribe links does not just annoy people, it teaches them to ignore it, and then it is worse than nothing because it has spent the credibility it needed for the one time it is right. Half the work here was making it quieter, not louder.