Projects

Phish-Signals: Heuristic Phishing Detection Engine

View project →

← Back to Projects

phish-signals is the detection engine that powers /phish-report on this site, pulled out into its own project so the logic isn't stuck behind a web form. It ships as two independent, complete implementations of the same detection logic: @farksecurity/phish-signals on npm, written in TypeScript, and phish-signals on PyPI, written in Python. Feed either one a URL string, a headers Map/dict, raw .eml text, or an Outlook .msg file, and you get back the same structured findings the hosted tool shows, as plain objects you can do whatever you want with. No network calls, no API key, nothing phones home in either language. Everything runs on data you already have in memory.

What it checks

URLs and domains. Typosquats, including lookalike substitutions like paypa1.com. Brand names planted in a subdomain or hyphenated into an unrelated domain, which is the shape most real phishing links take and which a typosquat check alone misses entirely, since the spelling is exact. Raw IP-literal links, links to executables, shorteners, and credential-harvesting keyword density in the surrounding text.

Punycode, decoded properly. A hand-written RFC 3492 decoder, not Node's deprecated punycode module, turns xn--pypal-4ve.com into pаypal.com and names which Unicode scripts got mixed together to build it. That distinction matters: a single Cyrillic character substituted into an otherwise-Latin word is a different trick from a word written entirely in another script that was chosen because it looks Latin, and conflating them means missing one or the other.

Headers. SPF, DKIM, and DMARC parsed straight from Authentication-Results, with a defense built in for a specific forgery: an attacker appending their own fake, fully-passing Authentication-Results header after the real one. The engine counts Received headers to work out how many Authentication-Results headers could legitimately exist and discards anything past that count. It also walks the Received chain looking for HELO/reverse-DNS mismatches and timestamps running backwards, and flags thread hijacking, a message carrying real reply/forward threading headers while failing hard authentication, which is a much more specific tell than ordinary phishing ever bothers faking.

Attachments and content. Dangerous and double extensions, MIME-type-versus-extension mismatches, macro-enabled documents, urgency and credential-harvesting language, link text that doesn't match its href, and data:/javascript: URI schemes that carry no hostname for a domain check to even look at. ZIP archives get their central directory listed, names and declared sizes only, without ever decompressing an entry.

Scoring. Findings come back as Signal objects, a category, a severity, a label, never a hand-picked weight. Within a category the strongest finding counts in full and the rest count for a fraction of that, so SPF, DKIM, and DMARC all failing registers as one fact rather than three stacked ones. No single category can push a verdict into the top band alone, which means a high-risk score requires evidence spread across more than one axis. Benign evidence, a clean aligned pass, no links in the message at all, lowers the score, but gets suppressed the moment anything high-severity has already fired.

Analyst output. Defanged and refanged IOC extraction, MITRE ATT&CK technique mapping, a generated Sigma rule built from whatever was actually observed in the specific message (labeled a starting point, not a finished rule, since email telemetry field names differ across every SIEM), a generated KQL query for Microsoft 365 Defender / Sentinel Advanced Hunting built the same way, against whichever of EmailEvents, EmailUrlInfo, EmailAttachmentInfo, UrlClickEvents, DeviceNetworkEvents, and DeviceFileEvents are actually relevant to what was found, and a JSON export of the full result.

Parsing. The layer with real dependencies, mailparser for raw .eml and pasted messages, @kenjiuno/msgreader for Outlook .msg, and jsqr/pngjs/jpeg-js for QR decoding out of embedded or attached images on the TypeScript side, with equivalent libraries doing the same jobs on the Python side. QR decoding is bounded on size, pixel count, and wall-clock time so a hostile image can't turn a single scan into a denial-of-service vector. All of it normalizes down to one structured shape that every check above consumes, regardless of which format or which language it started in.

Relationship to /phish-report

This isn't a rewrite or a similar tool, it's the actual engine. The TypeScript side started as a direct copy of the same code the site runs, and the site still runs its own copy today. But they version independently, so if you depend on either published package directly, pin a version rather than assuming it tracks the site in lockstep. The site's /phish-report route may sit a version ahead or behind whatever's published at any given moment, and it doesn't import either package directly.

How it ships

The project lives entirely at github.com/MasonC-402/phish-signals now, with its own git history, its own issues, its own release cadence, no longer mirrored out of this site's repo. It's laid out as typescript/ and python/, two independent, complete implementations rather than a core engine with a thin language binding on top. Each is a standalone installable package with its own test suite, its own docs, and its own version number, and they aren't expected to version in lockstep.

What keeps two separately written implementations from quietly drifting apart on what a given input actually means is conformance/: a set of language-neutral JSON vectors, input in, expected findings out, checked into the repo and run by a small test harness in each language. A vector that only one language's harness satisfies is a failing test in CI, not something that turns up later because someone happened to compare outputs by hand.

Each language publishes on its own tag-triggered pipeline through OIDC trusted publishing, no stored npm or PyPI tokens anywhere. The tag prefixes are kept distinct, npm-v* for the TypeScript release workflow and pypi-v* for the Python one, specifically so a release in one language never triggers the other's pipeline by accident. As of this writing that's put @farksecurity/phish-signals at 0.2.1 on npm and phish-signals at 0.2.2 on PyPI, published a version apart from each other because that's exactly what independent release cadences look like in practice.

What it doesn't do

No DNS lookups, no reputation data, no sandboxing, no attachment execution. Every check here is structural, reading what's already in the message rather than asking anything about the outside world. It can't tell you a domain was registered yesterday or that an address has never written to you before, and it never will by design. The phrase-matching content checks are English-only. None of this replaces a person reading the output. It's a heuristic engine, meant to hand a human or a downstream system a scored starting point, not a verdict to act on blind.

Get the TypeScript package from npm, the Python package from PyPI, or read the source on GitHub.