While testing lookalike detection for a hardware-wallet brand, the scanner returned seven high-confidence threats. Three of them were ordinary subdomains — `app.`, `auth.`, `billing.` — of a company with a similar name that has nothing to do with the brand and never claimed to.
They scored 0.85. Actionable. One approval away from a legal notice sent to a real business.
The cause
Credential words — *login*, *auth*, *secure*, *verify*, *app* — were being matched anywhere in the hostname. But a subdomain costs nothing and belongs to whoever already owns the domain. Only the registrable label, the part an attacker had to buy, carries intent.
`ledger-login.top` is a purchase. `login.ledger.ai` is a server name at a company that already existed.
The fix, and what it cost
Scoring now separates the registrable label from the subdomain prefix and weights them differently. Credential wording in a bought name scores heavily; the same word as a subdomain of an unrelated domain scores almost nothing.
On the same live data the false positives went from seven to zero, while genuine combosquats — `paypal-login-support.com`, `ledger-live-login.top` — kept their scores. It cost some recall on an edge case: an attacker who compromises a legitimate domain and hosts a phishing page on a subdomain of it now scores lower. We took that trade deliberately.
Why this is the error that matters
A missed clone is a bad day. A notice sent to a legitimate business is a liability the customer carries, under their name, on our advice. The asymmetry is total, and it is why detection tuning in this field should bias toward silence.
It is also why we publish removal rate rather than detection count. A detection count rewards exactly the behaviour that produced this bug.