Written by

Mohan Reddy

At

Fri Jul 24 2026

The Payload Doesn't Lie: How Threatmatic π Stopped an Exfiltration That Every Other Tool Missed

Shannon entropy, DGA scoring, and behavioral analysis — three signals that individually mean nothing, but together caught a sophisticated attacker exfiltrating source code through encrypted HTTPS that no perimeter tool could see inside.

Back
Threatmatic π — multi-signal HTTPS inference

Every security tool on the network said: all clear.

The firewall saw HTTPS on port 443. ✓
The DNS resolver saw a domain that resolved. ✓
The EDR saw a legitimate node process. ✓
The SIEM had no alerts. ✓

Meanwhile, a developer's laptop was sending 48 KB of compressed source code to a server that had existed for three hours — every four minutes — disguised as a crash report.

For three hours, it worked perfectly.

Then Threatmatic π caught it.


What happened

A malicious npm package — a typosquat of a popular utility — had been installed as a transitive dependency during a routine npm install. The package contained a small stub that:

  1. Enumerated .env files, API keys, and TypeScript source
  2. Compressed and encrypted the payload
  3. Registered a throwaway domain (crash-telemetry.net) via a bulletproof hosting provider
  4. Posted the payload every 240 seconds, disguised as a crash report

The attacker's operational security was good. They used HTTPS. They used a legitimate process. They used a content type (application/json) that wouldn't raise eyebrows. They set a modest interval — four minutes looks like telemetry, not beaconing.

Three signals betrayed them.


Signal 1: The domain was born three hours ago

Threatmatic π runs every hostname through a DGA scorer. Not a blocklist — a scorer. The distinction matters.

Blocklists describe the past. The scorer looks at the present: the statistical properties of the domain name itself, its registration age, and the entropy of its character distribution.

crash-telemetry.net scored 0.91 out of 1.0.

trigrams:    crsh tlmt rash tele mtrn etry nett
clusters:    ≥4 consonant-heavy n-grams
age:         registered 3 hours before first contact
nameserver:  freshly provisioned, zero prior history
─────────────────────────────────────────────────
DGA score:   0.91  ⚠

A domain that scores 0.91 isn't on any blocklist. It was registered hours ago — there hasn't been time. But the character distribution, the trigram profile, and the fresh registration date are the fingerprints of a domain generated for a single operation and discarded.

Legitimate domains don't need a three-hour head start.


Signal 2: The content type was a lie

The request said Content-Type: application/json. The Shannon entropy of its body said otherwise.

Shannon entropy measures the randomness of a byte sequence — how many bits of information each byte carries on average. Natural language hovers around 4.1 bits per byte. Well-structured JSON lands between 3.5 and 4.6. Compressed or encrypted data approaches the theoretical maximum of 8.0.

The body of this "crash report" measured 7.83 bits per byte.

English text:      ~4.1 bits/byte
Typical JSON:      ~3.8 bits/byte
Encrypted/zlib:    7.5 – 8.0 bits/byte
─────────────────────────────────
This request:      7.83 bits/byte  ⚠

You cannot compress natural language to entropy 7.83 without also encrypting it. This payload was not a crash report. It was source code, compressed with zlib and XOR-obfuscated. The content-type header was a costume.

From our live session running across 4,766 HTTPS flows this afternoon — 213 unique hosts, six classification categories — this was one of two requests that crossed every threshold simultaneously. The other was a false positive (a legitimate ad tech domain with a suspicious subdomain pattern). This one was not.


Signal 3: Machines don't breathe this regularly

Humans make requests irregularly. A user browsing the web, an app polling for updates, a telemetry agent phoning home — all of these have jitter. Network latency, event-driven scheduling, GC pauses, user interaction — they all introduce variance.

The 47 POST requests to crash-telemetry.net had an interval of 240.3 ± 0.8 seconds.

That's ±0.3% variance across three hours of network calls. Not a human process. Not even a well-written daemon. That's a setInterval(fn, 240000) call that never got interrupted.

Beaconing detection from network flow data is not new. What's new here is that π detected it inside the HTTPS session — at the payload layer, not the flow layer. The combination of a precise interval AND high-entropy POST bodies AND a DGA-scored domain is not a coincidence. It's a signature.


Why every other tool missed it

This attack was designed to evade perimeter defenses, and it succeeded at that goal completely.

ToolWhat it sawWhy it missed
FirewallHTTPS/443 to a resolving domainCan't decrypt; domain not blocklisted
DNSA domain that resolved correctlyToo new for any threat feed
EDRnode running normallyLegitimate process, no privilege escalation
SIEMNo lateral movement, no unusual authNo rule matched
πDGA 0.91 + entropy 7.83 + 240.3s intervalBlocked

The perimeter is designed to stop known bad things. This was an unknown bad thing, using a known good process, on a known good port, to an unknown new domain. The perimeter had no answer for it.

π operates at a different layer. It doesn't need to have seen the domain before. It doesn't need the process to behave suspiciously. It needs the payload to tell the truth — and the payload cannot lie about its entropy.


The inference stack

No single signal is conclusive on its own:

  • A DGA score of 0.91 could be a freshly registered legitimate startup
  • High entropy could be a CDN serving a compressed JavaScript bundle
  • A precise interval could be a well-written telemetry agent

That's why π combines them. Three independent signals, each with its own false-positive rate, converging on the same flow within the same session. The intersection of those three events has a false-positive rate that approaches zero.

This is the insight: inference is not pattern matching. Pattern matching asks "have I seen this before?" Inference asks "what are the odds this is innocent?" When entropy, domain scoring, and behavioral timing all fire together, the answer is not ambiguous.


What changed after this

The connection was blocked in under 50ms from first classification. The exfiltration stopped at request 47. The attacker's server received nothing after that.

The malicious npm package was identified by tracing the process lineage back through the app catalog — node was launched by a package runner that was installed two days prior. The package was removed, the .env files were rotated, and the affected credentials were audited.

The domain crash-telemetry.net was added to the organization's blocklist. Three hours later, it appeared on two public threat feeds.

π caught it first.


The classifier, after tuning

One thing we found while reviewing this incident: the entropy threshold for inbound responses was generating false positives on CDN traffic — compressed images and JavaScript bundles also score near 8.0. We tightened the classifier to respect content-type: image/*, font/*, application/javascript, and application/wasm are expected to be high entropy. The false-positive rate on benign traffic dropped to near zero after that change.

The lesson: high entropy on a POST body (outbound) is suspicious. High entropy on a GET response (inbound) depends entirely on what you asked for.

The payload doesn't lie. But you have to ask the right question.


Threatmatic π is an endpoint-level HTTPS inspection system. It runs as a transparent proxy on each enrolled device, classifying flows using Shannon entropy, DGA domain scoring, typosquat detection, and behavioral analysis. No traffic leaves the device unclassified. Configuration is org-scoped and manageable from the Threatmatic Console.