fix(chrome): drop heuristic in-process extractors, keep discovery only

User testing of --chrome-process-scan on a VMware Win10 Edge session
exposed the heuristic's structural limits. The password matcher captured
URL-path fragments as usernames (`internal/`, `api/v1/`), CJK noise from
random bytes read as UTF-16, and chrome.dll auth-flow constants
(login.microsoft.com etc) — 71 reported "passwords", 0 real. The cookie
matcher fared similarly: 99k hits whose top hosts looked real
(americanexpress.com, youtube.com, apartments.com) but whose individual
rows were dominated by UUID-prefixed concatenated hostnames, JS-token
names (`typeof`, `viz.mojom.GpuHostMessageHeader`) and minified-JS
values (`Symbol&&Symbol.`, `||void 0===t||t`).

Tighter and tighter filters reduced the noise (TLD allowlist; RFC 6265
token-shape names; URL-fragment + auth-noise + JS-keyword rejection;
mixed-class value enforcement; dedup) but each pass either still leaked
thousands of false positives or rejected real cookies. The structural
problem is that chrome process memory contains too many sequences that
look like `domain\0name\0value` without being one; the heuristic has no
way to tell a CanonicalCookie instance from a string-table entry.

Strip both extractors from process_scan. The flag now walks every
running chromium process through the page-table region enumerator and
emits one info-level log line per process ("PID/image/MiB resident")
plus a discovery total — useful as "a browser was active when the
snapshot was taken; here's how much RAM it had touched" intel, no fake
data. ChromeFindings stays empty for the memory side, so disk-side
results aren't polluted.

The fix going forward is the per-Chrome-version `CookieMonster` locator
signature (chrome.dll destructor pattern → vtable → heap scan → RB-tree
walk). The CanonicalCookie struct layouts and tree walker remain in
`cookie_monster.rs` ready for that work — only the locator pattern is
missing, and adding it is non-noisy (it either finds a real
CookieMonster instance or it doesn't).

Tests: 86/86 pass. VMware Win10 hybrid baseline (83 cookies + 1
password from disk) preserved with --chrome-process-scan toggled either
way. Docs updated to match: README "in-process memory scan" section
becomes "in-process discovery", architecture.md adds the rationale,
examples.md replaces the fake-results sample with a real -v discovery
trace.
This commit is contained in:
NK
2026-06-07 19:04:11 +02:00
parent ae464e92e3
commit 899da8acf1
6 changed files with 357 additions and 176 deletions
+21 -15
View File
@@ -179,19 +179,25 @@ primitives. Four extraction vectors are supported:
keyring via `ComposedResolver`. This unlocks profiles whose user password
isn't in LSA secrets, and recovers v20 keys that depend on
SYSTEM-context-user MKs which only the elevation service can produce.
- **In-process memory scan** (`--chrome-process-scan`, opt-in) — walks every
chrome.exe / msedge.exe / brave.exe in the snapshot, dumps each process's
mapped userland through the page-table walker, and runs heuristic
pattern matchers for `https://` URLs followed by username/password pairs
and ASCII cookie domains. ChromeKatz-style; this is the lever that
recovers what's *in flight* in browser memory — including the plaintext
passwords Edge ≤ 147 holds in memory for the whole session ([Rønning,
April 2026](https://www.threatlocker.com/blog/microsoft-edge-is-keeping-your-passwords-in-plaintext-memory-heres-what-that-actually-means)).
Results are heuristic and noisy; the disk-side path remains the
high-fidelity reference. The precise per-Chrome-version `CookieMonster`
locator (next iteration) will replace the heuristic with structured
extraction; `src/chrome/cookie_monster.rs` already carries the matching
`CanonicalCookie` struct layouts.
- **In-process discovery** (`--chrome-process-scan`, opt-in) — walks every
chrome.exe / msedge.exe / brave.exe in the snapshot through the
page-table region enumerator and logs each process's PID, image and
mapped userland size. Currently a discovery-only signal ("a browser was
active when the snapshot was taken; here's how much RAM it had
touched"); structured cookie / password extraction is queued behind
per-Chrome-version `CookieMonster` locator signatures (the
`CanonicalCookie` struct layouts are already in
`src/chrome/cookie_monster.rs`, only the locator pattern is missing).
An earlier heuristic-based extractor was retired because chrome process
memory contains huge amounts of minified-JavaScript string tables and
chrome.dll auth-flow constants that look syntactically identical to
cookie or credential data once isolated from their structural context —
every filter pass either still leaked thousands of false positives
(which the user would have to triage) or rejected real cookies too. The
upcoming locator-signature path will recover what's in flight in
browser memory accurately — including the plaintext passwords Edge ≤
147 holds in memory for the whole session ([Rønning, April
2026](https://www.threatlocker.com/blog/microsoft-edge-is-keeping-your-passwords-in-plaintext-memory-heres-what-that-actually-means)).
- **Memory-only** — limited; without disk access the encrypted SQLite files
are unreadable, so this path is mostly useful for pivoting MKs to a later
disk-mode run.
@@ -212,8 +218,8 @@ binaries still decrypt.
# Hybrid mem+disk: best yield for v20 ABE cookies
./vmkatz --chrome --disk disk.vmdk snapshot.vmsn
# Hybrid + ChromeKatz-style in-process scan of chrome.exe / msedge.exe
./vmkatz --chrome --chrome-process-scan --disk disk.vmdk snapshot.vmsn
# Hybrid + chromium process discovery (logs each running browser PID+size)
./vmkatz -v --chrome --chrome-process-scan --disk disk.vmdk snapshot.vmsn
# Structured output for tooling
./vmkatz --chrome --chrome-json disk.vmdk
+34 -21
View File
@@ -152,29 +152,42 @@ the one whose decrypted output validates against the expected layer shape.
This is defensive: `decrypt_blob` has no HMAC verify, so wrong MKs silently
produce garbage that is only caught by the next layer's parse.
### In-process memory scan (opt-in)
### In-process discovery (opt-in)
The `--chrome-process-scan` flag enables a fourth extraction vector that
runs alongside the disk path: for every running `chrome.exe` / `msedge.exe`
/ `brave.exe` / `vivaldi.exe` / `opera.exe` in the snapshot,
`paging::regions::enumerate_user_regions` walks the process's page tables
top-down to enumerate every mapped 4 KiB page in the canonical low half,
`ProcessMemory` reads those pages via the existing DTB + page-table
walker, and `heuristic::scan_heap_for_passwords` / `_cookies` pattern-match
the resulting bytes for `https://` UTF-16 + username/password triples and
ASCII cookie tuples. Hits are tagged `ChromeSource::Memory { pid, process }`
and merged into the same `ChromeFindings` the disk path produces.
The `--chrome-process-scan` flag walks every running `chrome.exe` /
`msedge.exe` / `brave.exe` / `vivaldi.exe` / `opera.exe` in the snapshot
through `paging::regions::enumerate_user_regions` (page-table top-down
traversal of the canonical low half) and reads each region via
`ProcessMemory`. Currently it emits one info-level log line per process
("PID/image/MiB mapped") and a discovery total; no entries are written
into `ChromeFindings`.
The vector is opt-in because the heuristic produces noise — `plausible_cookie`
gates on a small TLD allowlist and an RFC 6265 token-shape cookie name to
trim the worst false-positive classes, but the in-process scan still hits
debug strings, URL templates and HTTP cache entries. Its real value is
catching what isn't on disk yet: in-flight session cookies, autofill state,
and (for Edge ≤ 147) the full plaintext password vault that Edge keeps
mapped for the whole browser session. The struct layouts in
`cookie_monster.rs` (`CanonicalCookieChrome130`, `…Edge130Pb`, etc.) are
ready for the upcoming per-Chrome-version locator signature that will
replace the heuristic with structured `CookieMonster` walking.
The flag previously merged heuristic password / cookie hits from
`heuristic::scan_heap_for_passwords` / `_cookies` into the disk-side
findings. That path was retired because chrome process memory is filled
with minified-JavaScript string tables and chrome.dll auth-flow
constants that look syntactically identical to cookie or credential
bytes once isolated from their structural context. On VMware Win10 with
Edge browsing live the heuristic emitted ~99k "cookies" — top hosts
included real domains the user had visited (americanexpress.com,
youtube.com, apartments.com) but individual rows were dominated by
UUID-prefixed hostnames, JS-token cookie names (`typeof`,
`viz.mojom.GpuHostMessageHeader`) and minified-JS values
(`Symbol&&Symbol.`, `||void 0===t||t`). Every filter tightening pass
either still leaked thousands of false positives or rejected real
cookies too — so the honest choice is to ship discovery only and let
the next iteration replace it with structured walking.
The struct layouts in `cookie_monster.rs` (`CanonicalCookieChrome`,
`CanonicalCookieChrome130`, `CanonicalCookieEdge130`, plus the
`ProcessBoundString` variants for Chrome 130+ in-memory cookie value
encryption) are ready for the per-Chrome-version `CookieMonster` locator
signature: pattern-match the destructor in chrome.dll, resolve the
vtable from the resulting code address, scan the heap for objects whose
first qword equals that vtable, walk the `std::map` red-black tree from
the resulting CookieMonster. Until that signature work lands, the disk
DPAPI chain remains the high-fidelity reference and the only
structurally-reliable extractor.
See [`docs/plans/2026-06-05-chrome-module-design.md`](plans/2026-06-05-chrome-module-design.md)
for the full design spec.
+28 -21
View File
@@ -108,34 +108,41 @@ $ vmkatz --chrome --disk windows.vmdk snapshot.vmsn
[+] Chrome findings: 1 password, 83 cookies, 0 autofill
```
### In-process scan (`--chrome-process-scan`, opt-in)
### In-process discovery (`--chrome-process-scan`, opt-in)
Walks every running `chrome.exe` / `msedge.exe` / `brave.exe` in the
snapshot, dumps each one's mapped userland through the page-table walker,
and runs heuristic pattern matchers for `https://` URL + username/password
triples and ASCII cookie tuples. Hits are tagged with the originating PID
and process image, and merged into the same `ChromeFindings` document
the disk path produces:
snapshot through the page-table region enumerator and logs each
process's PID, image and resident memory size. Pair with `-v` to see
the discovery lines:
```
$ vmkatz --chrome --chrome-process-scan --disk windows.vmdk snapshot.vmsn
[INFO] [chrome-mem] PID 5688 msedge.exe (115 MiB): +68 passwords, +12345 cookies
[INFO] [chrome-mem] PID 7732 msedge.exe (7 MiB): +0 passwords, +1654 cookies
[INFO] [chrome-mem] scanned 5 chromium processes
...
$ vmkatz -v --chrome --chrome-process-scan --disk windows.vmdk snapshot.vmsn
[INFO] [chrome-mem] PID 5688 msedge.exe (115 MiB mapped userland) — discovery only; structured CookieMonster walking is queued behind per-Chrome-version locator signatures
[INFO] [chrome-mem] PID 7732 msedge.exe (7 MiB mapped userland) — discovery only; ...
[INFO] [chrome-mem] PID 8372 msedge.exe (23 MiB mapped userland) — discovery only; ...
[INFO] [chrome-mem] discovered 5 chromium process(es); no in-memory cookies/passwords emitted (heuristic was structurally unreliable, signature locator pending)
```
This is the lever that recovers in-flight values and the plaintext
password vault Edge ≤ 147 keeps mapped for the whole session — see
[Rønning's April 2026 disclosure](https://www.threatlocker.com/blog/microsoft-edge-is-keeping-your-passwords-in-plaintext-memory-heres-what-that-actually-means).
The flag's `findings` output is intentionally empty today: an earlier
heuristic-based extractor (search for `https://` UTF-16 plus the next
two strings; search for ASCII domains plus the next four strings) was
removed because chrome process memory is filled with minified-JavaScript
string tables and chrome.dll constants that look syntactically identical
to cookie or credential data once isolated from their structural context.
The triples it produced were dominated by URL-path-fragment usernames
(`internal/`, `api/v1/`) and JS-token cookie values (`Symbol&&Symbol.`,
`||void 0===t||t`) — every filter pass either still leaked thousands of
false positives or rejected real cookies too.
The scan is heuristic-driven (TLD allowlist + RFC 6265 token-shape cookie
names) and noisy; the disk DPAPI path remains the high-fidelity source.
The upcoming per-Chrome-version `CookieMonster` locator signature will
replace the heuristic with structured walking of the in-process cookie
store; the `CanonicalCookie` struct layouts that signature work targets
already live in `src/chrome/cookie_monster.rs` (ported from
[ChromeKatz](https://github.com/Meckazin/ChromeKatz)).
The proper fix is the per-Chrome-version `CookieMonster` locator
signature ChromeKatz uses: pattern-match the destructor in chrome.dll,
resolve the vtable, scan the heap for objects with that vtable, then
walk the `std::map` red-black tree to read each `CanonicalCookie`. The
[struct layouts in `src/chrome/cookie_monster.rs`](https://github.com/nikaiw/VMkatz/blob/dev/src/chrome/cookie_monster.rs)
are ready; only the locator pattern is missing. Once it lands the same
flag will return real cookies and the plaintext passwords Edge ≤ 147
keeps mapped for the whole session ([Rønning, April
2026](https://www.threatlocker.com/blog/microsoft-edge-is-keeping-your-passwords-in-plaintext-memory-heres-what-that-actually-means)).
### JSON output
+203 -11
View File
@@ -68,13 +68,138 @@ fn read_utf16le_until_nul(mem: &[u8], at: usize, max_chars: usize) -> Option<Str
Some(String::from_utf16_lossy(&units))
}
/// Hosts that appear in chrome/edge auth-flow string tables but are *not*
/// real saved-credential URLs. The heuristic finds `https://...` strings
/// anywhere in memory; without filtering these we get thousands of
/// false-positive triples whose "URL" is a Microsoft/Xbox auth endpoint
/// and whose "username" / "password" are whatever bytes happened to
/// follow in the DLL.
/// Hosts that appear in chrome/edge auth-flow and internal-API string
/// tables but are *not* real saved-credential URLs. The heuristic finds
/// `https://...` strings anywhere in memory; without filtering these we
/// get thousands of false-positive triples whose "URL" is a
/// Microsoft/Xbox/Apple auth endpoint or an Edge-internal API and whose
/// "username" / "password" are whatever bytes happened to follow.
const AUTH_NOISE_HOSTS: &[&str] = &[
"login.microsoft.com",
"login.microsoftonline.com",
"login.windows.net",
"login.live.com",
"xsts.auth.xboxlive.com",
"user.auth.xboxlive.com",
"device.login.microsoftonline.com",
"accounts.google.com",
"oauth.googleusercontent.com",
"appleid.apple.com",
];
/// Host suffixes that flag the URL as an internal Microsoft / chrome.dll
/// API endpoint rather than a user-saved credential URL.
const NOISE_HOST_SUFFIXES: &[&str] = &[
".cdp.microsoft.com", // Edge CDP / Connected Device Platform
".edgesv.microsoft.com", // Edge service backend
".windows.com", // generic MS svcs that show up in strings
".microsoftonline.com", // AAD / Office 365 backend
".live.com", // Xbox / Live backends
".googleusercontent.com", // Google CDN / OAuth content
".gstatic.com", // Google static asset CDN
".chrome.com", // Chrome telemetry / sync
"chromewebstore.googleapis.com",
"clients.google.com",
"update.googleapis.com",
];
fn url_host(url: &str) -> &str {
let s = url.strip_prefix("https://").unwrap_or(url);
s.split('/').next().unwrap_or(s)
}
fn host_is_auth_noise(url: &str) -> bool {
let host = url_host(url);
if AUTH_NOISE_HOSTS.iter().any(|h| host == *h) {
return true;
}
NOISE_HOST_SUFFIXES.iter().any(|suf| host.ends_with(suf))
}
/// Username heuristic: real saved usernames are email addresses or
/// alphanumeric handles. They never contain `/` (that's a URL path
/// fragment), never start with `http`, and aren't pure CJK noise.
fn looks_like_username(s: &str) -> bool {
if s.len() < 3 || s.len() > 128 {
return false;
}
// URL fragments captured after a null terminator: `internal/`,
// `api/v1/`, etc. Real usernames never contain `/`.
if s.contains('/') {
return false;
}
if s.starts_with("http://") || s.starts_with("https://") || s.contains("://") {
return false;
}
// 80%+ ASCII to reject CJK / random-bytes-as-UTF-16 noise.
let ascii_count = s.chars().filter(|c| c.is_ascii()).count();
if ascii_count * 100 / s.chars().count().max(1) < 80 {
return false;
}
// Real usernames are email-shaped or alphanumeric handles. Require at
// least 3 alphanumeric chars to drop pure-punctuation strings.
let alnum = s.chars().filter(|c| c.is_ascii_alphanumeric()).count();
if alnum < 3 {
return false;
}
// Email shape (`a@b.c` with a TLD-shaped tail) or all-tokenchars
// identifier (letters / digits / `_-.+`). Reject anything else.
let is_email = s.contains('@')
&& s.matches('@').count() == 1
&& s.split_once('@')
.map(|(local, domain)| !local.is_empty() && domain.contains('.'))
.unwrap_or(false);
let is_handle = s
.chars()
.all(|c| c.is_ascii_alphanumeric() || matches!(c, '_' | '-' | '.' | '+'));
is_email || is_handle
}
fn looks_like_password(s: &str) -> bool {
if s.len() < 6 || s.len() > 128 {
return false;
}
// Real saved passwords aren't URLs or URL fragments.
if s.contains("://") || s.starts_with('/') || s.starts_with("http") {
return false;
}
// 80%+ ASCII to reject CJK / random-bytes-as-UTF-16 noise.
let ascii_count = s.chars().filter(|c| c.is_ascii()).count();
if ascii_count * 100 / s.chars().count().max(1) < 80 {
return false;
}
// Real passwords contain at least one letter and one digit OR
// special character — pure-letter "passwords" of length 6+ in
// process memory are overwhelmingly debug strings, function names,
// or HTTP method tokens.
let has_letter = s.chars().any(|c| c.is_ascii_alphabetic());
let has_digit_or_special = s
.chars()
.any(|c| c.is_ascii_digit() || (c.is_ascii_punctuation() && c != '/'));
if !has_letter || !has_digit_or_special {
return false;
}
// No whitespace inside the value — real passwords occasionally have
// spaces, but those are very rare and indistinguishable from a
// function-arg-style "arg1 arg2" capture; better to drop them.
!s.chars().any(|c| c.is_whitespace() || c.is_control())
}
fn plausible(t: &PasswordTriple) -> bool {
t.url.starts_with("https://")
&& t.url.len() < 2048
&& !t.username.is_empty() && t.username.len() < 256
&& !t.password.is_empty() && t.password.len() < 256
&& !t.url.contains('\u{FFFD}')
&& !host_is_auth_noise(&t.url)
&& !t.username.contains('\u{FFFD}')
&& !t.password.contains('\u{FFFD}')
&& looks_like_username(&t.username)
&& looks_like_password(&t.password)
}
#[derive(Debug, Clone)]
@@ -84,10 +209,14 @@ pub struct CookieTriple {
pub value: String,
}
/// Scan ASCII host strings (domain-looking) followed by a cookie name and value
/// within a small window.
/// Scan ASCII host strings (domain-looking) followed by a cookie name and
/// value within a small window. Dedupes by `(host, name, value)` because
/// chrome.dll's string tables hold many copies of the same config key
/// triples (telemetry enum values, AAD scope names, etc).
pub fn scan_heap_for_cookies(mem: &[u8]) -> Vec<CookieTriple> {
let mut out = Vec::new();
let mut seen: std::collections::HashSet<(String, String, String)> =
std::collections::HashSet::new();
let mut i = 0usize;
while i < mem.len() {
if mem[i] == b'.' || mem[i].is_ascii_alphabetic() {
@@ -112,7 +241,14 @@ pub fn scan_heap_for_cookies(mem: &[u8]) -> Vec<CookieTriple> {
value: strs[1].clone(),
};
if plausible_cookie(&triple) {
out.push(triple);
let key = (
triple.host.clone(),
triple.name.clone(),
triple.value.clone(),
);
if seen.insert(key) {
out.push(triple);
}
}
}
i += len + 1;
@@ -175,8 +311,26 @@ fn host_has_common_tld(host: &str) -> bool {
COMMON_TLDS.iter().any(|t| *t == tld)
}
/// Returns true if `s` contains any substring that strongly suggests it's
/// a URL or domain fragment rather than a cookie name or value (`.com`,
/// `.net`, `://`, etc.). Used to drop heuristic captures where the
/// `read_ascii_cstr` loop walked past a struct boundary and concatenated
/// a UUID with a domain.
fn contains_url_fragment(s: &str) -> bool {
let lower = s.to_ascii_lowercase();
if lower.contains("://") {
return true;
}
for tld in [".com", ".net", ".org", ".io", ".co.", ".edu", ".gov", ".de.", ".fr.", ".uk.", ".cn"] {
if lower.contains(tld) {
return true;
}
}
false
}
fn looks_like_cookie_name(s: &str) -> bool {
if s.is_empty() || s.len() > 96 {
if s.len() < 2 || s.len() > 48 {
return false;
}
// Cookie names per RFC 6265: token characters
@@ -187,19 +341,57 @@ fn looks_like_cookie_name(s: &str) -> bool {
if !(first.is_ascii_alphabetic() || first == b'_') {
return false;
}
s.bytes().all(|b| {
if !s.bytes().all(|b| {
b.is_ascii_alphanumeric() || matches!(b, b'_' | b'-' | b'.' | b'~' | b'#' | b'$')
})
}) {
return false;
}
// Reject UUID-or-domain-fragment shapes. Real cookie names never
// embed `.com`-shaped fragments; if we see one, the heuristic
// walked past a struct boundary.
!contains_url_fragment(s)
}
fn looks_like_cookie_value(s: &str) -> bool {
if s.len() < 8 || s.len() > 4096 {
return false;
}
if contains_url_fragment(s) {
return false;
}
// Real cookie values are session IDs, base64 / hex / URL-encoded
// payloads, JWTs, etc. They almost always mix at least two of
// {letter, digit, special-char} — pure-letter values of length 8+
// are overwhelmingly debug strings or function names.
let has_letter = s.chars().any(|c| c.is_ascii_alphabetic());
let has_digit = s.chars().any(|c| c.is_ascii_digit());
let has_special = s
.chars()
.any(|c| c.is_ascii_punctuation() && !matches!(c, '/' | '\\'));
let classes = [has_letter, has_digit, has_special]
.iter()
.filter(|b| **b)
.count();
classes >= 2
}
fn host_for_noise_check(host: &str) -> &str {
host.trim_start_matches('.')
}
fn cookie_host_is_noise(host: &str) -> bool {
let h = host_for_noise_check(host);
AUTH_NOISE_HOSTS.iter().any(|n| h == *n)
|| NOISE_HOST_SUFFIXES.iter().any(|suf| h.ends_with(suf))
}
fn plausible_cookie(t: &CookieTriple) -> bool {
t.host.contains('.')
&& t.host.len() < 256
&& host_has_common_tld(&t.host)
&& !cookie_host_is_noise(&t.host)
&& looks_like_cookie_name(&t.name)
&& !t.value.is_empty()
&& t.value.len() > 8
&& t.value.len() < 8192
&& looks_like_cookie_value(&t.value)
}
#[cfg(test)]
+14 -10
View File
@@ -16,13 +16,16 @@
//! the elevation service produces and only LSASS retains.
//! - **VMFS reader** ([`runner::run_reader`]): used by the ESXi VMFS-6 raw
//! reader so SAM extraction and chrome discovery share one disk handle.
//! - **In-process scan** ([`process_scan::scan_chromium_processes`]):
//! - **In-process discovery** ([`process_scan::scan_chromium_processes`]):
//! opt-in via `--chrome-process-scan`. Walks every chromium process's
//! mapped userland through the page-table region enumerator and runs
//! heuristic pattern matchers for `https://` URL/username/password
//! triples and ASCII cookie tuples. Picks up in-flight secrets that
//! never reach disk — including the plaintext password vault Edge ≤ 147
//! keeps mapped for the whole session. Inspired by Meckazin/ChromeKatz.
//! mapped userland through the page-table region enumerator and logs a
//! discovery summary (PID, image, MiB resident). Earlier iterations
//! merged heuristic password/cookie hits into the disk findings; that
//! path was removed because chrome process memory contains too much
//! minified-JS string-table data that pattern-matches as cookies
//! without being one. Structured `CookieMonster` walking via per-
//! Chrome-version locator signatures is the queued follow-up; the
//! `CanonicalCookie` struct layouts ([`cookie_monster`]) are ready.
//!
//! ## Module layout
//!
@@ -54,13 +57,14 @@
//! scaffold).
//! - [`process_scan`] — `--chrome-process-scan` orchestrator: enumerate
//! chromium processes, dump each one's mapped userland through the
//! page-walk region enumerator, run the [`heuristic`] scanners, tag
//! findings with `ChromeSource::Memory { pid, process }`.
//! page-walk region enumerator, log a discovery summary. Returns an
//! empty [`types::ChromeFindings`] today; in-memory cookie/password
//! extraction is queued behind the per-Chrome-version locator
//! signature in `cookie_monster.rs`.
//! - [`cookie_monster`] — ported `CanonicalCookie` struct layouts +
//! `OptimizedString` reader + `std::map` RB-tree walker from
//! `ChromeKatz/CookieKatz/Memory.h`. Waiting for the per-Chrome-version
//! locator signature that will replace the heuristic scan with
//! structured `CookieMonster` traversal.
//! locator signature that will populate the in-memory extraction path.
//!
//! See [`docs/plans/2026-06-05-chrome-module-design.md`](../../docs/plans/2026-06-05-chrome-module-design.md)
//! for the original design spec.
+57 -98
View File
@@ -1,29 +1,41 @@
//! In-process scanner for chromium browser secrets.
//! In-process chromium memory scanner.
//!
//! Enumerates every chrome.exe / msedge.exe / brave.exe process in a memory
//! snapshot, classifies each (browser vs network-service vs other), walks
//! that process's mapped userland pages via the page-table-based region
//! enumerator, and runs the existing heuristic password/cookie scanners
//! over the collected bytes.
//! Enumerates every chrome.exe / msedge.exe / brave.exe / vivaldi.exe /
//! opera.exe in a memory snapshot, walks the process's mapped userland
//! through the page-walk region enumerator, and reports per-process
//! statistics (PID, image name, resident memory size) so an analyst knows
//! a browser was active when the snapshot was taken.
//!
//! When this scan is enabled (hybrid mode with both `--disk` and a memory
//! snapshot), its findings are merged with the disk-side findings under a
//! single [`crate::chrome::types::ChromeFindings`]. Each result is tagged
//! `ChromeSource::Memory { pid, process }` so the source is visible in the
//! output renderers.
//! Earlier iterations of this scanner also ran the
//! [`crate::chrome::heuristic`] ASCII / UTF-16 pattern matchers on the
//! collected bytes and merged hits into the disk-side findings. Validation
//! showed the heuristic produces structurally-unreliable triples: chrome
//! process memory contains huge amounts of minified-JavaScript string
//! tables, embedded HTML, and chrome.dll auth-flow constants that look
//! syntactically identical to cookie or credential bytes once isolated
//! from their structural context. With no way to distinguish a real
//! `CanonicalCookie` instance from a string table entry that *happens* to
//! pair a domain with an alphanumeric token, every filter we tried either
//! still leaked thousands of false positives or rejected real cookies
//! too.
//!
//! Per-version `CanonicalCookie` struct decoding (`cookie_monster::read_cookie`)
//! is not yet driven here — locating CookieMonster instances precisely
//! requires per-Chrome-major-version byte signatures that are queued as a
//! follow-up. Today's scanner uses the existing
//! [`crate::chrome::heuristic`] ASCII / UTF-16 patterns which work across
//! versions but produce flatter cookie / password records.
//! The correct fix is the per-Chrome-version `CookieMonster` locator
//! signature ChromeKatz uses — pattern-match the destructor in chrome.dll,
//! resolve the vtable, scan the heap for objects with that vtable, then
//! walk the `std::map` red-black tree to read each `CanonicalCookie`. The
//! struct layouts that signature work produces are already in
//! [`crate::chrome::cookie_monster`]. Until that locator lands, this
//! module deliberately ships *no* heuristic memory cookie or password
//! extraction — emitting only a discovery list ("browser is running, this
//! is the PID and how much RAM it has touched") is more honest than
//! pretending to recover credentials from cell tower auth-flow noise.
//!
//! When the flag is set, the discovery summary is logged at info level
//! and an empty [`ChromeFindings`] is returned so the downstream merge
//! step doesn't grow false positives.
use crate::chrome::heuristic::{scan_heap_for_cookies, scan_heap_for_passwords};
use crate::chrome::memory::is_chromium_image;
use crate::chrome::types::{
Browser, BrowserProfile, ChromeFindings, ChromeSource, Cookie, SavedPassword,
};
use crate::chrome::types::ChromeFindings;
use crate::error::Result;
use crate::memory::{PhysicalMemory, VirtualMemory};
use crate::paging::regions::enumerate_user_regions;
@@ -40,54 +52,51 @@ const MAX_BYTES_PER_PROCESS: usize = 1024 * 1024 * 1024;
/// scan budget.
const MAX_BYTES_PER_REGION: usize = 64 * 1024 * 1024;
/// Run the in-process chromium scan across every chrome.exe / msedge.exe
/// / brave.exe / vivaldi.exe / opera.exe in `processes`. Reads each
/// process's mapped userland through `phys` + the process DTB and runs
/// both the password (`https://` UTF-16) and cookie (ASCII-domain) scans
/// on it.
///
/// We don't yet read process command lines, so every chromium process is
/// treated as if it could hold either kind of secret. In practice the
/// main browser process holds passwords and the network-service utility
/// holds cookies — scanning renderers is wasted work, but they don't
/// contain the patterns we look for so they yield zero false positives.
/// Per-role gating via PEB ProcessParameters → CommandLine is a follow-up.
/// Walk every chromium process in `processes`, read its mapped userland,
/// and emit one info-level log line per process describing what was
/// found. Returns an empty [`ChromeFindings`] because the heuristic
/// extractors were retired (see the module docstring). The function
/// remains as the host for the future signature-based locator: once
/// `cookie_monster.rs` has per-Chrome-version patterns wired in, real
/// cookie / password findings will be returned here.
pub fn scan_chromium_processes<P: PhysicalMemory>(
phys: &P,
processes: &[Process],
) -> Result<ChromeFindings> {
let mut findings = ChromeFindings::default();
let mut scanned = 0;
for proc in processes {
if !is_chromium_image(&proc.name) {
continue;
}
let buf = match collect_process_memory(phys, proc.dtb) {
Ok(b) => b,
let mib = match collect_process_memory(phys, proc.dtb) {
Ok(b) => b.len() / 1024 / 1024,
Err(e) => {
log::info!(
"[chrome-mem] PID {} {} heap read failed: {}",
proc.pid, proc.name, e
"[chrome-mem] PID {} {} region scan failed: {}",
proc.pid,
proc.name,
e
);
continue;
}
};
let mib = buf.len() / 1024 / 1024;
let pw_before = findings.passwords.len();
let cookie_before = findings.cookies.len();
harvest(proc.pid as u32, &proc.name, &buf, &mut findings);
log::info!(
"[chrome-mem] PID {} {} ({} MiB): +{} passwords, +{} cookies",
"[chrome-mem] PID {} {} ({} MiB mapped userland) — \
discovery only; structured CookieMonster walking is queued \
behind per-Chrome-version locator signatures",
proc.pid,
proc.name,
mib,
findings.passwords.len() - pw_before,
findings.cookies.len() - cookie_before,
mib
);
scanned += 1;
}
log::info!("[chrome-mem] scanned {} chromium processes", scanned);
Ok(findings)
log::info!(
"[chrome-mem] discovered {} chromium process(es); no in-memory \
cookies/passwords emitted (heuristic was structurally unreliable, \
signature locator pending)",
scanned
);
Ok(ChromeFindings::default())
}
/// Collect every mapped userland page in a process's address space into a
@@ -120,53 +129,3 @@ fn collect_process_memory<P: PhysicalMemory>(phys: &P, dtb: u64) -> Result<Vec<u
}
Ok(out)
}
fn browser_from_image(image: &str) -> Browser {
let lower = image.to_ascii_lowercase();
match lower.as_str() {
"msedge.exe" => Browser::Edge,
"brave.exe" => Browser::Brave,
"opera.exe" => Browser::Opera,
"vivaldi.exe" => Browser::Vivaldi,
_ => Browser::Chrome,
}
}
fn memory_profile(pid: u32, image: &str) -> BrowserProfile {
BrowserProfile {
browser: browser_from_image(image),
user: String::new(),
profile_name: String::new(),
path: format!("memory:pid={}", pid),
}
}
fn harvest(pid: u32, image: &str, mem: &[u8], out: &mut ChromeFindings) {
let profile = memory_profile(pid, image);
let src = ChromeSource::Memory {
pid,
process: image.to_string(),
};
for t in scan_heap_for_passwords(mem) {
out.passwords.push(SavedPassword {
profile: profile.clone(),
url: t.url,
username: t.username,
password: t.password,
source: src.clone(),
});
}
for c in scan_heap_for_cookies(mem) {
out.cookies.push(Cookie {
profile: profile.clone(),
host: c.host,
name: c.name,
value: c.value,
path: "/".into(),
expires: None,
http_only: false,
secure: false,
source: src.clone(),
});
}
}