95 Commits
Author SHA1 Message Date
416rehmanandClaude Opus 5 6030ae113c feat(decompile): record whether a device exists without the hardware
Many drivers create their device object only once their hardware is
enumerated. On a machine without it the device never appears, so nothing
the driver exposes can be reached - and in the report that is
indistinguishable from the driver not being vulnerable. Confirming a
flagged driver meant loading it and finding out by hand.

Decompilation now records which function calls IoCreateDevice and
whether that function is DriverEntry or something it calls. A driver
created on the load path can be exercised on any machine that will load
it; one created elsewhere, typically a PnP add-device or start-device
callback, needs the physical device. Neither is recorded when the call is
not found, rather than guessing.

Resolved through the symbol table instead of decompiling every function
and matching text, so a call ghidra renders differently in C is still
found, and the pass costs nothing.

The report needs no change to show it: the processor declares the value
in `provides` and the bundled pipeline names it under `report.columns`,
which is the existing route for anything a pipeline wants surfaced. The
report stays unaware of what a device object is.

The decision is a plain function so it can be checked without ghidra, and
the whole path was run against a real driver: null.sys resolves to
on-load-path, which matches \.\NUL opening on any machine.

Closes #20

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:55:34 -06:00
416rehmanandClaude Opus 5 696a46b0e1 feat(engine): tell running out of time apart from going wrong
A stage that hit its budget was recorded as failed. Because the failure
is stored against the sample, a resumed run skipped it, so those samples
were excluded from every later stage for good. They are the largest ones
in a corpus, which are often the ones worth looking at, so the run
quietly biased itself away from its best targets and the report filed
them with genuine errors.

Running out of time is now its own outcome, carrying the budget it
exceeded. TimeoutError is an OSError, so the handler for it has to come
before the general one or it is caught as something breaking; a test
pins that ordering, and another drives a slow stage through the runner
to check what actually gets recorded.

`--retry-timeouts` attempts those samples again with a longer budget,
`--timeout-multiplier` deciding how much longer. Only the stages that ran
out of time are forgotten, so everything else the run produced is kept
and the retry costs the time those stages need and nothing else.

The report gives them their own group with the budget each exceeded, and
counts them apart from both errors and clear results, so a corpus that
was never finished is not reported as one that came back clean.

Closes #19

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:46:34 -06:00
416rehmanandClaude Opus 5 6aa9320eb2 fix(report): tell the two kinds of filter apart
The bar held seven identical chips in one run. Three of them narrow by
what the run concluded and four by what a person recorded, and the two
combine rather than being alternatives, so presenting them as one set
misread both what they do and how they work together.

They are now separate groups with a rule between them and the second one
named. The review chips take a different shape and carry a status dot
rather than a different colour, because colour already means outcome
here and a review is not an outcome. Stacked on a narrow screen the rule
between them turns with them.

The bar measures its own height for the table head to park under, and it
was only remeasured when the window resized. It also changes height when
its controls wrap, which a live run causes by itself as new outcomes
appear and add chips, leaving the head parked against a stale figure.
It now watches the bar itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:36:44 -06:00
416rehmanandClaude Opus 5 3fd55acc43 feat(report): record a result that was checked and did not hold up
A review had two outcomes: confirmed, or something outstanding. That left
nowhere to put the most expensive result there is - taking a finding to a
test machine, running it, and watching it not happen. Recorded as
outstanding it reads as unfinished work; recorded as nothing at all it is
indistinguishable from a result nobody has opened.

Reviews now have a third outcome for a result that did not hold up. It
opens the same note the outstanding state does, because what was run and
what happened instead is the part worth keeping, and the summary counts
each outcome so a run can be read as what is known rather than what was
alleged.

Closes #25

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:27:26 -06:00
416rehmanandClaude Opus 5 523ed1a738 feat(report): let the reader record what they have actually checked
A pipeline records what it concluded. Whether a person has verified that
conclusion is a separate claim, and nothing in a run can make it, so the
report now carries marks the reader adds: confirmed, or outstanding with
a note naming what is still unproven.

Deliberately generic. What counts as outstanding is the reader's
business, and reproducing a crash, proving a precondition or reading a
diff are the same shape of unfinished work as far as the report is
concerned, so it only provides somewhere to put it.

Marks show against each row and can be filtered on, including the results
nobody has reviewed yet. The report is a file, so they are kept in the
browser and keyed on the pipeline and target rather than the path, which
means a rerun over the same corpus keeps them. Exporting writes marks.json;
saved next to the report it is read back at generation time, so the marks
render for anyone who opens it and reach inventory.csv alongside
everything else instead of being stranded in one person's browser.

Closes #24

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:15:11 -06:00
416rehmanandClaude Opus 5 ccca456331 docs(decompile): record what decompilation writes and guard its shape
Decompilation output grows silently: it stays correct while getting
enormous, so the regression is only visible to someone who happens to
look at the directory size. The reference pipeline documentation now says
what the stage writes per driver and what it should total, with the
figures from the corpus where storing the routine per IOCTL code turned
28 MB into 734 MB.

The entry a code decodes to is now built by its own function, so a test
can hold it to carrying the code and its decoded fields and nothing that
could grow. One case covers a field large enough to be a routine appearing
under any name, since that is the shape of the regression rather than any
particular field.

Closes #22

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:01:35 -06:00
416rehmanandClaude Opus 5 6a6badf81b fix(scan): stop one bad batch from discarding a whole corpus of results
The scanner ran every sample through a single semgrep invocation, so any
failure in that run cost every result. On a real driver pack that meant
1,380 drivers and six minutes of work producing nothing, with the scan
stage gating assessment so the run quietly ended there.

Samples are now scanned in batches, configurable per stage and defaulting
to 200. A batch that times out or comes back unreadable marks only its own
samples failed; the rest still deliver their findings.

A scan that read none of the files it was given is also no longer recorded
as a clean result. semgrep lists the files it opened, and reading none of
them means the scan never looked, which is a different outcome from
looking and finding nothing - the second is safe to pass to assessment and
the first is not. Each sample now also records how many of its files were
submitted and how many were read, so a sample reporting nothing because
its files went unread can be told apart in the report.

Closes #21

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 17:52:35 -06:00
416rehmanandClaude Opus 5 5c0fbda75f fix(test): check the bundled prompt without needing a full environment
The test covering the shipped pipeline's prompt ran validate_pipeline,
which also insists on a Ghidra install and a signed-in model backend.
Neither says anything about whether the prompt is correct, so the test
passed only on a machine already set up to run the pipeline and failed
everywhere else.

It now resolves the pipeline's stages and checks the prompt against what
they declare, which is the property being tested and needs no
environment. Verified against a machine with neither dependency present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 15:02:46 -06:00
416rehmanandClaude Opus 5 1c937fa36e feat(pipeline): catch a prompt asking for values no stage produces
A prompt is rendered from whatever the stages before it recorded. A name
none of them produce renders as nothing, so the model is asked to judge
an empty payload and answers about the emptiness -- and that answer is
stored as the verdict. The run reports success and nothing indicates the
code was never seen.

Processors now declare the names they make available in `provides`: the
keys they record, plus one per artifact they write. `deepzero validate`
walks the stages in order and checks each prompt against what its
predecessors declare, reporting an unknown name along with the names that
were available instead. A value is only offered to stages after the one
recording it, so a prompt cannot reach its own stage's output.

This also closes a gap between validation and rendering. A prompt named
as a bare filename next to the pipeline passed validation, because the
file exists, but rendering only resolved references containing a path
separator -- so at run time the filename itself was sent as the whole
prompt. Both now resolve a reference the same way, and one shared rule
turns an artifact path into the name a prompt uses, so what is checked is
what gets rendered.

Documents the values a prompt receives, and how a processor declares the
ones it adds.

Closes #23

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 14:57:20 -06:00
416rehmanandClaude Opus 5 5d58b6afaf ci: split the pipeline by what each job actually needs
Every job in the matrix installed a JDK and downloaded a
several-hundred-megabyte Ghidra archive, then ran the linter and the
security scan again. Neither of those reads the interpreter, and only one
test reads Ghidra, so both costs were paid once per python version for
no additional signal.

The work is now split three ways:

- lint runs once, and installs the tools and the package's own light
  dependencies rather than everything under full
- the version matrix runs the suite without Ghidra, which is every test
  but one, and so needs no JVM at all
- a single integration job drives the real Ghidra install

Ghidra is cached between runs, keyed on the version and build it pins,
and unpacked under HOME so restoring it needs no privileges. Its
checksum is still verified whenever it is downloaded. The test that needs
it now carries a marker, so selecting it is declarative instead of a path
spelled out in the workflow.

A newer push cancels an in-flight run for the same branch, and the
linters are pinned to a minor so their own releases cannot fail a build
on their own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 14:41:37 -06:00
416rehmanandClaude Opus 5 d2fc5cea5a ci: test against python 3.13 and 3.14
The project declares support from 3.11 upward, so 3.13 and 3.14 were
already installable but never exercised. Both now run in the matrix,
alongside the versions that were there before.

A failing version no longer cancels the others. Across four interpreters,
knowing which ones broke is most of the answer, and stopping at the first
failure threw that away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 14:32:56 -06:00
416rehmanandClaude Opus 5 93619be980 refactor(report): name the two assessment outcomes for what they are
"Needs review" did not say what kind of review was missing. These are
results a scan flagged that no assessment stage has confirmed or
dismissed, so they are now "Needs assessment", matching the vocabulary
the rest of the report already uses.

That left two labels a reader could not tell apart, because the outcome
next to it was called "Not assessed" while describing results an
assessment had in fact produced a verdict for. What is missing there is
a verdict the pipeline recognises, so it is now "Unclear verdict".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 14:29:38 -06:00
416rehmanandClaude Opus 5 9aafd6a79c feat(report): rank results in one list and rebuild the palette
The index stacked a separate table per outcome, so sorting one of them
said nothing about the others even though risk ranks across the whole
corpus. Results now sit in a single list ordered by risk, with the
outcome as a filter alongside the name search, so a sort applies to
everything currently on screen.

Severity counts critical separately from high. The two differ by ten
times in the ranking, so folding them together left the column unable to
explain the order it was sorted in.

Colour now marks outcomes and nothing else: severity tiers, the outcome
of each result, and a corpus bar that dims to whatever the filter is
showing. Light and dark are built as one system, and both meet WCAG AA
on every surface a colour appears on, including the raised background
under a hovered row.

A live run reloads the page from script rather than a meta refresh,
keeping the reader's filter and scroll position instead of returning
them to the top every twenty seconds.

Fixes:

- the table head never stayed in view, because the scroll wrapper around
  the table was acting as the scroll container
- sort arrows rendered as an unrelated glyph
- results with no recorded verdict printed a raw HTML entity
- samples an assessment could not classify were counted as clear while
  also being listed separately as unassessed

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 14:25:40 -06:00
416rehmanandClaude Opus 5 9dc0fb388c fix: say whether a run is still going, not just what it last claimed
A run recorded its status when it started and never revisited it. A pipeline
that is killed, or whose machine restarts, never gets to write a closing
status, so "running" stayed on record indefinitely. The report repeated it,
kept its auto-refresh going, and presented a run that had ended hours earlier
as though results were still arriving.

A run now records which process is writing it, and the report checks that
process before repeating the claim. A run that recorded an outcome is taken at
its word; only one still claiming to be in progress is checked. Where the
answer cannot be established - another machine, a process that cannot be
queried - the report says so rather than guessing.

The page itself also ages. A live run rewrites it every few seconds, so once
that stops the gap between the page and the reader's clock gives it away, and
the status changes to say the page is no longer being updated.

The liveness probe deliberately avoids signalling the process: on Windows the
usual existence check terminates the target instead of testing it, which would
have made a report refresh kill the run it was describing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 11:21:41 -06:00
416rehmanandClaude Opus 5 e687c75e17 fix: show results as they land, and explain every outcome label
Results were only written out once fifty samples had finished, so a slow
stage left them invisible: a run that had already found and recorded a
vulnerable driver still reported none, because the finding had not been
written to the run's state yet. An interruption also discarded up to fifty
finished samples. State is now written on a short timer as well, so the
report reflects what has actually been found.

Each outcome label now explains on hover what it means and how a sample
came to be in it, including that a filtered item was excluded before
analysis finished rather than judged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:23:42 -06:00
416rehmanandClaude Opus 5 a2b471ec4d fix: say whether a sample was examined or excluded
The summary counted every sample that was not vulnerable as clean, which
included files a stage had excluded before any analysis ran. On a real
driver pack that put 950 files that are not kernel drivers, and were never
looked at, in the same count as drivers that were analysed and came back
clear.

Those are now separate: 'Clear' means analysed all the way through with
nothing flagged, and 'Filtered out' means a stage excluded it, naming the
stage that did so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:19:57 -06:00
416rehmanandClaude Opus 5 1ff9f7c79e feat: report a vulnerability with what is needed to reproduce it
An assessment used to end in prose, so acting on it meant reading the
paragraphs and reconstructing the request by hand. It now states plainly
that a separate step will try to reproduce whatever it reports on a real
machine, and a vulnerable verdict ends with the values that step needs:
the device to open, the IOCTL code, the input buffer laid out field by
field with values that exercise the flaw, the expected output size, and
the observable result that tells a real hit apart from the driver merely
accepting the request.

The assessment is also given the IOCTL codes recovered from the dispatch
routine, which it is asked to name but previously had to guess at.

A reduce stage ranks the samples it keeps, and later stages now work
through them in that order. Assessment of a large corpus takes hours, so
stopping early now leaves the highest-ranked samples already done rather
than an arbitrary subset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:16:54 -06:00
416rehmanandClaude Opus 5 caa3b0d05b fix: give the model the driver code it is asked to assess
The assessment prompt referred to values the pipeline does not produce.
Jinja renders an unknown name as empty text, so the prompt ended with
"Payload:" and nothing after it: every driver was assessed with no
decompiled code, no handler name and no findings list. The model said so
in its own answers - that there was nothing to analyse and any finding it
reported would be invented - and those verdicts were recorded as results.

The prompt now uses the names the pipeline actually records, so a driver
is assessed with its dispatch routine and its findings. On a driver from
a real pack this is the difference between a 13 character prompt and a
20,800 character one.

A prompt naming something the pipeline does not produce is now an error
that says which name was wrong and lists what is available, instead of
quietly asking the model to judge nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:04:29 -06:00
416rehmanandClaude Opus 5 0a9146b1f1 fix: stop scanning the same routine once per IOCTL code
Decompilation wrote one file per IOCTL code and copied the whole dispatch
routine into every one of them. Every code in a driver is handled by that
same routine, which is already written once to dispatch_ioctl.c.

On a real driver pack this produced 682 MB across 10,305 files where the
unique content was 18 MB, and the scanner then read the same function up
to 252 times for a single driver - counting every finding again for each
code, so a driver with many codes looked far worse than one with few and
ranked above it.

Each per-code file now records what the code decodes to and points at the
routine that handles it. Scan input for the same pack drops from 703 MB to
roughly 18 MB, and a finding is counted once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:03:46 -06:00
416rehmanandClaude Opus 5 210f49f43b fix: scan large corpora without losing every result
Scanning a real driver pack failed on all 1380 drivers at once. The
scanner read the results document back through a pipe, and at that size
the read fails, so a scan that had already run for six minutes produced
nothing and the whole assessment stage was skipped.

The scanner now has semgrep write its results to a file and reads that,
which does not depend on how large the document is. A launch failure also
no longer claims semgrep is missing when it is installed and something
else went wrong - the real error is reported.

Decompilation stored the whole dispatch routine again for every IOCTL code
it found, which produced 52 MB of output for a single driver and 734 MB
across the pack, nearly all of it the same text repeated. The routine is
recorded once, as it already was alongside it.

Assessment also reads artifacts into the prompt. A file far larger than
the context budget is now skipped with a warning instead of being parsed
and loaded on every sample, matching how oversized source files were
already handled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 10:03:46 -06:00
416rehmanandClaude Opus 5 0b1f89966c feat: assess every driver that survives scanning, worst first
The bundled kernel-driver pipeline kept only the ten highest-scoring
drivers before assessment and discarded the rest, so most of what
scanning flagged was never looked at. It now orders every surviving
driver by how many findings it has and assesses all of them, worst
first, and the stage is named for what it does.

A ranking stage asked to keep `0` items used to drop every sample and
report success, which made "no limit" the one value that silently
analysed nothing. Zero or fewer now keeps everything, still in metric
order, and the stage says how many it kept and dropped either way.

Assessment is pinned to a full model id rather than a short alias, so
the pipeline cannot quietly switch generations when an alias moves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:54:07 -06:00
416rehmanandClaude Opus 5 f2f8fe4ed9 feat: review a run's results in the browser with deepzero report
A run left its output spread across work/<pipeline>/samples/<id>/ as
per-sample state, findings, decompiled sources and assessments. Reviewing
thousands of those by hand is not practical, so `deepzero run` now prints
a link to a report and keeps it up to date as results land, and
`deepzero report` rebuilds it at any time.

The report answers one question first: what is vulnerable. Items an
assessment stage marked vulnerable lead the page, then items with
findings but no confirmed verdict, then anything that errored. Each item
links to its own page carrying the assessment, every finding with the
code it matched, and links to the artifacts on disk.

It is built from what a pipeline actually recorded rather than from any
one domain's field names, so a source-code review over repositories
renders as well as a kernel-driver review. A pipeline can shape the
presentation with an optional `report:` block - what to call one item,
which stage data key holds the verdict, which values mean vulnerable,
and which columns to surface - and every field has a default.

Output is layered so it stays usable on a large corpus: index.html holds
the triage summary at a bounded size, items/<id>.html covers everything
worth reading, inventory.csv carries every item for a spreadsheet, and
findings.jsonl carries every finding one per line. When a listing is
capped the page says what was capped and where the rest is.

Pages are self-contained with no network access, readable in light and
dark, keyboard navigable, and escape all analysed content.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:29:26 -06:00
416rehmanandClaude Opus 5 64f0003fe2 fix: keep already-analysed samples in the pipeline when a run resumes
A stage can report that a sample needs no work because its output is
already on disk. That sample was then treated as filtered out and never
reached any later stage, so resuming an interrupted run quietly analysed
fewer samples than a fresh run would - drivers dropped out before
scanning and assessment with nothing to indicate it.

Work that is already done now counts as passed and the sample continues,
which is what MapProcessor.should_skip documents. A stage that did no
work also records why in its own field, so a routine skip is no longer
reported as an error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:29:21 -06:00
416rehmanandClaude Opus 5 190aeb5158 feat: assess drivers with Opus in the bundled loldrivers pipeline
Vulnerability assessment of decompiled kernel drivers is the most
demanding step in the pipeline, so it now defaults to claude-code/opus.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:28:39 -06:00
416rehmanandClaude Opus 5 07c85f38b0 fix: tell the user how to fix Claude Code auth when it is not signed in
An auth failure surfaced the CLI's raw message ("401 OAuth access token
has been revoked") with no indication of what to do about it. DeepZero now
relays that message with the remedy attached: run `claude` in a terminal
and sign in.

DeepZero deliberately does not inspect Claude Code's credential store to
predict this. Whatever auth the CLI has is the auth DeepZero uses; if it
has none, the CLI reports it and we pass that along. Auth environment
variables are left untouched so the CLI authenticates exactly as it
normally would.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:28:33 -06:00
416rehmanandClaude Opus 5 d6023d612d feat: add a preflight check that catches an unauthenticated LLM before a run
A pipeline using Claude Code only calls the LLM in its final stage, so a
signed-out CLI wasn't discovered until after decompilation and scanning
had already run - potentially hours of wasted work. Passing `--preflight`
to `deepzero run` now sends one tiny probe to the configured backend
first and stops immediately with a clear "not authenticated" message if
it fails. It is opt-in and off by default, so normal runs spend nothing
extra, and a transient hiccup like a rate limit does not block the run.

Closes #14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 23:14:33 -06:00
416rehmanandClaude Opus 5 a9ed18f601 fix: cap decompilation concurrency so large runs don't exhaust memory
With `parallel: 0`, the decompile stage auto-scaled to one worker per
CPU core, and each worker is a separate Ghidra JVM holding a full
program database. On a many-core machine that meant dozens of JVMs and,
on a real driver corpus, exhausted RAM.

`parallel: 0` now derives a memory-safe worker count (roughly one worker
per 4 GiB of RAM, bounded by cores and a hard cap) instead of using
every core. An explicit `parallel: N` is still honored as-is, and the
ceiling can be raised on a large machine with `max_parallel:` in the
decompile stage config. Processors can now declare a concurrency ceiling
that only applies to auto-scaling.

Closes #12.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 23:14:33 -06:00
416rehmanandClaude Opus 5 dc2fee1d50 fix: give a clear error when the vulnerability scan can't find its rules
Addresses PR review feedback:

- The scan no longer falls back to an unintended config path when its
  rules directory can't be resolved; it fails fast with an explicit
  "rules_dir not found" instead of silently scanning with no rules.
- Per-call options passed to the LLM are now forwarded to the backend
  rather than dropped, so controls like `timeout=` take effect for the
  Claude Code CLI backend instead of being silently ignored. An unused
  internal flag that gated this is removed.
- Made a rules-resolution test independent of the working directory so
  it can't flake on a stray local `rules/` directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 23:00:37 -06:00
416rehmanandClaude Opus 5 9b41502198 feat: run the bundled loldrivers pipeline on Claude Code out of the box
The pipeline defaulted to vertex_ai/gemini-2.5-pro, which needs Vertex AI
credentials most users do not have, so a fresh checkout failed validation
before doing any work. It now defaults to claude-code/sonnet and runs
against a locally signed-in Claude Code with no cloud API keys. Override
per run with `-m` as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 22:56:39 -06:00
416rehmanandClaude Opus 5 ca3c0307c7 ci: exclude docs from ruff so a ruff upgrade cannot fail the build
CI installs the latest ruff, and newer versions format python snippets
embedded in markdown. That made `ruff format --check .` fail on
docs/en/extensibility/building-custom.md purely from a ruff upgrade,
with no source change. Documentation examples are prose, not build
artifacts, so they are now out of the formatter's scope.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:27:45 -06:00
416rehmanandClaude Opus 5 ec7ea0068b fix: stop the driver vulnerability scan from silently finding nothing
When run from the repo root, the scan pointed at the wrong rules
directory: the check that approves a run used one path while the run
itself used another. semgrep then loaded no rules and reported every
driver as clean, so real findings were silently missed. Both now resolve
the rules the same way, and a scan that loads no rules fails loudly
instead of looking like "no vulnerabilities found".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:20:29 -06:00
416rehmanandClaude Opus 5 cab855c278 fix: explain why a vulnerability scan failed instead of failing silently
The scanner treated any unexpected exit code from semgrep as fatal and
reported only its (often empty) error stream, producing an opaque
failure with no cause. It now trusts a complete scan result whatever the
exit code - semgrep can finish the scan yet exit oddly on Windows - and,
when a run genuinely fails, reports the exit code and error text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:16:46 -06:00
416rehmanandClaude Opus 5 981c6845ab fix: keep runs from crashing when printing results on Windows
DeepZero prints status glyphs such as checkmarks and arrows. On a legacy
Windows console, or whenever output was redirected to a file, printing
them raised an encoding error that aborted the run while writing the
final summary - after all the analysis work had already completed.
Output is now written as utf-8 with unrepresentable characters escaped
rather than fatal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:16:46 -06:00
416rehmanandClaude Opus 5 dd0e861e06 fix: decompile drivers that expose IOCTL handlers
Decompilation crashed for every driver that actually had IOCTL codes -
precisely the drivers worth investigating - while drivers with none
completed, which hid the problem. Writing the per-IOCTL output failed
inside Ghidra's Jython interpreter, which rejects byte strings on a
utf-8 text stream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:16:45 -06:00
416rehmanandClaude Opus 5 9330427649 fix: show a clear error when a required pipeline setting is missing
An unset ${VAR} with no default previously survived as the literal text
"${VAR}", which then read as a real value and produced confusing errors
such as "ghidra not found at ${GHIDRA_INSTALL_DIR}" instead of saying
the setting was required. Unset variables now expand to empty, matching
${VAR:-default} and ordinary shell behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:07:35 -06:00
416rehmanandClaude Opus 5 d83d1711e0 feat: run pipelines with your Claude Code subscription, no API key required
Pipelines can now use the Claude Code you already have installed and
signed in, so LLM analysis works without a third-party API key or
per-token billing:

    model: claude-code            # default model
    model: claude-code/sonnet     # or /opus, or a full model name

DeepZero never handles your credentials - it runs your own `claude`
binary in headless mode and inherits its existing sign-in. If the CLI is
missing or not signed in, validation says so before a run starts.

The pipeline feeds untrusted decompiled code to the model, so the
integration denies tools and external servers, sends the prompt over
stdin rather than the command line, and hides any ANTHROPIC_API_KEY from
the CLI so your subscription is used and not a metered API. Usage limits
retry with backoff; sign-in failures stop immediately with a clear
message.

Under the hood, LLM backends are selected from the model string by a
registry, so another agent CLI can be added later without touching the
engine, pipelines, or stages. Existing API-key models are unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 21:07:34 -06:00
416rehman 06e68c13ac refactor: simplify logging color schema and enforce proper signal handling encapsulation 2026-04-16 19:21:36 -04:00
416rehman 6ff7ad10b3 docs: use mark_interrupted 2026-04-16 19:16:44 -04:00
416rehman 603d01acef build: add python-dotenv to base dependencies 2026-04-16 19:13:10 -04:00
416rehman 9e71093e0f Merge branch 'refactor/cli-telemetry-and-metrics' of github.com:416rehman/DeepZero-Agentic-Vulnerability-Research-Pipeline into refactor/cli-telemetry-and-metrics
# Conflicts:
#	README.md
2026-04-16 19:13:10 -04:00
416rehman 98c3d6d25c docs: prominently mark REST API as WIP / experimental in CLI usage section 2026-04-16 19:09:54 -04:00
416rehman 82a4d2a886 docs: prominently mark REST API as WIP / experimental in CLI usage section 2026-04-16 19:09:15 -04:00
416rehman 908f30474e style: auto-format python code via ruff format 2026-04-16 19:06:20 -04:00
416rehman d99e4d333f style(cli): use crc32 hashing and expanded palette for wider color variance in logger module prefixes 2026-04-16 19:05:06 -04:00
416rehman ae48a94660 style(cli): dynamically color structural module prefix via deterministic sum-hashing 2026-04-16 19:04:01 -04:00
416rehman 8ac8ad9361 style(cli): use hard brackets for logger module prefix 2026-04-16 19:03:00 -04:00
416rehman 310114ba5c style(cli): condense and professionalize structural log formatting 2026-04-16 19:02:03 -04:00
416rehman 036bb1c303 fix(runner): flush active manifest metrics and state.json on forced sigint hard-kills 2026-04-16 18:59:21 -04:00
416rehman e7eca62cf3 style(cli): enforce uncolored dim style for unstarted status rows 2026-04-16 18:57:06 -04:00
416rehman 1a5ce38c04 fix(cli): bypass manifest cache and calculate stats directly from live sample files 2026-04-16 18:56:17 -04:00
416rehman 5de08efe69 style(cli): dim unstarted stages and zero values in status table to match tui 2026-04-16 18:54:04 -04:00
416rehman d60e30d1c1 fix(cli): align status table to unconditionally display all pipeline stages in sequence 2026-04-16 18:53:01 -04:00
416rehman c4e467cad4 feat(cli): calculate running stage stats dynamically from live manifest for real-time status accuracy 2026-04-16 18:51:11 -04:00
416rehman 05f621ed00 fix(cli): uniformly load environment variables in dry-run commands 2026-04-16 18:48:19 -04:00
416rehman 66269a8b0f fix(cli): gracefully handle validation and authorization exceptions without dumping tracebacks 2026-04-16 18:46:36 -04:00
416rehman 4640b6c055 docs: streamline quickstart and eliminate fragmented installation section 2026-04-16 18:42:29 -04:00
416rehman 7664d5c5ec fix(stage): import threading for type hints 2026-04-16 18:25:09 -04:00
416rehman 5b0e0ea5ba style: apply ruff auto-fixes and formatting 2026-04-16 18:24:53 -04:00
416rehman 7d67f93dcb refactor: complete processor architecture and pipeline engine overhaul
Includes major rework to Ghidra decompile block generation, PE ingest validation, core stage engine typing, resilient runtime state persistence, unit tests coverage, and completely rebuilt TUI engine telemetry.
2026-04-16 18:18:32 -04:00
416rehman 4485608e76 feat(ui): complete overhaul of CLI dashboard and pipeline visualizations
This introduces a beautiful vertical timeline layout for the pipeline stages, stripping out the fragile horizontal flow sequence that wrapped poorly on narrow terminals. All components received a modern, clean visual polish (via rich boxes and explicit hex color theming) and the documentation assets have been entirely rebuilt using high-DPI vector SVGs.
2026-04-16 18:17:20 -04:00
416rehman 3de5cba747 fix: explicit internal dev bindings mapping pytest-asyncio directly against Github Actions environment containers 2026-04-16 00:47:12 -04:00
416rehman 151522883b chore: remove hanging local ghidra output from version control 2026-04-16 00:45:14 -04:00
416rehman 82a451c6e0 chore: include explicitly formatted MIT License aligning with downstream repository config defaults 2026-04-16 00:44:37 -04:00
416rehman 438288963c docs: strip explicitly opinionated terminology and genericize README scope correctly 2026-04-16 00:42:39 -04:00
416rehman 49a749f229 docs: Completely overwrite README to precisely standardize project architecture and workflow execution schema 2026-04-16 00:41:03 -04:00
416rehman 4cdb38fee7 Enforce Ruff formatting constraints rigidly inside GitHub Actions CI limits 2026-04-16 00:37:42 -04:00
416rehman 8b8c9be438 Refactoring processor validate lifecycle configurations 2026-04-16 00:35:35 -04:00
416rehman b8cd2055cc fix(ghidra): sync internal timeout to prioritize global StageSpec timeouts 2026-04-15 23:50:02 -04:00
416rehman 294886299a docs: add .env.example template for quick setup 2026-04-15 23:39:45 -04:00
416rehman 6afd6969dc docs: overhaul README with updated features and architecture 2026-04-15 23:37:14 -04:00
416rehman 8764640c9f refactor: resolve pipeline coupling and technical debt 2026-04-15 23:26:20 -04:00
416rehman 841363866b refactor: drop #nosec tag and switch ghidra decompile processor to native asyncio execution 2026-04-15 22:31:57 -04:00
416rehman aeb2a088bb desloppify: ruff autofixes for unused imports 2026-04-15 22:27:08 -04:00
416rehman e20a9e0778 desloppify: fix unused imports and security issues 2026-04-15 22:26:15 -04:00
416rehman 37483c1998 Merge main into pipeline-architecture, resolving conflict by deleting legacy ghidra_runner.py 2026-04-15 20:48:49 -04:00
416rehman 5ab9892ded chore: finalize pipeline architecture and stabilize CI 2026-04-15 20:46:49 -04:00
416rehman a5b104900d chore: resolve static analysis vulnerabilities strictly without suppressions 2026-04-15 20:39:15 -04:00
416rehman 98a17fa116 Refactor engine APIs: introduce generic pipeline enums and processor state abstractions 2026-04-15 20:33:01 -04:00
416rehman fae8541f68 fix: globally silence tool info logs during progress loop 2026-04-15 17:00:33 -04:00
416rehman 65397ddff9 feat: max concurrency pooling and rich progress bars 2026-04-15 16:36:27 -04:00
416rehman 370bd44244 ci: add github actions workflow for testing and linting 2026-04-15 15:10:31 -04:00
416rehman 91015ea697 chore: add root dir to pytest pythonpath to resolve tools imports in test suite 2026-04-15 15:09:11 -04:00
416rehman dae7ceb6df desloppify: resolve security vulnerabilities per strict structural rules without suppressions 2026-04-15 15:02:31 -04:00
416rehman 28eaad4385 desloppify: Refactored runner broad exceptions 2026-04-15 14:43:06 -04:00
416rehman 08d743efe3 cleanup: remove unused imports in runner.py after process.py extraction 2026-04-13 21:41:01 -04:00
416rehman 9776aae88c refactor: extract subprocess utils to engine/process.py, rename _run_inner, add --model to resume
- extract run_subprocess_with_kill and kill_process_tree to engine/process.py (design coherence)
- rename _run_inner to _execute_pipeline_stages (naming quality)
- add --model/-m option to resume command for API coherence with run
- add network exposure warning when serve --host is non-localhost
2026-04-13 21:22:36 -04:00
416rehman 71e5e0fcb7 cleanup: remove unused imports in test_ghidra_decompile.py and test_stages_llm.py 2026-04-13 21:17:34 -04:00
416rehman 988bab2919 test coverage: add tests for stages/llm.py and ghidra_decompile tool, fix pe_ingest B110
- test_stages_llm: 13 tests covering process flow, caching, classification, template vars, artifact loading
- test_ghidra_decompile: 7 tests covering error paths, successful/failed decompilation, cache skip
- fix imphash extraction B110: add debug logging and noqa annotation
- total: 129 tests (up from 74)
2026-04-13 20:57:13 -04:00
416rehman d015fcd4f8 cleanup: fix unused imports in test files, add _load_env dotenv guard 2026-04-13 19:53:36 -04:00
416rehman 5f37f2554f test coverage: add 35 tests for providers (llm, decompiler) and tools (loldrivers_filter, pe_ingest)
- test_providers_llm: LLMProvider properties, completion, retry, rate limiting, token errors, import guard
- test_providers_decompiler: analyzeHeadless finder, cache hit/corrupt, ghidra execution
- test_loldrivers_filter: DB loading, process logic, skip/pass verdicts
- test_pe_ingest: directory discovery, extensions, recursion, subdirs, metadata
- also: add _load_env helper for dotenv ImportError guard in cli.py
2026-04-13 16:44:52 -04:00
416rehman 9bdaca85fd type safety: add LLMProtocol and GlobalConfig TypedDict to replace Any types in stage/runner/cli 2026-04-13 15:18:24 -04:00
416rehman 20d42d2fcf desloppify: fix security issues, remove unused imports, extract CLI helper, add exit codes
- fix B110: replace bare except:pass with specific exceptions + debug logging
- fix B104: default API host 0.0.0.0 -> 127.0.0.1
- fix B324: add usedforsecurity=False to MD5 hash
- fix B701: add jinja2 autoescape
- extract _build_runner helper to deduplicate CLI setup
- add SystemExit(1) to all CLI error paths
- remove 19 unused imports across 10 files
- strict score: 18.7 -> 70.9
2026-04-13 14:15:58 -04:00
416rehman 1f7e5b827f breadth-first core engine 2026-04-08 01:01:12 -04:00
416rehman 3f058a106e chore: remove .agent/workflows from repo 2026-04-06 22:16:22 -04:00
416rehman 77d315a894 Initial commit of DeepZero pipeline 2026-04-06 21:37:28 -04:00