Many drivers create their device object only once their hardware is
enumerated. On a machine without it the device never appears, so nothing
the driver exposes can be reached - and in the report that is
indistinguishable from the driver not being vulnerable. Confirming a
flagged driver meant loading it and finding out by hand.
Decompilation now records which function calls IoCreateDevice and
whether that function is DriverEntry or something it calls. A driver
created on the load path can be exercised on any machine that will load
it; one created elsewhere, typically a PnP add-device or start-device
callback, needs the physical device. Neither is recorded when the call is
not found, rather than guessing.
Resolved through the symbol table instead of decompiling every function
and matching text, so a call ghidra renders differently in C is still
found, and the pass costs nothing.
The report needs no change to show it: the processor declares the value
in `provides` and the bundled pipeline names it under `report.columns`,
which is the existing route for anything a pipeline wants surfaced. The
report stays unaware of what a device object is.
The decision is a plain function so it can be checked without ghidra, and
the whole path was run against a real driver: null.sys resolves to
on-load-path, which matches \.\NUL opening on any machine.
Closes#20
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A stage that hit its budget was recorded as failed. Because the failure
is stored against the sample, a resumed run skipped it, so those samples
were excluded from every later stage for good. They are the largest ones
in a corpus, which are often the ones worth looking at, so the run
quietly biased itself away from its best targets and the report filed
them with genuine errors.
Running out of time is now its own outcome, carrying the budget it
exceeded. TimeoutError is an OSError, so the handler for it has to come
before the general one or it is caught as something breaking; a test
pins that ordering, and another drives a slow stage through the runner
to check what actually gets recorded.
`--retry-timeouts` attempts those samples again with a longer budget,
`--timeout-multiplier` deciding how much longer. Only the stages that ran
out of time are forgotten, so everything else the run produced is kept
and the retry costs the time those stages need and nothing else.
The report gives them their own group with the budget each exceeded, and
counts them apart from both errors and clear results, so a corpus that
was never finished is not reported as one that came back clean.
Closes#19
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The bar held seven identical chips in one run. Three of them narrow by
what the run concluded and four by what a person recorded, and the two
combine rather than being alternatives, so presenting them as one set
misread both what they do and how they work together.
They are now separate groups with a rule between them and the second one
named. The review chips take a different shape and carry a status dot
rather than a different colour, because colour already means outcome
here and a review is not an outcome. Stacked on a narrow screen the rule
between them turns with them.
The bar measures its own height for the table head to park under, and it
was only remeasured when the window resized. It also changes height when
its controls wrap, which a live run causes by itself as new outcomes
appear and add chips, leaving the head parked against a stale figure.
It now watches the bar itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A review had two outcomes: confirmed, or something outstanding. That left
nowhere to put the most expensive result there is - taking a finding to a
test machine, running it, and watching it not happen. Recorded as
outstanding it reads as unfinished work; recorded as nothing at all it is
indistinguishable from a result nobody has opened.
Reviews now have a third outcome for a result that did not hold up. It
opens the same note the outstanding state does, because what was run and
what happened instead is the part worth keeping, and the summary counts
each outcome so a run can be read as what is known rather than what was
alleged.
Closes#25
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A pipeline records what it concluded. Whether a person has verified that
conclusion is a separate claim, and nothing in a run can make it, so the
report now carries marks the reader adds: confirmed, or outstanding with
a note naming what is still unproven.
Deliberately generic. What counts as outstanding is the reader's
business, and reproducing a crash, proving a precondition or reading a
diff are the same shape of unfinished work as far as the report is
concerned, so it only provides somewhere to put it.
Marks show against each row and can be filtered on, including the results
nobody has reviewed yet. The report is a file, so they are kept in the
browser and keyed on the pipeline and target rather than the path, which
means a rerun over the same corpus keeps them. Exporting writes marks.json;
saved next to the report it is read back at generation time, so the marks
render for anyone who opens it and reach inventory.csv alongside
everything else instead of being stranded in one person's browser.
Closes#24
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Decompilation output grows silently: it stays correct while getting
enormous, so the regression is only visible to someone who happens to
look at the directory size. The reference pipeline documentation now says
what the stage writes per driver and what it should total, with the
figures from the corpus where storing the routine per IOCTL code turned
28 MB into 734 MB.
The entry a code decodes to is now built by its own function, so a test
can hold it to carrying the code and its decoded fields and nothing that
could grow. One case covers a field large enough to be a routine appearing
under any name, since that is the shape of the regression rather than any
particular field.
Closes#22
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The scanner ran every sample through a single semgrep invocation, so any
failure in that run cost every result. On a real driver pack that meant
1,380 drivers and six minutes of work producing nothing, with the scan
stage gating assessment so the run quietly ended there.
Samples are now scanned in batches, configurable per stage and defaulting
to 200. A batch that times out or comes back unreadable marks only its own
samples failed; the rest still deliver their findings.
A scan that read none of the files it was given is also no longer recorded
as a clean result. semgrep lists the files it opened, and reading none of
them means the scan never looked, which is a different outcome from
looking and finding nothing - the second is safe to pass to assessment and
the first is not. Each sample now also records how many of its files were
submitted and how many were read, so a sample reporting nothing because
its files went unread can be told apart in the report.
Closes#21
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The test covering the shipped pipeline's prompt ran validate_pipeline,
which also insists on a Ghidra install and a signed-in model backend.
Neither says anything about whether the prompt is correct, so the test
passed only on a machine already set up to run the pipeline and failed
everywhere else.
It now resolves the pipeline's stages and checks the prompt against what
they declare, which is the property being tested and needs no
environment. Verified against a machine with neither dependency present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A prompt is rendered from whatever the stages before it recorded. A name
none of them produce renders as nothing, so the model is asked to judge
an empty payload and answers about the emptiness -- and that answer is
stored as the verdict. The run reports success and nothing indicates the
code was never seen.
Processors now declare the names they make available in `provides`: the
keys they record, plus one per artifact they write. `deepzero validate`
walks the stages in order and checks each prompt against what its
predecessors declare, reporting an unknown name along with the names that
were available instead. A value is only offered to stages after the one
recording it, so a prompt cannot reach its own stage's output.
This also closes a gap between validation and rendering. A prompt named
as a bare filename next to the pipeline passed validation, because the
file exists, but rendering only resolved references containing a path
separator -- so at run time the filename itself was sent as the whole
prompt. Both now resolve a reference the same way, and one shared rule
turns an artifact path into the name a prompt uses, so what is checked is
what gets rendered.
Documents the values a prompt receives, and how a processor declares the
ones it adds.
Closes#23
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every job in the matrix installed a JDK and downloaded a
several-hundred-megabyte Ghidra archive, then ran the linter and the
security scan again. Neither of those reads the interpreter, and only one
test reads Ghidra, so both costs were paid once per python version for
no additional signal.
The work is now split three ways:
- lint runs once, and installs the tools and the package's own light
dependencies rather than everything under full
- the version matrix runs the suite without Ghidra, which is every test
but one, and so needs no JVM at all
- a single integration job drives the real Ghidra install
Ghidra is cached between runs, keyed on the version and build it pins,
and unpacked under HOME so restoring it needs no privileges. Its
checksum is still verified whenever it is downloaded. The test that needs
it now carries a marker, so selecting it is declarative instead of a path
spelled out in the workflow.
A newer push cancels an in-flight run for the same branch, and the
linters are pinned to a minor so their own releases cannot fail a build
on their own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The project declares support from 3.11 upward, so 3.13 and 3.14 were
already installable but never exercised. Both now run in the matrix,
alongside the versions that were there before.
A failing version no longer cancels the others. Across four interpreters,
knowing which ones broke is most of the answer, and stopping at the first
failure threw that away.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Needs review" did not say what kind of review was missing. These are
results a scan flagged that no assessment stage has confirmed or
dismissed, so they are now "Needs assessment", matching the vocabulary
the rest of the report already uses.
That left two labels a reader could not tell apart, because the outcome
next to it was called "Not assessed" while describing results an
assessment had in fact produced a verdict for. What is missing there is
a verdict the pipeline recognises, so it is now "Unclear verdict".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The index stacked a separate table per outcome, so sorting one of them
said nothing about the others even though risk ranks across the whole
corpus. Results now sit in a single list ordered by risk, with the
outcome as a filter alongside the name search, so a sort applies to
everything currently on screen.
Severity counts critical separately from high. The two differ by ten
times in the ranking, so folding them together left the column unable to
explain the order it was sorted in.
Colour now marks outcomes and nothing else: severity tiers, the outcome
of each result, and a corpus bar that dims to whatever the filter is
showing. Light and dark are built as one system, and both meet WCAG AA
on every surface a colour appears on, including the raised background
under a hovered row.
A live run reloads the page from script rather than a meta refresh,
keeping the reader's filter and scroll position instead of returning
them to the top every twenty seconds.
Fixes:
- the table head never stayed in view, because the scroll wrapper around
the table was acting as the scroll container
- sort arrows rendered as an unrelated glyph
- results with no recorded verdict printed a raw HTML entity
- samples an assessment could not classify were counted as clear while
also being listed separately as unassessed
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds the release-please GitHub Action (python release type), config, and
manifest pinned at the current 0.2.0. The version lives in both
pyproject.toml and src/deepzero/__init__.py; the latter is bumped via
extra-files plus an x-release-please-version annotation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Runs are namespaced <work>/<pipeline>/<corpus-key>/ so different targets
through the same pipeline no longer share a sample store or report. The
corpus key is <basename>-<8 hex of sha256(resolved path)>, so a rerun of
the same target resumes in place while distinct targets stay isolated.
status and report resolve to the newest corpus run; --clean purges only
that corpus's run directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A run recorded its status when it started and never revisited it. A pipeline
that is killed, or whose machine restarts, never gets to write a closing
status, so "running" stayed on record indefinitely. The report repeated it,
kept its auto-refresh going, and presented a run that had ended hours earlier
as though results were still arriving.
A run now records which process is writing it, and the report checks that
process before repeating the claim. A run that recorded an outcome is taken at
its word; only one still claiming to be in progress is checked. Where the
answer cannot be established - another machine, a process that cannot be
queried - the report says so rather than guessing.
The page itself also ages. A live run rewrites it every few seconds, so once
that stops the gap between the page and the reader's clock gives it away, and
the status changes to say the page is no longer being updated.
The liveness probe deliberately avoids signalling the process: on Windows the
usual existence check terminates the target instead of testing it, which would
have made a report refresh kill the run it was describing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Results were only written out once fifty samples had finished, so a slow
stage left them invisible: a run that had already found and recorded a
vulnerable driver still reported none, because the finding had not been
written to the run's state yet. An interruption also discarded up to fifty
finished samples. State is now written on a short timer as well, so the
report reflects what has actually been found.
Each outcome label now explains on hover what it means and how a sample
came to be in it, including that a filtered item was excluded before
analysis finished rather than judged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The summary counted every sample that was not vulnerable as clean, which
included files a stage had excluded before any analysis ran. On a real
driver pack that put 950 files that are not kernel drivers, and were never
looked at, in the same count as drivers that were analysed and came back
clear.
Those are now separate: 'Clear' means analysed all the way through with
nothing flagged, and 'Filtered out' means a stage excluded it, naming the
stage that did so.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An assessment used to end in prose, so acting on it meant reading the
paragraphs and reconstructing the request by hand. It now states plainly
that a separate step will try to reproduce whatever it reports on a real
machine, and a vulnerable verdict ends with the values that step needs:
the device to open, the IOCTL code, the input buffer laid out field by
field with values that exercise the flaw, the expected output size, and
the observable result that tells a real hit apart from the driver merely
accepting the request.
The assessment is also given the IOCTL codes recovered from the dispatch
routine, which it is asked to name but previously had to guess at.
A reduce stage ranks the samples it keeps, and later stages now work
through them in that order. Assessment of a large corpus takes hours, so
stopping early now leaves the highest-ranked samples already done rather
than an arbitrary subset.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The assessment prompt referred to values the pipeline does not produce.
Jinja renders an unknown name as empty text, so the prompt ended with
"Payload:" and nothing after it: every driver was assessed with no
decompiled code, no handler name and no findings list. The model said so
in its own answers - that there was nothing to analyse and any finding it
reported would be invented - and those verdicts were recorded as results.
The prompt now uses the names the pipeline actually records, so a driver
is assessed with its dispatch routine and its findings. On a driver from
a real pack this is the difference between a 13 character prompt and a
20,800 character one.
A prompt naming something the pipeline does not produce is now an error
that says which name was wrong and lists what is available, instead of
quietly asking the model to judge nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Decompilation wrote one file per IOCTL code and copied the whole dispatch
routine into every one of them. Every code in a driver is handled by that
same routine, which is already written once to dispatch_ioctl.c.
On a real driver pack this produced 682 MB across 10,305 files where the
unique content was 18 MB, and the scanner then read the same function up
to 252 times for a single driver - counting every finding again for each
code, so a driver with many codes looked far worse than one with few and
ranked above it.
Each per-code file now records what the code decodes to and points at the
routine that handles it. Scan input for the same pack drops from 703 MB to
roughly 18 MB, and a finding is counted once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Scanning a real driver pack failed on all 1380 drivers at once. The
scanner read the results document back through a pipe, and at that size
the read fails, so a scan that had already run for six minutes produced
nothing and the whole assessment stage was skipped.
The scanner now has semgrep write its results to a file and reads that,
which does not depend on how large the document is. A launch failure also
no longer claims semgrep is missing when it is installed and something
else went wrong - the real error is reported.
Decompilation stored the whole dispatch routine again for every IOCTL code
it found, which produced 52 MB of output for a single driver and 734 MB
across the pack, nearly all of it the same text repeated. The routine is
recorded once, as it already was alongside it.
Assessment also reads artifacts into the prompt. A file far larger than
the context budget is now skipped with a warning instead of being parsed
and loaded on every sample, matching how oversized source files were
already handled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The bundled kernel-driver pipeline kept only the ten highest-scoring
drivers before assessment and discarded the rest, so most of what
scanning flagged was never looked at. It now orders every surviving
driver by how many findings it has and assesses all of them, worst
first, and the stage is named for what it does.
A ranking stage asked to keep `0` items used to drop every sample and
report success, which made "no limit" the one value that silently
analysed nothing. Zero or fewer now keeps everything, still in metric
order, and the stage says how many it kept and dropped either way.
Assessment is pinned to a full model id rather than a short alias, so
the pipeline cannot quietly switch generations when an alias moves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run left its output spread across work/<pipeline>/samples/<id>/ as
per-sample state, findings, decompiled sources and assessments. Reviewing
thousands of those by hand is not practical, so `deepzero run` now prints
a link to a report and keeps it up to date as results land, and
`deepzero report` rebuilds it at any time.
The report answers one question first: what is vulnerable. Items an
assessment stage marked vulnerable lead the page, then items with
findings but no confirmed verdict, then anything that errored. Each item
links to its own page carrying the assessment, every finding with the
code it matched, and links to the artifacts on disk.
It is built from what a pipeline actually recorded rather than from any
one domain's field names, so a source-code review over repositories
renders as well as a kernel-driver review. A pipeline can shape the
presentation with an optional `report:` block - what to call one item,
which stage data key holds the verdict, which values mean vulnerable,
and which columns to surface - and every field has a default.
Output is layered so it stays usable on a large corpus: index.html holds
the triage summary at a bounded size, items/<id>.html covers everything
worth reading, inventory.csv carries every item for a spreadsheet, and
findings.jsonl carries every finding one per line. When a listing is
capped the page says what was capped and where the rest is.
Pages are self-contained with no network access, readable in light and
dark, keyboard navigable, and escape all analysed content.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A stage can report that a sample needs no work because its output is
already on disk. That sample was then treated as filtered out and never
reached any later stage, so resuming an interrupted run quietly analysed
fewer samples than a fresh run would - drivers dropped out before
scanning and assessment with nothing to indicate it.
Work that is already done now counts as passed and the sample continues,
which is what MapProcessor.should_skip documents. A stage that did no
work also records why in its own field, so a routine skip is no longer
reported as an error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vulnerability assessment of decompiled kernel drivers is the most
demanding step in the pipeline, so it now defaults to claude-code/opus.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An auth failure surfaced the CLI's raw message ("401 OAuth access token
has been revoked") with no indication of what to do about it. DeepZero now
relays that message with the remedy attached: run `claude` in a terminal
and sign in.
DeepZero deliberately does not inspect Claude Code's credential store to
predict this. Whatever auth the CLI has is the auth DeepZero uses; if it
has none, the CLI reports it and we pass that along. Auth environment
variables are left untouched so the CLI authenticates exactly as it
normally would.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A pipeline using Claude Code only calls the LLM in its final stage, so a
signed-out CLI wasn't discovered until after decompilation and scanning
had already run - potentially hours of wasted work. Passing `--preflight`
to `deepzero run` now sends one tiny probe to the configured backend
first and stops immediately with a clear "not authenticated" message if
it fails. It is opt-in and off by default, so normal runs spend nothing
extra, and a transient hiccup like a rate limit does not block the run.
Closes#14.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
With `parallel: 0`, the decompile stage auto-scaled to one worker per
CPU core, and each worker is a separate Ghidra JVM holding a full
program database. On a many-core machine that meant dozens of JVMs and,
on a real driver corpus, exhausted RAM.
`parallel: 0` now derives a memory-safe worker count (roughly one worker
per 4 GiB of RAM, bounded by cores and a hard cap) instead of using
every core. An explicit `parallel: N` is still honored as-is, and the
ceiling can be raised on a large machine with `max_parallel:` in the
decompile stage config. Processors can now declare a concurrency ceiling
that only applies to auto-scaling.
Closes#12.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Addresses PR review feedback:
- The scan no longer falls back to an unintended config path when its
rules directory can't be resolved; it fails fast with an explicit
"rules_dir not found" instead of silently scanning with no rules.
- Per-call options passed to the LLM are now forwarded to the backend
rather than dropped, so controls like `timeout=` take effect for the
Claude Code CLI backend instead of being silently ignored. An unused
internal flag that gated this is removed.
- Made a rules-resolution test independent of the working directory so
it can't flake on a stray local `rules/` directory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The pipeline defaulted to vertex_ai/gemini-2.5-pro, which needs Vertex AI
credentials most users do not have, so a fresh checkout failed validation
before doing any work. It now defaults to claude-code/sonnet and runs
against a locally signed-in Claude Code with no cloud API keys. Override
per run with `-m` as before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI installs the latest ruff, and newer versions format python snippets
embedded in markdown. That made `ruff format --check .` fail on
docs/en/extensibility/building-custom.md purely from a ruff upgrade,
with no source change. Documentation examples are prose, not build
artifacts, so they are now out of the formatter's scope.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When run from the repo root, the scan pointed at the wrong rules
directory: the check that approves a run used one path while the run
itself used another. semgrep then loaded no rules and reported every
driver as clean, so real findings were silently missed. Both now resolve
the rules the same way, and a scan that loads no rules fails loudly
instead of looking like "no vulnerabilities found".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The scanner treated any unexpected exit code from semgrep as fatal and
reported only its (often empty) error stream, producing an opaque
failure with no cause. It now trusts a complete scan result whatever the
exit code - semgrep can finish the scan yet exit oddly on Windows - and,
when a run genuinely fails, reports the exit code and error text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
DeepZero prints status glyphs such as checkmarks and arrows. On a legacy
Windows console, or whenever output was redirected to a file, printing
them raised an encoding error that aborted the run while writing the
final summary - after all the analysis work had already completed.
Output is now written as utf-8 with unrepresentable characters escaped
rather than fatal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Decompilation crashed for every driver that actually had IOCTL codes -
precisely the drivers worth investigating - while drivers with none
completed, which hid the problem. Writing the per-IOCTL output failed
inside Ghidra's Jython interpreter, which rejects byte strings on a
utf-8 text stream.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An unset ${VAR} with no default previously survived as the literal text
"${VAR}", which then read as a real value and produced confusing errors
such as "ghidra not found at ${GHIDRA_INSTALL_DIR}" instead of saying
the setting was required. Unset variables now expand to empty, matching
${VAR:-default} and ordinary shell behaviour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pipelines can now use the Claude Code you already have installed and
signed in, so LLM analysis works without a third-party API key or
per-token billing:
model: claude-code # default model
model: claude-code/sonnet # or /opus, or a full model name
DeepZero never handles your credentials - it runs your own `claude`
binary in headless mode and inherits its existing sign-in. If the CLI is
missing or not signed in, validation says so before a run starts.
The pipeline feeds untrusted decompiled code to the model, so the
integration denies tools and external servers, sends the prompt over
stdin rather than the command line, and hides any ANTHROPIC_API_KEY from
the CLI so your subscription is used and not a metered API. Usage limits
retry with backoff; sign-in failures stop immediately with a clear
message.
Under the hood, LLM backends are selected from the model string by a
registry, so another agent CLI can be added later without touching the
engine, pipelines, or stages. Existing API-key models are unaffected.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>