deepzero report
A run left its output spread across work/<pipeline>/samples/<id>/ as per-sample state, findings, decompiled sources and assessments. Reviewing thousands of those by hand is not practical, so `deepzero run` now prints a link to a report and keeps it up to date as results land, and `deepzero report` rebuilds it at any time. The report answers one question first: what is vulnerable. Items an assessment stage marked vulnerable lead the page, then items with findings but no confirmed verdict, then anything that errored. Each item links to its own page carrying the assessment, every finding with the code it matched, and links to the artifacts on disk. It is built from what a pipeline actually recorded rather than from any one domain's field names, so a source-code review over repositories renders as well as a kernel-driver review. A pipeline can shape the presentation with an optional `report:` block - what to call one item, which stage data key holds the verdict, which values mean vulnerable, and which columns to surface - and every field has a default. Output is layered so it stays usable on a large corpus: index.html holds the triage summary at a bounded size, items/<id>.html covers everything worth reading, inventory.csv carries every item for a spreadsheet, and findings.jsonl carries every finding one per line. When a listing is capped the page says what was capped and where the rest is. Pages are self-contained with no network access, readable in light and dark, keyboard navigable, and escape all analysed content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Automated vulnerability research pipeline engine
Define pipelines as YAML. DeepZero handles orchestration, parallelism, fault tolerance, and state.
- 🔗 Pipeline-as-YAML - chain ingest, filter, transform, and LLM-assess stages declaratively
- ⚡ Parallel execution - ThreadPoolExecutor with configurable concurrency per stage
- 💾 Resumable runs - atomic per-sample state on disk; Ctrl+C and re-run to pick up where you left off
- 🤖 LLM integration - Jinja2 prompt templates with any LLM provider via LiteLLM
- 🌐 REST API (WIP) - query run state and sample data over HTTP (currently experimental and incomplete)
- 🧩 Extensible - write custom processors as Python classes, reference them by path in YAML
📚 Documentation
DeepZero features extensive, exhaustive documentation covering architecture, pipeline schemas, CLI references, and custom processor development.
👉 Read the Official Documentation here
⚡️ Quickstart
DeepZero requires a target corpus of files to analyze and a pipeline configuration detailing how to process them.
-
Clone & Install (Python 3.11+)
git clone https://github.com/416rehman/DeepZero.git cd DeepZero pip install -e . -
Configure Environment
cp .env.example .env -
Run a Pipeline
deepzero run C:\drivers -p .\pipelines\loldrivers\pipeline.yaml
For detailed setup instructions and example corpora, see the Quickstart Documentation.
📁 Repository Structure
src/deepzero/
├── api/ # REST API (starlette)
├── engine/ # orchestration, state persistence, pipeline execution
└── stages/ # built-in processors (map, reduce, ingest)
processors/ # external processors (shipped as examples)
├── ghidra_decompile/ # ghidra headless decompiler (MapProcessor)
├── loldrivers_filter/ # loldrivers.io hash exclusion filter (MapProcessor)
├── pe_ingest/ # PE header parser and driver metadata extractor (IngestProcessor)
└── semgrep_scanner/ # semgrep batch scanner (BulkMapProcessor)
pipelines/
└── loldrivers/ # BYOVD kernel driver vulnerability research pipeline
├── pipeline.yaml
├── assessment.j2 # LLM prompt template
└── rules/ # semgrep rules
docs/ # Jekyll-based GitHub Pages documentation
tests/ # pytest suite
🤝 Contributing
CI runs on Python 3.11 and 3.12 via GitHub Actions.
Run linting and security checks before submitting:
ruff check . && ruff format --check . && bandit -ll -ii -c pyproject.toml -r .
Please refer to the Contributing Guide and the Code of Conduct before submitting pull requests.
📄 License
DeepZero is released under the MIT License.