mirror of
https://github.com/SpecterOps/Nemesis
synced 2026-06-08 12:36:42 +00:00
6ec0a3b61a
* upgrade to dapr postgresv2 statestore * actually make it v2 * dapr state table name, cleanup subscriptions/globals * proper exceptions * formatting/lint * Refactor workflow tracking and improve activity input handling - Extract workflow tracking logic into dedicated WorkflowTrackingService - Simplify activity signatures to accept specific parameters instead of generic dicts - Remove unused asyncio event loop references from enrichment modules - Update YaraRuleManager initialization and method names * re-added workflow tracking in the DB * update uvicorn prod options * enrichment work parallelism, convert queues from broadcast to task queues * Refactor pubsub and improve workflow parallelism - Split Dapr pubsub Dapr yaml components into topic-specific queues (alerting, dotnet, dpapi, files, noseyparker, workflow_monitor) - Update all Dapr volume mounts to reference new topic-specific pubsub components - Converted queues to task queues - Use YAML anchors to reduce duplication for file-enrichment replicas - Pass asyncpg pool to enrichment modules instead of creating connections - Add asyncpg_pool parameter throughout chromium and enrichment module analyzers - Update VSCode workspace (removed InspectAssembly, renamed dotnet_api to dotnet_service) - Added curl commands for Jaeger API to performance docs to help with perf troubleshooting - Created common.queues module to centralize pubsub/topic names (eases future refactoring) * worker mods * Workflow performance tuning, fix pubsub config, CLI arg changes - Fix pubsub deleteWhenUnused typo (deletedWhenUnused) - Add LOG_LEVEL environment variable support across services - CLI: Rename --repeat to --times, add --max-files option - Increase files pubsub prefetchCount from 25 to 50 - Add MAX_PARALLEL_WORKFLOWS configuration - Fix DotNetAssemblyAnalysis null handling with field validators - Update dashboard to show cumulative files/findings over time - Add RUST_LOG environment variable support to noseyparker - Update CHANGELOG for 2.1.4 release notes * Dapr 1.16.2 and use db transactions - Upgrade all Dapr containers from 1.16.1 to 1.16.2 - Reduce enrichment parallelism default from 25 to 5 workflows - Reduce healthcheck intervals from 10s to 5s for alerting and document conversion - Fixed DPAPI eventing to use new pubsubs - Refactor file_linking database operations to use atomic upserts and avoid deadlocks - Add WriteOnceViolationError handling in DPAPI masterkey analyzer - Wrap database operations in transactions for enrichment storage and plaintext indexing - Fix postgres notification handler closure variable capture * remove unused start_time * Scheduler persistence, workflow concurrency tuning, and config cleanup - Add volume for Dapr scheduler and init service - Add scheduler dependency to file enrichment service - Add async workflow client libraries - Format and cleanup compose.yaml (spacing, indentation, empty lines) * Migrate file_enrichment to async Dapr client and optimize Dockerfile - Use async DaprClient where possible in file_enrichment - Improve Dockerfile caching - Add asyncpg connection pool helper and fix typo in secret store name - Include VS Code debug configuration for document_conversion - Remove unused dapr_client from DpapiBlobAnalyzer - Clean up activity return types and better handle exceptions * Enrichment tracking for NoseyParker and logging cleanup - Add workflow_id to NoseyParkerInput and NoseyParkerOutput models - Remove workflow lookup query in noseyparker subscription handler - Adjust jaeger_perf_stats.sh output formatting and precision - Add type hints for async functions * noseyparker scanner perf, tracing for update_enrichment_results --------- Co-authored-by: Lee Chagolla-Christensen <lee@localhost>
Document Conversion Service
A microservice for the Nemesis platform that handles document processing, text extraction, and file format conversion.
Purpose
This service processes uploaded files to extract textual content, convert documents to PDF format, and extract strings from binary files. It serves as a key component in making file contents searchable and viewable within the Nemesis platform.
Features
- Text extraction: Extract readable text from various document formats using Apache Tika
- PDF conversion: Convert Office documents and other formats to PDF using Gotenberg
- String extraction: Extract printable strings from binary files
- Encryption detection: Automatically detect and skip encrypted or protected files
- Parallel processing: Execute multiple extraction methods simultaneously using Dapr workflows
- Container support: Process files extracted from archives and containers
Document Processing Pipeline
The service uses a Dapr workflow to process files through three parallel activities:
- Tika text extraction: Extracts formatted text from documents
- String extraction: Pulls ASCII/Unicode strings from binary files
- PDF conversion: Converts supported formats to PDF for viewing
Supported File Types
- Text extraction: Office documents, PDFs, HTML, RTF, and other text-based formats
- PDF conversion: Microsoft Office files, LibreOffice documents, text files
- String extraction: Any binary file format
Configuration
The service integrates with external components:
- Apache Tika: Java-based text extraction engine
- Gotenberg: Document to PDF conversion service via HTTP API
- Minio: Object storage for input files and generated outputs
- PostgreSQL: Workflow state and transform metadata storage
Workflow Behavior
Files are processed only if they meet specific criteria:
- Not already plaintext
- Not encrypted or password protected
- Either original submissions or container-extracted files
- Not existing transforms to avoid duplicate processing
Health Monitoring
GET /healthz: Comprehensive health check including database connectivity, JVM status, and workflow runtime verification