LLMRed is a modular AI red-teaming and security assessment platform. It combines deterministic security probes, evidence collection, framework mapping, governance assessment, and reporting under explicit engagement safety controls.
Engagement configuration
|
v
Capability profiler ---> STRIDE + AI-domain threat model
|
v
Target client + audit wrapper
|
v
Technique registry ---> execution recorder ---> coverage records
| |
v v
Recon / Access / Execution / Impact modules
|
v
Attack + benign control ---> evaluator ---> repeated-trial statistics
|
v
Optional surface adapters ---> channel / memory / sink / MCP / multimodal / RAG tests
|
v
Quarantined artifacts ---> supply-chain trust gate ---> approve / reject / review
|
v
Sanitized findings + NIST outcome evidence/profile + project-score caps
|
v
HTML / PDF / JSON / comparison reports
|
v
Remediation plan ---> conclusive retest ---> verified / failed / unresolved
|
v
Version-bound regression gate
The technique registry contains stable identifiers, phase, risk level, prerequisites, and whether a test writes to the target. Phase modules own execution order. The execution recorder isolates failures so one broken technique does not abort the assessment and records one of:
skippedexecuted_no_findingfindinginconclusiveerror
A comparison is considered divergent only when both targets have comparable finding or executed_no_finding records. Missing, skipped, inconclusive, and errored tests cannot be described as blocked controls.
The next architecture stages add:
- An engagement-policy boundary in front of execution for scope, consent, budgets, and dangerous-test approval. Implemented: static host/CIDR/port validation, time/runtime windows, thread-safe request/token reservations, separate write/DoS consent, run-bound policy hashing, cleanup verification, and recovery ledger.
- An evidence service for redaction, hashing, encryption, and reproducibility bundles.
- An evaluator layer separating attack generation from deterministic/LLM/human judgment. Implemented foundation: typed attack/control trials, explicit success/failure/review/inconclusive outcomes, evaluator provenance, repeated-trial ASR, Wilson intervals, control-positive and adjusted rates, and human-review states. The jailbreak library is migrated; remaining legacy modules still require migration.
- Capability discovery and STRIDE-AI threat modeling that select applicable techniques. Implemented: tri-state evidence with provenance, explicit declarations, capability prerequisites, distinct
not_applicablecoverage, classic STRIDE categories plus separate AI/ML domains, and HTML/JSON/forensic outputs. - Additional RAG, memory, multimodal, agent, tool, MCP, and supply-chain modules. Implemented foundation: typed target adapters and executable tests cover external channels, persistent memory, sandboxed tool sinks, MCP catalogs, multimodal injection, and RAG provenance. The separate supply-chain gate performs non-loading artifact inspection, provenance/integrity checks, dependency/SBOM review, provider assessment, and optional static scanners. Concrete target adapters and registry promotion remain engagement-specific.
- Implemented foundation: a constrained planner chooses only zero-argument tools from a sealed registry. The trusted runner independently enforces explicit technique/risk allowlists, phase and capability prerequisites, write/DoS consent, engagement budgets, repetitions, step/error limits, cleanup-failure stops, and Critical-finding stops. Planner context contains sanitized outcome summaries rather than target evidence. Distributed workers and an external planning-service adapter remain future extensions.
Report risk and assessment assurance are separate dimensions. Finding severity cannot hide missing coverage, and missing coverage cannot be interpreted as a passed control. Versioned project crosswalks relate findings to OWASP LLM 2025, MITRE ATLAS 2026.06, the complete NIST AI RMF 1.0 Core with an informative NIST AI 600-1 profile reference, Google SAIF, and CWE while preserving the distinction between taxonomy relevance and compliance evidence.
Remediation records are bound to the SHA-256 of their source JSON report.
Regression comparison operates on stable technique IDs and typed execution
states. It may verify a remediation only from finding to a conclusive
executed_no_finding; skipped, errored, missing, and inconclusive retests remain
unresolved. This favors defensible evidence over optimistic status automation.
The complete official ATLAS 2026.06 YAML is a pinned, checksum-verified data
source. atlas_catalog.py provides matrix-wide queries and relationship
resolution. atlas_mapping.py contains only platform-specific mapping judgments,
including multiple references and exactness limitations. Reports never invent
tactic names or technique labels; those fields come from the pinned official
catalog.
- The current recorder wraps legacy list-returning modules. A returned empty list means the module completed without emitting a finding, but only module-level typed results can distinguish every network-level inconclusive case. Migrating each module to return a native result object is therefore still required.
- Calibration thresholds are engagement policy, not universal truth. Small trial counts have wide confidence intervals; reports expose that uncertainty rather than treating
temperature=0as proof of determinism. External LLM judges are opt-in and must identify provider, model, and judge-prompt version. - Capability inference is intentionally conservative. A missing schema or adapter leaves a capability
unknown; only an explicit unsupported declaration can make a techniquenot_applicable. Architecture-owner declarations should later be replaced or corroborated with observed API/data-flow evidence. - Specialized adapters deliberately move target-specific transport outside generic test logic. This increases per-engagement adapter work but prevents simulated inputs from being misreported as tests of a real ingestion/action boundary. Every persistent adapter must expose deletion/rollback verification; failures are retained in the recovery ledger.
- The registry centralizes metadata but leaves test implementations modular. This duplicates a small amount of phase ordering data in exchange for avoiding imports and side effects in reporting/policy code.
- The audit hash chain is tamper-evident within a retained run, not externally anchored. Production evidentiary use requires WORM or independently timestamped storage.
- Supply-chain promotion is deliberately outside the intake process. This requires a second registry integration step, but prevents parser/scanner uncertainty from causing an automatic publish. SafeTensors validation reduces serialization risk and is never treated as behavioral or architectural assurance.
- Agentic orchestration improves adaptive ordering but adds planner complexity. Authorization therefore remains deterministic and outside the planner. The current single-process runner favors auditability and reproducibility over parallelism; a distributed version will need durable state, idempotency, isolation, and queue-backed recovery handling.
These boundaries should be revisited when the platform supports distributed workers, long-running campaigns, or multi-user engagement management; those cases will require persistent run storage and a queue rather than the current in-process recorder.