The Problem with Raw Scanner Output
Run Semgrep and Bandit on the same Python repo and you’ll get the same SQLi finding at app.py:42 twice โ different rule name, same vulnerability. Add OWASP ZAP for dynamic analysis and the noise multiplies further. Before you can triage what matters, you’re deduplicating spreadsheets.
P1 solves this.
Architecture
[Semgrep JSON] [Bandit JSON] [ZAP XML]
โ โ โ
parsers/base.py (unified Finding schema)
โ
core/dedup.py โ key: {cwe}:{file}:{line}
โ
core/scorer.py โ risk_score 0-10
โ
core/llm.py โ hermes3:70b false-positive filter
โ
output/sarif_report.py โ SARIF 2.1.0
CWE-Based Deduplication
Rule names are inconsistent across tools. python.lang.security.injection.tainted-sql-string (Semgrep) and bandit.B608 (Bandit) both describe CWE-89 at the same location. Use the CWE as the canonical identifier:
def _dedup_key(f: Finding) -> str:
if f.cwe:
return f"{f.cwe}:{f.file}:{f.line}"
return f"{f.rule_id}:{f.file}:{f.line}"
On a test Django app: 400 raw findings โ 180 unique issues.
Risk Scoring
Each finding gets a risk_score 0โ10:
_SEVERITY_RISK = {"CRITICAL": 9.0, "HIGH": 7.0, "MEDIUM": 5.0, "LOW": 2.5, "INFO": 1.0}
_CWE_BUMP = {"CWE-89": 1.5, "CWE-79": 1.0, "CWE-22": 1.0}
_DYNAMIC_BUMP = 1.0 # ZAP findings confirmed exploitable
Local LLM Filter
The model sees: rule ID, CWE, severity, file path, and 10 lines of code context. Returns:
{"verdict": "false_positive", "reason": "input is sanitized at line 38"}
All on-prem via a local Ollama endpoint (http://localhost:11434 by default, override with OLLAMA_HOST). No source code leaves the machine.
SARIF Export
python main.py --target http://localhost:8080 --source ./src --format sarif
# โ findings.sarif (upload to GitHub Security tab)
likely_fp findings get a suppressions block in the SARIF โ GitHub Code Scanning treats them as dismissed, keeping the dashboard clean.