Skip to main content

Overview

What is ClusterHawk?

ClusterHawk is the threat infrastructure analysis platform built for CTI teams, threat hunters, and SOC analysts. It finds malicious actor networks before they appear in reputation feeds.

Infrastructure Profiling
Threat Detection
Predictive Models
Infrastructure Mapping
Anomaly Detection
Integrations

Core Concepts

How ClusterHawk thinks about infrastructure

ClusterHawk analyzes the metadata behind IP addresses (service stacks, TLS configurations, certificates, ports) to surface the relationships and shared posture that reveal a common operator. This helps security teams detect potential threats before they become active.

Unlike traditional IP reputation services that rely solely on historical data, ClusterHawk uses a deterministic ensemble pipeline to identify and predict potentially malicious infrastructure, even when specific IPs haven't been previously flagged. No AI theater: same input, same clusters, same reasoning, every run. The decision logic is exposed through SHAP/LIME on every cluster.

Validated in the open: our published casework independently reproduced eSentire's ClickFix chain, Black Lotus Labs' SystemBC botnet, and Trellix's SideWinder campaign, and added tiers, fingerprints, and anomalies each original report didn't carry. Most importantly, the methodology reliably surfaces emerging threats with no prior reputation indicators, closing the gap between first appearance and first detection.

ClusterHawk's analysis is deep and behavior-driven, rather than a surface-level scan. That depth turns raw clustering into actionable intelligence, helping SOC, Threat Intelligence, and Hunting teams see and stop complex threats before they take hold.

How it works

How ClusterHawk Works

1. Data Ingestion & Infrastructure Profiling

ClusterHawk processes IP datasets through a weighted ensemble clustering methodology that combines multiple approaches. Unlike standard clustering, our weighted approach prioritizes the most effective algorithms for your specific dataset. The system handles datasets from 100 to 5000+ addresses and identifies patterns that are impossible to detect manually through infrastructure analysis.

2. Evaluation & Labeling

Our system employs an evaluation methodology to assess cluster quality and applies custom labeling rules to identify malicious clusters. This extends threat detection to previously unknown IP addresses sharing characteristics with known threats. Configure your own labeling rules or use pre-defined templates, all evaluated using our documented scoring system.

3. Model Training & Prediction

ClusterHawk creates custom models based on specific attack patterns and infrastructure behaviors. These models predict whether new IP addresses are likely malicious, with confidence scores for different threat types (e.g., "63% confidence for APT29 phishing, 73% for Lazarus Group").

4. Analysis

Detailed analysis views include clustering results, infrastructure relationships, neighborhood analysis, and relationship and feature analysis. Explore how IP addresses relate to each other and move between clusters over time. Clustering also reveals infrastructure deployment lifecycles (from staging and provisioning through full operational deployment to decommissioning), which helps predict where new infrastructure will appear next. Statistical analysis includes organizational profiling, geographic distribution, vulnerability analysis with CVE/EPSS integration, and malware indicators.

5. Noise Intelligence Mining

Transform traditional 'noise' clusters into actionable intelligence through adaptive re-clustering and similarity analysis. This capability uncovers emerging threats and sophisticated adversary infrastructure designed to evade conventional detection methods.

6. GPU-Accelerated Processing

GPU acceleration speeds up analysis while maintaining analytical depth. Typical jobs complete in 20 to 45 minutes; the largest datasets take a few hours.

7. Structural Anomaly Detection

Anomaly detection analyzes the topological structure of clustering results using mathematical concepts including structural tensor analysis and topological charges. It identifies anomalous patterns without training data through multi-scale analysis and explainable intelligence.

8. Bidirectional Predictive Defense

When an attacker compromises a server, they leave fingerprints: new services, modified TLS configurations, installed tooling like beacons or reverse proxies. The compromised server now carries a hybrid fingerprint: its original service stack plus the attacker's additions, which may also introduce new vulnerabilities. These victim IOCs can be clustered to build profiles of what compromised infrastructure looks like, then used to identify other servers that are already exploited or share the same vulnerable characteristics that made exploitation possible in the first place. Combined with attacker infrastructure profiling, this creates bidirectional defense: it detects the attacker's C2 network on one side and finds compromised or at-risk assets on the other.

Integrations

Integrations & Exports

ClusterHawk fits into the tools your team already uses. Export intelligence in industry-standard formats, query results via API, or push detection rules straight to your SIEM.

STIX 2.1 Export

Export analysis results as a STIX 2.1 bundle containing infrastructure objects, malware SDOs, indicators, and relationship mappings. Bundles are ready for direct import into OpenCTI and MISP, turning every analysis into structured, shareable threat intelligence.

Learn more in the docs

MISP Export

Export in MISP-native JSON format for direct import into your MISP instance. Cluster relationships, annotation objects, and IOC flags are preserved natively, so no STIX conversion is needed. MISP auto-correlates exported IPs with other events in your instance.

Learn more in the docs

Public Prediction API

Programmatically submit IP lists for prediction against your trained models using a REST API with API-key authentication. Automate bulk IP analysis and integrate ClusterHawk predictions directly into your SOC workflows, SOAR playbooks, or custom tooling. Available on Team tier and above.

Learn more in the docs

SIEM & Detection Rules

Auto-generate SIEM detection rules from cluster fingerprints, including JARM (TLS), JA3S (SSL), and HASSH (SSH) signatures. Rules are based on infrastructure behavior, not static IOCs that rotate, and can be deployed directly to any metadata signature-aware SIEM.

Learn more in the docs

Ready to Get Started?

Start profiling threat infrastructure today: uncover actor networks, track emerging campaigns, and get early warning before indicators go public. Choose a plan that fits your team's scale.

Sign Up Now