Skip to main content

Platform Documentation

Learn how to use ClusterHawk for IP clustering and threat detection

Search Documentation

1
Submit IPs

Upload your IP addresses of interest through our secure interface. Our platform handles datasets up to 5000 addresses.

2
Analysis

Our deterministic ensemble pipeline analyzes patterns, identifies relationships, and generates threat intelligence automatically — same input, same clusters, same reasoning, every run.

3
Receive reports

Get comprehensive threat intelligence reports with IOCs, YARA rules, and hunting queries.

4
Execute hunting queries

Use our automated hunting query execution service to validate findings and monitor for new threats.

User Guide

Pipeline Tier Comparison


Pipeline Tier Comparison

All four analysis pipeline tiers (Core, Deep, Advanced, and Neural Network) share the same analytical foundation: the same multi-method clustering ensemble, the same anomaly detection, the same labeling and reporting. What differs is how deeply the system analyzes IP metadata before profiling.

What All Tiers Share
  • Multi-method clustering ensemble with dynamic weighted voting
  • Proprietary anomaly detection pipeline
  • Automated actor labeling with confidence scores
  • Noise intelligence mining (automatic re-analysis of noise clusters)
  • Validation query generation
  • STIX 2.1 and MISP exports, 3D visualization, automated reports
  • Cluster profiling with stability metrics
The Analytical Difference: Feature Representation

The tiers differ in how deeply the system analyzes IP metadata before profiling. Higher tiers detect increasingly subtle and complex infrastructure relationships, at the cost of longer processing time and more compute. Each tier builds on a proprietary analysis methodology validated through academic research.

Core Infrastructure Profiling

Available from: Analyst tier (€499/month) and above

The Analyst tier at €499/month is designed as an entry point for teams to experience ClusterHawk's analytical depth on real data before scaling up. Custom quotas are available through [email protected].

Optimized for feature granularity. The system preserves fine-grained indicators across your dataset, maintaining sensitivity to rare threat actor signatures: even subtle patterns that distinguish small clusters of 5 IP addresses within a dataset of 400.

  • Best for: Evaluating ClusterHawk's capabilities, initial triage, known-pattern matching, smaller datasets
  • Strengths: Highest feature granularity, sensitive to rare threat actor signatures and behavioral fingerprints
  • Trade-offs: May miss complex, multi-dimensional relationships between infrastructure attributes. Less effective on highly heterogeneous datasets

Core infrastructure profiling produces meaningful results on datasets as small as 50 IPs, though cluster stability and separation improve significantly above 150 IPs. At smaller scales, expect fewer distinct clusters and wider confidence ranges. The results are still actionable for triage, but higher-confidence groupings emerge with more data.

Deep Infrastructure Profiling

Available from: Team tier (€899/month) and above

Goes beyond surface-level grouping to capture how infrastructure clusters relate to each other structurally. The system self-tunes its analysis parameters for each dataset, revealing campaign-level relationships that simpler approaches miss.

  • Best for: Mixed-type data, infrastructure mapping, campaign-level pattern discovery, datasets where Core profiling misses important groupings
  • Strengths: Better structural coherence, self-tuning analysis, reveals how clusters relate to each other, not just what is inside them
  • Trade-offs: Some granular indicator visibility is traded for structural quality
Advanced Infrastructure Profiling

Available from: Professional tier (€1,799/month) and above

Learns which combinations of IP attributes are most discriminating for your specific dataset before profiling, automatically adapting its analytical depth to the data. This dataset-specific optimization catches infrastructure patterns that fixed-approach methods miss entirely.

  • Best for: Sophisticated threat infrastructure, datasets where surface features are misleading, high-stakes analysis requiring maximum cluster quality
  • Strengths: Discovers complex, multi-order relationships between IP characteristics. Dataset-specific optimization. Deepest analytical capability for unsupervised analysis
  • Trade-offs: Requires sufficient data volume (100+ IPs recommended). Lower direct interpretability
Neural Network Models

Available from: Professional tier (€1,799/month) and above

A deep learning model that learns which infrastructure features matter most for distinguishing clusters in your specific dataset. Produces calibrated confidence scores you can use for threshold-based operational decisions. The model automatically optimizes itself for your data.

  • Best for: Large datasets (500+ IPs), repeated analysis of similar infrastructure types. A model trained on a full Professional-tier dataset (2,500 IPs) produces high-quality predictions reusable across many subsequent runs at minimal cost. Also ideal for operationally critical decisions requiring calibrated confidence
  • Strengths: Learns dataset-specific feature importance. Calibrated confidence scores for threshold-based decisions. Designed for train-once, predict-many workflows
  • Trade-offs: No standalone unsupervised mode (always includes training). Most compute-intensive during training. Requires 500+ IPs for meaningful results
Quick Decision Guide
  • Quick triage of a new dataset → Core
  • Detect infrastructure patterns beyond individual IOCs → Deep or above
  • Investigate sophisticated threat actor infrastructure → Advanced
  • Maximize cluster quality for complex datasets → Advanced
  • Run many predictions on similar datasets → Neural Network
  • Need calibrated confidence scores → Neural Network
  • Small datasets (< 100 IPs) → Core
  • APT campaign hunting → Deep or Advanced

Not sure? Start with Core / Basic clustering: it's the right choice for a first look at any dataset.