Skip to main content

Platform Documentation

Learn how to use ClusterHawk for IP clustering and threat detection

Search Documentation

1
Submit IPs

Upload your IP addresses of interest through our secure interface. Our platform handles datasets up to 5000 addresses.

2
Analysis

Our deterministic ensemble pipeline analyzes patterns, identifies relationships, and generates threat intelligence automatically — same input, same clusters, same reasoning, every run.

3
Receive reports

Get comprehensive threat intelligence reports with IOCs, YARA rules, and hunting queries.

4
Execute hunting queries

Use our automated hunting query execution service to validate findings and monitor for new threats.

User Guide

Choosing the Right Pipeline


Choosing the Right Pipeline
Choosing the Right Pipeline for Your Needs

Different pipeline types are optimized for specific use cases:

Not sure? Start with Core / Basic clustering: it's the right choice for a first look at any dataset.

  • Core Infrastructure Profiling: Ideal for initial exploration of IP datasets without model training. Use this when:
    • You're analyzing a new dataset for the first time
    • You need preliminary assessment of a dataset
    • You want to identify basic groupings without predictive capabilities
  • Deep Infrastructure Profiling: Provides improved pattern discovery with structural relationship analysis. Use this when:
    • Core profiling misses important patterns in your data
    • You need better results but have resource constraints
    • Your dataset contains moderately complex relationships
  • Advanced Infrastructure Profiling: Best when you need the deepest analytical capability without saving a model. Use this when:
    • You need higher quality clusters with more nuanced patterns
    • You want to analyze complex relationships between IPs
    • You don't need to reuse the analysis approach on new data
  • Core Profile Model Training: Choose this when you need to create a reusable model for moderate-complexity IP datasets. Use this when:
    • You want to analyze similar datasets repeatedly
    • You need predictive capabilities but have resource constraints
    • Your dataset has clear patterns but moderate complexity
  • Deep Profile Model Training: Optimal for creating models that capture more sophisticated structural patterns. Use this when:
    • Core models miss important relationships in your data
    • You need stronger structural analysis with detection capability
    • Your dataset contains moderately complex infrastructure relationships
  • Advanced Profile Model Training: Optimal for complex datasets with subtle patterns. Use this when:
    • You're analyzing sophisticated threat infrastructure
    • You need maximum accuracy for high-stakes security decisions
    • Your dataset contains subtle or complex relationships
  • Neural Network Pipeline: Best for discovering non-linear patterns in very large datasets. Use this when:
    • You have large datasets with complex interrelationships
    • Standard clustering methods miss important patterns
    • You need to detect subtle anomalies within similar IP behaviors
IP Requirements by Pipeline Type

Each pipeline type has specific IP address requirements for optimal performance and feature availability:

  • All Pipelines: A minimum of 10 IP addresses is required for clustering to run successfully (works with as few as 2 IPs, but with limited results). At least 25 IPs are recommended for automatic report generation and noise analysis.
  • Training Pipelines (Core/Deep/Advanced): A minimum of 100+ IPs is recommended for achieving acceptable model performance, depending on the dataset.
  • Neural Network Pipelines: At least 500 IPs are recommended for effective deep learning model performance.
  • Recommended Job Size: For discovery and best clustering results, submit 200-700 IPs per job. If you already have validated/proven datasets or are running on validated datasets and need exhaustive coverage, you can use higher counts up to your plan limits (with proportionally longer runtimes).
Performance Considerations

Training pipelines require approximately 1.5x the time of clustering pipelines due to the additional model training phase.

Note: Processing times may vary based on data complexity and pipeline type.