DataFocus Unstructured · Sizing Tool
Estimate the duration of a Discovery scan. Enter your hardware capacity, data volume, and number of enabled classification types. The result is based on measured benchmark tests.
Detection is CPU-bound. Cores are aggregated across servers. The scan engine runs a single shared scan across the entire server group.
It is rarely a bottleneck. Worker count is limited only when memory is below 1.5 GB per core.
Very small files (millions of tiny objects) shift the cost from reading to scheduling and slow down the scan. Large files have almost no additional impact on duration.
Text extraction is fast, while OCR on scanned pages is several times slower. The number of findings does not affect duration because detection cost is based on bytes processed.
Benchmark tests were run with the default catalog and approximately 20 classification types. Each additional enabled rule adds processing overhead per byte. This effect is modeled below.
The calculation is shown transparently so you can verify it.
| Configuration | Data | Measured | Model |
|---|---|---|---|
| 16 cores · 1 server | 1 TB | 28.8 hr | ~27 hr |
| 16 cores · 1 server | 500 GB | 14.2 hr | ~14 hr |
| 48 cores · 4 servers | 1 TB | 11.8 hr | ~11 hr |
The model column shows CPU processing time only and excludes small-file scheduling overhead. Therefore, the live estimate may be slightly higher for a specific file count.