Before Installation
The installation of Data Focus consists of 8 key steps. For a successful installation, each step must be completed in the correct order.
You must have sudo rights to execute the installation scripts. Using the root user for installation is not recommended.
The steps are outlined below.
1. Docker Installation
2. Extracting Packages
3. Loading Packages
4. Hostname Configuration
5. HTTPS Configuration
6. Deployment
7. Keycloak Configuration
8. Monitoring
Environment-Specific Considerations¶
Choose your deployment type for specific requirements
Quick Setup Requirements:
- Minimum hardware specifications are sufficient
- HTTPS configuration mandatory
- Standard VM settings work fine
- No special backup requirements
Critical Production Requirements:
- Enhanced hardware specifications recommended
- HTTPS configuration mandatory
- Special virtualization and storage considerations required
Supported Browser Versions¶
Data Focus is tested against the current and the three preceding major versions of each browser below. Keeping the browser up to date is enough — no specific build number is required.
| Browser | Supported versions |
|---|---|
| Google Chrome | Current major version and the three before it |
| Mozilla Firefox | Current major version and the three before it |
| Microsoft Edge | Current major version and the three before it |
| Safari (macOS) | Current major version and the three before it |
Note
Browsers on this list update themselves, so the supported range moves with them. Older major versions may still work but are not tested.
Network Requirements¶
Data Focus is a web application that operates both at the backend and frontend. We utilize Nginx as a reverse proxy to manage traffic to our applications, including those running in Docker containers.
Ports Required
- HTTP:
80 - HTTPS:
443 - SSH:
22
Ensure that these ports are open and properly configured to allow seamless communication between clients and the application.
Outbound access to the data to be scanned
The ports above cover access to Data Focus. Scanning also needs a route from the Data Focus server to the systems that hold the data, and those systems must be resolvable by the server's DNS. The sources in scope and the ports they use are listed in Access to Data Sources.
Hardware Requirements¶
Select a configuration to view details
- CPUs:
16 - Memory:
64 GiB - Disk:
500 GiB - Use Case: POC, and small-scale setups
- CPUs:
32 - Memory:
80 GiB - Disk:
1TB SSD or NVME Storage - Use Case: Production deployments
Storage Specifications¶
Volume Requirements
For all deployments, we need two dedicated volumes on the host:
| Mount Path | Purpose | Recommended Allocation |
|---|---|---|
/datafocus |
Application data storage | 40% of total storage |
/var/lib/docker |
Docker storage | 40% of total storage |
| Unallocated | Reserved for future use | 20% of total storage |
These volumes must be dedicated to the OS, Docker, and application operations. This configuration ensures that even if the VM fails to boot, you can continue operations on a new VM by mounting these volumes.
Note
The recommended configuration ensures optimal performance and scalability for your applications.
Disk Sizing¶
The figures above cover the installation itself and a PoC-sized scan. Beyond that, disk usage grows with the number of findings produced, not with the size of the data that was scanned.
Scanned files are never copied to the server, so a large file share does not consume disk by itself. What consumes disk is one row per finding in the database, plus Kafka retention for the records in transit. A densely populated source can therefore produce far more findings than a much larger but sparsely populated one.
Sizing beyond a PoC
For scans larger than the PoC scope, size the disk from the expected finding volume rather than from the volume of data to be scanned, and review it with the Kafein team before the scan starts. As a reference point, a dense 70 GB corpus has produced close to 20 million findings.
Operating System Prerequisites¶
The following host-level settings are expected on the server before deployment. They are not configured by the installation scripts.
Dedicated Server¶
Data Focus should be the only workload on this server
The scan workers watch system-wide memory usage, not their own. When total memory usage on the host stays above 85% for a sustained period, a worker shuts itself down to prevent an out-of-memory condition, and the scan stops making progress.
This means that unrelated software on the same server — monitoring agents, backup jobs, log collectors, other applications — can stop a scan even when Data Focus itself is within its limits. Allocate the server exclusively to Data Focus.
Shared Memory (/dev/shm)¶
PostgreSQL uses /dev/shm for parallel query execution. Docker's default of 64 MB is not enough and causes parallel queries to fail with a disk-full error.
The bundled Compose file already sets this:
shm_size: ${POSTGRES_SHM_SIZE:-2gb}
Note
No host change is required when the provided Compose file is used. If the database is deployed through your own orchestration, allocate at least 2 GB of shared memory to the PostgreSQL container.
Time Synchronization¶
Keep the server clock synchronized
Kafka, Keycloak tokens and TLS certificate validation are all sensitive to clock drift, and the resulting failures are difficult to diagnose. Ensure an NTP client (chrony or systemd-timesyncd) is installed, enabled and synchronized before installation:
timedatectl status
The output should report System clock synchronized: yes.
Swap¶
Keeping swap disabled — or vm.swappiness set low — is recommended. The memory protection described above measures physical memory, so swapping does not prevent a worker from shutting down; it only makes the system slower before it happens.
Kernel Setting for OpenSearch (Optional)¶
Full text search over findings is served by OpenSearch, which is disabled by default. When it is enabled (OPENSEARCH_ENABLED=true), raise the host's memory map limit:
sudo sysctl -w vm.max_map_count=262144
# persist across reboots:
echo "vm.max_map_count=262144" | sudo tee -a /etc/sysctl.conf
Note
The container runs in single-node mode, so it starts even with the Linux default of 65530. Raising the limit prevents memory map exhaustion once the index grows. File descriptor and memory lock limits for this container are already set in the Compose file and need no host change.
Production-Specific Requirements¶
Additional Requirements for Production Deployments
If you're deploying Data Focus in a production environment, please review and implement the following additional requirements
Virtual Infrastructure Settings¶
Virtualization Configuration
Resource Allocation:
- CPU Overcommit: Maximum 2:1 ratio (recommended 1.5:1)
- Memory Overcommit: Disabled or maximum 1.2:1 ratio
- Vertical Scaling: Ensure resource allocation supports vertical scaling capabilities
- Resource Reservation: Reserve minimum required CPU and memory
Backup and Snapshot Strategy¶
Data Protection Requirements
Mount Disk Snapshots:
- Both
/datafocusand/var/lib/dockervolumes must be snapshot-capable - Pre-deployment snapshots: Mandatory before any updates
- Regular snapshots: Daily incremental, weekly full snapshots
- Retention policy: Minimum 14 days for production data
- Recovery testing: Verify snapshot restore procedures regularly
Security and Network Hardening¶
Production Security
- HTTPS Configuration: Mandatory for production
- SSL Certificates: Use valid certificates (avoid self-signed)
- Firewall Rules: Restrict access to required ports only
- Network Security: Implement proper network segmentation