Skip to content

Before Installation

The installation of Data Focus consists of 8 key steps. For a successful installation, each step must be completed in the correct order.

You must have sudo rights to execute the installation scripts. Using the root user for installation is not recommended.

The steps are outlined below.

1. Docker Installation
2. Extracting Packages
3. Loading Packages
4. Hostname Configuration
5. HTTPS Configuration
6. Deployment
7. Keycloak Configuration
8. Monitoring

Environment-Specific Considerations

Choose your deployment type for specific requirements

Quick Setup Requirements:

  • Minimum hardware specifications are sufficient
  • HTTPS configuration mandatory
  • Standard VM settings work fine
  • No special backup requirements

Critical Production Requirements:

  • Enhanced hardware specifications recommended
  • HTTPS configuration mandatory
  • Special virtualization and storage considerations required

See detailed production requirements below

Supported Browser Versions

Data Focus is tested against the current and the three preceding major versions of each browser below. Keeping the browser up to date is enough — no specific build number is required.

Browser Supported versions
Google Chrome Current major version and the three before it
Mozilla Firefox Current major version and the three before it
Microsoft Edge Current major version and the three before it
Safari (macOS) Current major version and the three before it

Note

Browsers on this list update themselves, so the supported range moves with them. Older major versions may still work but are not tested.

Network Requirements

Data Focus is a web application that operates both at the backend and frontend. We utilize Nginx as a reverse proxy to manage traffic to our applications, including those running in Docker containers.

Ports Required

  • HTTP: 80
  • HTTPS: 443
  • SSH: 22

Ensure that these ports are open and properly configured to allow seamless communication between clients and the application.

Outbound access to the data to be scanned

The ports above cover access to Data Focus. Scanning also needs a route from the Data Focus server to the systems that hold the data, and those systems must be resolvable by the server's DNS. The sources in scope and the ports they use are listed in Access to Data Sources.

Hardware Requirements

Select a configuration to view details

  • CPUs: 16
  • Memory: 64 GiB
  • Disk: 500 GiB
  • Use Case: POC, and small-scale setups
  • CPUs: 32
  • Memory: 80 GiB
  • Disk: 1TB SSD or NVME Storage
  • Use Case: Production deployments

Storage Specifications

Volume Requirements

For all deployments, we need two dedicated volumes on the host:

Mount Path Purpose Recommended Allocation
/datafocus Application data storage 40% of total storage
/var/lib/docker Docker storage 40% of total storage
Unallocated Reserved for future use 20% of total storage

These volumes must be dedicated to the OS, Docker, and application operations. This configuration ensures that even if the VM fails to boot, you can continue operations on a new VM by mounting these volumes.

Note

The recommended configuration ensures optimal performance and scalability for your applications.

Disk Sizing

The figures above cover the installation itself and a PoC-sized scan. Beyond that, disk usage grows with the number of findings produced, not with the size of the data that was scanned.

Scanned files are never copied to the server, so a large file share does not consume disk by itself. What consumes disk is one row per finding in the database, plus Kafka retention for the records in transit. A densely populated source can therefore produce far more findings than a much larger but sparsely populated one.

Sizing beyond a PoC

For scans larger than the PoC scope, size the disk from the expected finding volume rather than from the volume of data to be scanned, and review it with the Kafein team before the scan starts. As a reference point, a dense 70 GB corpus has produced close to 20 million findings.

Operating System Prerequisites

The following host-level settings are expected on the server before deployment. They are not configured by the installation scripts.

Dedicated Server

Data Focus should be the only workload on this server

The scan workers watch system-wide memory usage, not their own. When total memory usage on the host stays above 85% for a sustained period, a worker shuts itself down to prevent an out-of-memory condition, and the scan stops making progress.

This means that unrelated software on the same server — monitoring agents, backup jobs, log collectors, other applications — can stop a scan even when Data Focus itself is within its limits. Allocate the server exclusively to Data Focus.

Shared Memory (/dev/shm)

PostgreSQL uses /dev/shm for parallel query execution. Docker's default of 64 MB is not enough and causes parallel queries to fail with a disk-full error.

The bundled Compose file already sets this:

YAML
shm_size: ${POSTGRES_SHM_SIZE:-2gb}

Note

No host change is required when the provided Compose file is used. If the database is deployed through your own orchestration, allocate at least 2 GB of shared memory to the PostgreSQL container.

Time Synchronization

Keep the server clock synchronized

Kafka, Keycloak tokens and TLS certificate validation are all sensitive to clock drift, and the resulting failures are difficult to diagnose. Ensure an NTP client (chrony or systemd-timesyncd) is installed, enabled and synchronized before installation:

Bash
timedatectl status

The output should report System clock synchronized: yes.

Swap

Keeping swap disabled — or vm.swappiness set low — is recommended. The memory protection described above measures physical memory, so swapping does not prevent a worker from shutting down; it only makes the system slower before it happens.

Kernel Setting for OpenSearch (Optional)

Full text search over findings is served by OpenSearch, which is disabled by default. When it is enabled (OPENSEARCH_ENABLED=true), raise the host's memory map limit:

Bash
sudo sysctl -w vm.max_map_count=262144
# persist across reboots:
echo "vm.max_map_count=262144" | sudo tee -a /etc/sysctl.conf

Note

The container runs in single-node mode, so it starts even with the Linux default of 65530. Raising the limit prevents memory map exhaustion once the index grows. File descriptor and memory lock limits for this container are already set in the Compose file and need no host change.


Production-Specific Requirements

Additional Requirements for Production Deployments

If you're deploying Data Focus in a production environment, please review and implement the following additional requirements

Virtual Infrastructure Settings

Virtualization Configuration

Resource Allocation:

  • CPU Overcommit: Maximum 2:1 ratio (recommended 1.5:1)
  • Memory Overcommit: Disabled or maximum 1.2:1 ratio
  • Vertical Scaling: Ensure resource allocation supports vertical scaling capabilities
  • Resource Reservation: Reserve minimum required CPU and memory

Backup and Snapshot Strategy

Data Protection Requirements

Mount Disk Snapshots:

  • Both /datafocus and /var/lib/docker volumes must be snapshot-capable
  • Pre-deployment snapshots: Mandatory before any updates
  • Regular snapshots: Daily incremental, weekly full snapshots
  • Retention policy: Minimum 14 days for production data
  • Recovery testing: Verify snapshot restore procedures regularly

Security and Network Hardening

Production Security

  • HTTPS Configuration: Mandatory for production
  • SSL Certificates: Use valid certificates (avoid self-signed)
  • Firewall Rules: Restrict access to required ports only
  • Network Security: Implement proper network segmentation