Storage

Storage platforms that are not allowed to fail.

We design, build and operate block, file and object storage on commodity hardware: highly available, multi-tenant and entirely as code. We have run Ceph in production since 2016.

The challenge

Data keeps growing, retention rules get stricter and traditional storage ties you to one vendor. At the same time, storage is the layer whose failure takes everything else down with it. Mistakes in failure domains, sizing or networking only show when it matters most.

What we do

Capabilities in detail

Architecture and design

Failure domains and data placement, replication or erasure coding, pool and placement-group planning, topologies across several availability zones.

Block, file and object

RBD for OpenStack, Kubernetes and hypervisors, NVMe/TCP and iSCSI; CephFS with NFS exports; S3 on RGW with storage classes, versioning and lifecycle rules.

Multi-tenant S3

Separation through namespaces, restricted credentials, placement targets and quotas per tenant; a clear model for identities, roles and bucket policies; repeatable onboarding.

Compliance and retention

Object Lock (WORM), versioning and lifecycle rules as technical controls for retention requirements.

Capacity and performance

Growth models from usable-to-raw ratio, replication overhead and rebuild reserve; node sizing across CPU, memory, NVMe and network; tuning by measurement, not assumption.

Storage networking

Separate networks for client, cluster, management and out-of-band traffic, port and protocol matrices, MTU requirements, load-balancer and VIP design, S3 endpoints with BGP and ECMP.

Backup and disaster recovery

Cross-zone replication, active/active load distribution, disaster recovery and business continuity concepts, automated restores that are exercised regularly.

Operations and fault analysis

Upgrades and rolling maintenance, adding and draining nodes, rebalancing; diagnosing degraded placement groups, slow requests, drive and network faults.

Migration and enterprise storage

NetApp ONTAP alongside Ceph and as the starting point for a move to software-defined storage, planned vendor-neutral.

PoC, validation, documentation

Hardware PoC before procurement, test campaigns for stability, scaling, redundancy and performance, architecture documents traceable to numbered requirements.

Approach

From requirements to operations.

  1. 01

    Requirements

    Performance, capacity, availability and compliance are stated in measurable terms.

  2. 02

    PoC

    Hardware and design are tested under real load before anything is bought.

  3. 03

    Design and sizing

    Failure domains, network, capacity and growth are defined and documented.

  4. 04

    Build as code

    Unattended installation with cephadm, Ansible and GitOps, including air-gapped environments.

  5. 05

    Acceptance

    Test campaigns against criteria agreed up front: stability, redundancy, performance.

  6. 06

    Operations

    Monitoring, capacity reports, upgrades and fault analysis, with runbooks for your team.

Technology

  • Ceph
  • cephadm
  • Rook-Ceph
  • RGW / S3
  • RBD
  • CephFS
  • NVMe-oF
  • NFS-Ganesha
  • iSCSI
  • NetApp ONTAP
  • Ansible
  • Terraform
  • Prometheus
  • Grafana
  • MAAS

Services in detail

Insights

Let's talk about your infrastructure.

An architecture review, a new environment from the ground up or support in operations: talk directly to the engineers who will deliver it.