Ceph RGW Deep Dive: Building an S3 Endpoint with BGP and ECMP Load Distribution

Modern storage systems increasingly need to present a highly available and horizontally scalable S3 interface. Ceph RGW has emerged as one of the most flexible S3 gateways available, but to truly...

A Scalable S3 Ingress Architecture Using Ceph RGW, HAProxy and VyOS

Modern storage systems increasingly need to present a highly available and horizontally scalable S3 interface. Ceph RGW has emerged as one of the most flexible S3 gateways available, but to truly unlock its potential you need a resilient ingress layer that can distribute traffic across multiple gateway nodes without relying on traditional active standby patterns such as VRRP.

In this deep dive, we design and implement a compact, reproducible S3 ingress architecture that uses BGP and Equal Cost Multipath Routing (ECMP) to distribute traffic across multiple HAProxy instances. The goal is to create a fully active active S3 entry point where every gateway node participates simultaneously in request handling. This model mirrors large scale production deployments, but is built entirely inside a KVM based lab environment to make it easy to reproduce the concepts and understand how the components interact.

All networking is virtualized on top of KVM to create a reproducible, portable lab environment. The setup is intentionally simplified and serves as a controlled test installation to demonstrate the underlying mechanics and end to end behavior of the architecture.

1. Overview of the Architecture

The design consists of:

  • A three node Ceph cluster, each running a colocated RGW instance.
  • Two HAProxy nodes forming the S3 ingress tier.
  • One VyOS router acting as the routing layer that performs BGP peering and ECMP flow distribution.
  • One client host that uploads objects through the S3 VIP.

The key idea is simple. Each HAProxy node advertises the same S3 VIP into BGP. VyOS receives two equally valid routes for that VIP and installs both as next hops. With ECMP enabled, VyOS load balances flows across both HAProxy nodes without any shared state or failover protocol.

The result is a clean, fully distributed design: no floating IPs, no keepalived, no leader node, and no single point of failure on the S3 ingress layer.

2. Logical Network Topology

The environment uses two logical networks:

  • A Service Fabric Network (100.64.0.0/16) that interconnects VyOS and the two HAProxy nodes.
  • A Client Access Network (10.90.0.0/24) that connects the client to VyOS.

The RGW nodes live behind the HAProxy tier and are not directly routed.

3. BGP and ECMP in Practice

Each HAProxy node runs FRR and advertises the S3 VIP as a host route:

network 10.80.0.10/32

Both nodes advertise the same route and use the same AS number. From VyOS’s perspective, the two advertisements are fully equal. When ECMP is enabled, VyOS installs both next hops into the routing table:

10.80.0.10/32
nexthop 100.64.0.11
nexthop 100.64.0.12

This is true Layer 3 active active load balancing. There is no leader election. There is no VIP floating. The routing fabric automatically balances the flows.

To confirm ECMP behavior on VyOS:

ip route get 10.80.0.10

Running the command repeatedly shows alternating next hops as the kernel hashes the flow keys.

4. How Multipart Uploads Enable Parallelism

S3 multipart upload is inherently parallel. When a client uploads a file with several gigabytes in size, it typically splits the object into many
independent parts. Each part results in a separate HTTP PUT request and therefore a separate TCP connection.

The ECMP router treats each connection independently. Some connections are hashed to HAProxy1, others land on HAProxy2. This perfectly aligns with Ceph RGW’s architecture where each RGW node is stateless and writes directly to the Ceph object store. The parts converge when the final multipart completion request is made.

The sequence looks like this:

This is not session sharing. It is flow based parallelization, and it fits perfectly with the multipart model.

5. Verifying Load Distribution

On VyOS
Check BGP state:

show ip bgp summary

Inspect route installation:

show ip route 10.80.0.10

Test ECMP hashing:

ip route get 10.80.0.10

On HAProxy
List active sessions:

ss -tn sport = :443

Query HAProxy statistics:

echo "show stat" | socat stdio /var/run/haproxy.sock

On RGW Nodes
Track real time ingress:

tail -f /var/log/ceph/ceph-client.rgw.*.log

Count active connections:

ss -tn | grep :7480 | wc -l

On Ceph
Inspect multipart state:

radosgw-admin multipart list --bucket <bucket>

Inspect bucket stats:

radosgw-admin bucket stats --bucket <bucket>

This combination of tools provides full visibility into each step of the traffic flow.

6. Failure Scenarios

This architecture gracefully handles node failures without any floating IP mechanism.

HAProxy Failure
If one HAProxy node goes offline, it simply withdraws its BGP advertisement. VyOS removes that next hop from the ECMP set. All new connections automatically go to the surviving HAProxy.

RGW Failure
Since RGWs are stateless and access Ceph directly, losing one affects only a subset of multipart parts. HAProxy routes around the missing backend.

VyOS Failure
Because the core logic sits in VyOS, it should run on a reliable host. If you want redundancy, you can extend the model with a pair of VyOS routers and an upstream BGP speaker.

7. Summary

This lab demonstrates a clean, scalable, and highly parallel S3 ingress architecture using Ceph RGW, HAProxy, and BGP ECMP. By letting the routing fabric carry the VIP and distribute flows, the need for shared state or failover protocols is eliminated. Each HAProxy node participates actively, and the multipart nature of S3 uploads ensures strong parallelism across the entire path.

The result is a design that mirrors large scale production deployments yet remains accessible and reproducible using nothing more than KVM, VyOS, HAProxy, and a small Ceph cluster. It is a practical example of how modern distributed object storage can be exposed in a resilient and load balanced way using pure Layer 3 techniques.

Hands-On: Full Configuration for VyOS, HAProxy and FRR

This section shows the exact configuration used in the lab. The goal is not to provide a production hardened blueprint, but to give a reproducible reference that demonstrates how BGP, ECMP, HAProxy and Ceph RGW work together end to end.

1. Network Layout and Assigned Addresses

Service Fabric Network (between VyOS and HAProxy nodes)

Service Fabric Network (between VyOS and HAProxy nodes)
VyOS eth0 100.64.0.10/16
HAProxy1 eth0 100.64.0.11/16
HAProxy2 eth0 100.64.0.12/16

Client Network (between Client and VyOS)
VyOS eth0 10.90.0.1/24
Client br-sfn 10.90.0.10/24

VIP announced by both HAProxy nodes

10.80.0.10/32

2. VyOS Configuration (Routing Layer)

The VyOS router receives BGP announcements from both HAProxy nodes and installs both as equal-cost routes.

Enter configuration mode:

configure

Set basic system parameters:

set system host-name vyos
set system ipv6 disable

Interface configuration

set interfaces ethernet eth0 address 100.64.0.10/16
set interfaces ethernet eth0 address 10.90.0.1/24

BGP configuration (AS 65000)

set protocols bgp 65000 router-id 1.1.1.0
set protocols bgp 65000 parameters bestpath as-path multipath-relax
set protocols bgp 65000 maximum-paths ebgp 2

BGP neighbors (both HAProxy nodes)

set protocols bgp 65000 neighbor 100.64.0.11 remote-as 65001
set protocols bgp 65000 neighbor 100.64.0.11 address-family ipv4-unicast
set protocols bgp 65000 neighbor 100.64.0.12 remote-as 65001
set protocols bgp 65000 neighbor 100.64.0.12 address-family ipv4-unicast

Commit everything:

commit
save
exit

Verify:

show ip bgp summary
show ip route 10.80.0.10
ip route get 10.80.0.10

If ECMP is working, you will see:

10.80.0.10
* via 100.64.0.11
* via 100.64.0.12

3. FRR Configuration on HAProxy Nodes

Each HAProxy node runs FRR and advertises the same VIP.

FRR config on HAProxy1
File: /etc/frr/frr.conf

frr version 8.4
frr defaults traditional
hostname haproxy1
service integrated-vtysh-config

router bgp 65001
bgp router-id 100.64.0.11
no bgp ebgp-requires-policy
no bgp default require-policy
neighbor 100.64.0.10 remote-as 65000
address-family ipv4 unicast
neighbor 100.64.0.10 activate
network 10.80.0.10/32
no bgp network import-check
redistribute connected
exit-address-family
!

FRR config on HAProxy2
File: /etc/frr/frr.conf

frr version 8.4
frr defaults traditional
hostname haproxy2
service integrated-vtysh-config

router bgp 65001
bgp router-id 100.64.0.12
no bgp ebgp-requires-policy
no bgp default require-policy
neighbor 100.64.0.10 remote-as 65000
address-family ipv4 unicast
neighbor 100.64.0.10 activate
network 10.80.0.10/32
no bgp network import-check
redistribute connected
exit-address-family
!

Restart FRR:

sudo systemctl restart frr

Verify BGP:

vtysh -c "show ip bgp summary"

Both should be in Established state.

4. HAProxy Configuration

Each HAProxy node acts as a reverse proxy for the RGW instances. The backend servers are the RGW ports on each Ceph node.

File: /etc/haproxy/haproxy.cfg

global
maxconn 10000
log /dev/log local0

defaults
log global
mode http
option httplog
timeout connect 5s
timeout client 300s
timeout server 300s
frontend s3_frontend
bind *:80
bind *:443 ssl crt /etc/haproxy/certs/s3.pem
default_backend rgw_backend
backend rgw_backend
balance roundrobin
server rgw1 10.<ceph-node1-ip>:7480 check
server rgw2 10.<ceph-node2-ip>:7480 check
server rgw3 10.<ceph-node3-ip>:7480 check

Reload:

sudo systemctl reload haproxy

5. Client-Side Routing

The client routes the S3 VIP and the service fabric network through VyOS, without touching the default gateway.

ip addr add 10.90.0.10/24 dev br-sfn
ip route add 10.80.0.10/32 via 10.90.0.1
ip route add 100.64.0.0/16 via 10.90.0.1

6. Testing the End-to-End Path

Upload a 10 GB test file using the MinIO client:

mc cp file_10g.bin s3/test-bucket

During the upload:

  • ss -tn sport = :443 on HAProxy shows flows arriving on both nodes.
  • ip route get 10.80.0.10 on VyOS alternates between next hops.
  • RGW logs show multipart PUTs arriving on several nodes in parallel.

Load Distribution in Practice

To validate the end-to-end behavior, the HAProxy statistics view provides a clear picture of how traffic is distributed across both gateway nodes and all three RGW backends. During a large multipart upload, each HAProxy instance receives an active share of incoming S3 requests, and each request is forwarded to one of the RGW daemons behind it.

Below is the statistics view from the first HAProxy node, captured in the middle of a 10 GB multipart upload:

HAProxy Node 1 (H1)  -  Active Sessions and Backend Utilization

This snapshot shows multiple active frontend connections and an even distribution across the three RGW backends. Because S3 multipart uploads create numerous parallel PUT requests, the load naturally spreads across all available nodes.

The second HAProxy instance receives its own share of flows. Since VyOS installs both next hops via ECMP, each TCP session is independently hashed and forwarded to either gateway.

HAProxy Node 2 (H2)  -  Active Sessions and Backend Utilization

The behavior closely matches the first node. Both HAProxy instances act as fully active ingress points, and all three RGWs receive multipart PUT requests concurrently.

Together, the two screenshots demonstrate the exact goal of this design: traffic is evenly balanced across both HAProxy nodes and all RGW daemons, with no shared IP failover mechanism and no single node acting as a bottleneck.

This confirms full ECMP load distribution.

Closing Thoughts

This lab shows how much efficiency you gain when you stop treating load balancing as a shared IP problem and instead let the routing layer handle it. By advertising the VIP through BGP and allowing VyOS to install multiple equal-cost paths, the S3 ingress tier becomes fully active. Both HAProxy nodes contribute live capacity, and the RGWs handle traffic in parallel without any coordination beyond what Ceph already provides.

The environment is intentionally small, but the principles scale cleanly. Whether you operate a handful of gateways or a larger edge footprint, the behavior remains predictable and the failure domains stay simple. One of the most compelling aspects of this pattern is its minimalism. There are no clustering daemons, no state synchronization mechanisms, and no leader election logic. The routing fabric carries the VIP, HAProxy terminates TLS and distributes traffic, and Ceph preserves the object semantics underneath.

If you want to extend the lab further, you can introduce a second VyOS router, test more advanced BGP policies, or explore different HAProxy balancing strategies. All of these variations build on the same foundation demonstrated here, and every component remains easy to reason about as the design grows.

Let's talk about your infrastructure.

An architecture review, a new environment from the ground up or support in operations: talk directly to the engineers who will deliver it.