Rebuilding my homelab architecture

My homelab has changed quite a bit over time. What started as a fairly simple setup with a few virtual machines, Docker containers, internal services, and an OPNsense VM gradually grew into a larger environment with Proxmox, Kubernetes, internal DNS, S3 storage, monitoring, VPN access, and more externally exposed services.
The old setup worked well enough, but a lot of it had grown organically. Services were added when I needed them, networking changed over time, and some components ended up doing more than they were originally intended to do. The next version of the homelab is therefore mostly about making the existing setup more structured, more redundant, and easier to expand.
The old setup
The previous environment already used a three-node Proxmox cluster as the main virtualization platform. Most workloads ran as virtual machines or containers, while a five-node k3s cluster handled containerized applications.
Networking was centered around an HPE Aruba 2530-48G switch, with OPNsense running as a virtual router and firewall. Cloudflare handled external DNS and tunneling for services that needed to be reachable from the internet.
Storage consisted of a mix of local storage, ZFS, S3-compatible storage, and NAS storage depending on the workload.
This setup was useful because it exposed a few weak points fairly quickly. Running routing and firewalling inside the same virtual infrastructure it was protecting worked, but it also created unnecessary dependencies. Storage for infrastructure and storage for user files also had very different requirements, and treating them as one problem made things less clear.
The new setup separates those responsibilities much more clearly.
The new architecture
The main virtualization platform remains a three-node Proxmox cluster.
Each node currently has:
- 8 CPU cores
- 16 GB RAM
- Proxmox VE
The biggest change here is storage. The cluster will use Ceph as shared storage for virtual machines and containers. This allows workloads to move between nodes and makes Proxmox HA much more useful.
Instead of a VM being tied to disks inside one particular host, its storage is available across the cluster.
That also means the cluster can tolerate individual node failures much more cleanly than before.
Network overview
The diagram above shows the main separation between networking, virtualization, storage, Kubernetes, and dedicated compute.
Networking
The biggest networking change is the removal of the virtual OPNsense router.
It will be replaced by two Check Point 1450 appliances configured in a failover setup.
The Check Point pair will handle:
- WAN connectivity
- Firewalling
- Routing between VLANs
- VPN access
- Network segmentation
Moving routing out of the Proxmox cluster removes an awkward dependency where the virtual infrastructure partly depended on a router running inside that same infrastructure.
The HPE Aruba 2530-48G remains the main switch for the environment and connects the different physical systems and VLANs together.
This also makes the network topology more similar to a traditional enterprise setup, with dedicated routing and firewalling in front of the compute infrastructure.
Proxmox and Ceph
The three Proxmox nodes will form the main highly available compute platform.
Ceph will provide shared storage between the nodes, while Proxmox handles VM management, clustering, and HA.
This storage is specifically intended for infrastructure workloads such as:
- Virtual machine disks
- Container disks
- Kubernetes nodes
- Infrastructure services
One thing I wanted to avoid this time was mixing this storage with large amounts of user data.
VM storage and user file storage have different access patterns and different priorities, so they will be separated.
Separate Ceph storage for user data
A second Ceph environment will be built around the 1 TB drives available in the cluster.
Instead of storing virtual machine disks, this storage will primarily contain user files.
The storage will be exposed over S3, with Cloudreve providing the user-facing frontend.
The path will roughly look like:
Cloudreve → S3 → Ceph object storage
This gives applications an S3-compatible interface while keeping the underlying storage distributed across multiple machines.
Separating the two Ceph use cases also makes the design easier to understand.
One storage environment exists for infrastructure.
The other exists for data.
Kubernetes
The five-node k3s cluster will remain the main platform for containerized services.
The k3s nodes run on top of the Proxmox environment and host most of the services that do not require a dedicated virtual machine.
This includes things such as:
- Cloudreve
- Monitoring
- Internal applications
- Web services
- Supporting infrastructure
Keeping Kubernetes virtualized also makes the nodes relatively disposable. The important parts are the cluster configuration, persistent storage, and application definitions rather than the individual VM itself.
The longer-term goal is to keep as much configuration as possible reproducible rather than depending on manually configured systems.
Dedicated 2U compute server
Not everything fits particularly well inside the main cluster.
For heavier workloads I am currently building a separate 2U server running NixOS.
The system currently has:
- 16 CPU cores
- 64 GB RAM
- RTX 4060
- 2 × 2 TB drives in RAID
- NixOS
This machine will be used for workloads where raw compute performance or direct access to a GPU is more important than high availability.
Planned workloads include:
- Video encoding
- Game streaming
- GPU workloads
- File serving
- Compute-heavy tasks
- Windows VM
- Ubuntu VM
The RTX 4060 is especially useful for workloads such as hardware video encoding and game streaming.
This server is intentionally separate from the Proxmox cluster. If it is unavailable, the core infrastructure should continue running normally.
Cloudflare and external services
Cloudflare will continue to sit between parts of the homelab and the public internet.
It currently provides things such as:
- External DNS
- Cloudflare tunnels
- Controlled access to internal services
Not every workload runs inside the homelab either.
Some services, particularly client websites and other public-facing workloads, continue to run on external cloud infrastructure.
This keeps public workloads separate from the home network while still allowing DNS and access to be managed from the same place.
What changed from the old setup
The individual technologies are not radically different from what I was already using. Proxmox, Kubernetes, Cloudflare, S3 storage, VLANs, and self-hosted applications were already part of the environment.
The main change is how they are divided.
The previous setup made it easy to add another VM, container, disk, or service, but over time those additions created dependencies that were not always obvious.
The new design tries to keep those responsibilities separate:
Check Point handles the network edge.
Aruba handles switching.
Proxmox handles general virtualization.
Ceph handles distributed infrastructure storage.
k3s handles containerized workloads.
Ceph S3 handles user data.
Cloudreve provides the file frontend.
NixOS 2U server handles high-compute and GPU workloads.
Cloudflare handles external DNS and tunneling.
That separation is probably the biggest thing I took from running the previous setup. A homelab can become surprisingly complex once storage, networking, virtualization, and applications all start depending on each other.
The new setup is mostly an attempt to keep those boundaries clearer while still leaving enough room to experiment.
There is still quite a lot to build, especially around Ceph, failover, VLAN design, monitoring, backups, and the 2U server, so this architecture will probably change again as the rebuild progresses.