Capacity, Power, Thermal & Expansion Envelopes
Known compute, memory, storage, network, switch-port, and PoE limits, plus the electrical and thermal measurements required before expansion.
Capacity is constrained by assigned hardware, physical ports, storage topology, and the need to protect the edge firewalls from competing workloads. Inventory limits are known; whole-site electrical load, UPS runtime, cooling capacity, and measured thermal headroom are not yet recorded.
Electrical and Thermal Envelopes Are Unmeasured
No source records circuit ratings, PDU allocation, UPS capacity or runtime, idle and peak wattage, rack inlet temperature, sustained exhaust temperature, cooling capacity, or alarm thresholds. Hardware inventory is not evidence that the room or power path can support simultaneous peak load.
Known Resource Envelope
| Resource | Current or target envelope | Expansion boundary |
|---|---|---|
| Multiport 10 Gb NICs | Six total; five X710-T4 cards at Site B and one XL710/X710-class card in sa-stor-01 | All six are allocated; expansion requires procurement or an explicit reallocation |
| ThinkPad 10 Gb NICs | Two X550-T2 cards, one per Site A ThinkPad | No spare X550-T2 is recorded |
| Server RAM pool | Seventeen 32 GB DDR4 ECC RDIMMs for Supermicro systems | Eleven more 32 GB sticks are needed to bring every non-5049 Supermicro system to 128 GB |
sa-stor-01 RAM | 96 GB installed | Supported documented options are 192 GB, 288 GB, or 384 GB; do not mix RDIMM and LRDIMM |
| Site A ZFS | 8–12 Samsung SM863 1.92 TB SSDs in mirror vdevs | Expansion stays in even-drive mirror pairs; usable capacity is half of raw mirror capacity |
| Site B Ceph | Five nodes, 4–6 OSDs per node, 20–30 OSDs total, replication size 3 | Site-local only; reserve 2–4 enterprise SSDs as cold spares |
| Enterprise SSD pool | 20 SM863, 10 Micron 5200 MAX, 8 Micron 5300 Pro, all 1.92 TB | Allocation must preserve the cold-spare target and workload endurance roles |
| Boot media | Seven recorded 512 GB M.2/SATA M.2 devices plus the existing sa-stor-01 boot drive | Do not consume 1.92 TB enterprise SSDs for boot unless forced |
| Site B core ports | 41 of 48 sb-sw-01 ports assigned | Seven physical ports remain; VLAN and NIC design still governs whether they are usable |
Site A Switching and PoE Envelope
Each USW-Pro-XG-10-PoE provides ten 10 GbE RJ45 ports and two 10 Gb SFP+ ports. The RSTP triangle consumes all six SFP+ endpoints.
| Switch | Assigned RJ45 ports | Reserved expansion or recovery | SFP+ state |
|---|---|---|---|
sa-sw-01 | Ports 1–6 | Ports 7–9 disabled spares; port 10 disabled emergency admin | Ports 11–12 used by fabric |
sa-sw-02 | Ports 1–5 | Ports 6–9 disabled spares; port 10 disabled emergency admin | Ports 11–12 used by fabric |
sa-sw-03 | Ports 1–6 | Port 7 future AP; ports 8–9 disabled spares; port 10 disabled emergency admin | Ports 11–12 used by fabric |
PoE is disabled on every non-PoE endpoint. sa-ap-01 is the only enabled target PoE load: PoE++ 802.3bt at approximately 29 W on sa-sw-03 port 6. The future AP port remains disabled with PoE off until used.
PoE Budget Is Not the Site Power Budget
The AP's approximate draw is documented, but switch PoE capacity, aggregate switch draw, server load, UPS sizing, and circuit headroom are not. A second AP requires both a port-profile change and a verified power budget.
Workload Capacity Guardrails
| System class | Suitable workload | Capacity guardrail |
|---|---|---|
| E200 edge hosts | OPNsense, OpenBao, UOS, DNS helper, small proxy, monitoring agent | No heavy databases, Ceph OSDs, storage-heavy VMs, or heavy Kubernetes workers |
| Site A ThinkPads | CI, development, GPU/AI/media, light VMs or workers | No IPMI; recovery from a failed network or thermal event may require physical access |
sa-stor-01 | ZFS, PBS-A, DNS, monitoring, databases, core VMs | AQC107 does not carry management or Corosync; protect storage and control services from NIC instability |
| Site B compute nodes | Ceph, Kubernetes/OpenShift, distributed compute | Ceph public and cluster traffic stay on separate directly attached networks |
The E200 rule is a reliability envelope, not a utilization target. Spare CPU and I/O on the edge preserve routing responsiveness for the entire site.
Power and Thermal Evidence Gaps
| Measurement | Current record | Required before material expansion |
|---|---|---|
| Branch-circuit capacity | Not recorded | Circuit voltage, breaker rating, continuous-load allowance, and shared loads |
| PDU and outlet mapping | Not recorded | Device-to-outlet map, PDU rating, and phase/bank balance where applicable |
| UPS capacity and runtime | Not recorded | Rated VA/W, battery health, load percentage, runtime at current and projected peak |
| Device idle and peak draw | Not recorded | Measured watts per server, switch, storage chassis, AP, and WAN equipment |
| Rack thermal state | Not recorded | Inlet and exhaust temperatures at idle and sustained load |
| Cooling headroom | Not recorded | Room/rack heat-removal capacity and acceptable operating range |
| Thermal alerting | Not recorded | Sensor source, warning/critical thresholds, notification path, and response owner |
| Expansion reserve | Not recorded | Maximum additional watts and heat load without circuit, UPS, or cooling changes |
Expansion Gates
- Confirm the target port, VLAN profile, NIC path, and failure-domain effect.
- Measure present electrical load and thermal state under sustained representative workload.
- Verify UPS runtime and circuit headroom with the proposed device included.
- Preserve edge-host CPU/I/O reserve, storage cold spares, and the no-spare NIC constraint.
- Add monitoring and recovery ownership before treating new capacity as production.
Hardware quantities and wiring detail remain in per-site inventory, NIC allocation, RAM & storage allocation, and the site port maps.
Disaster Recovery Objectives, Ownership, RPO & RTO
Recovery mechanisms, technical ownership, implementation state, and objective gaps for network configuration, workloads, secrets, DNS, UniFi control, and whole-site failure.
Blast Radius & Recovery Matrix
How WAN, edge, switching, storage, control-plane, DNS, backup, and inter-site failures propagate, what remains available, and which recovery path applies.