The infrastructure atlas

Know the system beneath the system.

A practical, connected guide to Kubernetes, networking, GPU infrastructure, and the platform decisions around them.

An abstract connected cloud, cluster, storage and GPU infrastructure topology
Connected systems, one mental model
06Core domains
01Connected learning path
0Fluff layers

Start here

Follow the signal

Choose a domain, then move through the concepts that make it work in the real world.

Kubernetes

The cluster as a control system

Start with the control plane, then trace how desired state turns into running workloads.

Featured path

From manifest to running pod

  1. 01API server receives desired state
  2. 02Scheduler selects a node
  3. 03Kubelet reconciles the workload
  4. 04CNI gives the pod a network identity
Continue into networking

First command

kubectl
kubectl get pods -A

# What to inspect
STATUS     READY     NODE
Running    1/1       worker-02

Use a simple cluster view to build the habit of asking where a workload is and why it landed there.

Networking

Packets need a story too.

Understand the difference between pod connectivity, service routing, and the fabric underneath both.

CNI

Pod identity

How interfaces, IPs, and policy enter the cluster.

kube-proxy

Service routing

How a stable virtual IP reaches an eligible endpoint.

eBPF

Programmable path

When data plane work moves closer to the kernel.

VXLAN

Overlay reach

How networks cross boundaries while preserving isolation.

GPU infrastructure

Compute is a topology problem.

A GPU is not just a schedulable resource. Its locality to CPU, network, storage, and peer GPUs decides how useful it is.

01

Node topology CPU, memory, PCIe, and GPU placement.

02

GPU fabric NVLink, NVSwitch, and peer-to-peer bandwidth.

03

Cluster fabric InfiniBand and Ethernet paths between nodes.

AI platform

Make scarce compute usable.

The platform is where scheduling policy, user workflows, observability, and data management meet.

Tenancy

Fair access

Quotas and guardrails make demand visible before it becomes an incident.

Scheduling

Useful placement

Align resource requests with actual hardware topology and workload shape.

Operations

Clear ownership

Surface health, cost, and capacity without turning every user into an operator.

Infrastructure

The physical layer still has opinions.

Servers, storage, and data center networking are part of the application experience. Treat them as first-class architecture.

Storage latencyNetwork fabricServer profilesPower and coolingFailure domains

Labs

Turn the map into muscle memory.

Short, composable exercises are coming next: start local, inspect the system, deliberately break one thing, then repair it.

Share a lab idea

Search KubeAtlas