Featured path
From manifest to running pod
- 01API server receives desired state
- 02Scheduler selects a node
- 03Kubelet reconciles the workload
- 04CNI gives the pod a network identity
The infrastructure atlas
A practical, connected guide to Kubernetes, networking, GPU infrastructure, and the platform decisions around them.
Start here
Choose a domain, then move through the concepts that make it work in the real world.
Control planes, workloads, services, storage, and the failure modes between them.
Explore cluster foundations → 02Packet paths from pod to fabric: CNI, kube-proxy, eBPF, and overlays.
Trace the packet → 03Scheduling, topology, NVLink, InfiniBand, and the realities of high-density compute.
Map accelerated compute → 04Multi-tenancy, workloads, data, and the platform layer that turns hardware into a service.
Design the platform → 05Storage, servers, fabrics, and the physical decisions that shape application behavior.
See the foundations → 06Small, repeatable environments for moving from a diagram to an operating intuition.
Build something →Kubernetes
Start with the control plane, then trace how desired state turns into running workloads.
Featured path
First command
kubectlkubectl get pods -A
# What to inspect
STATUS READY NODE
Running 1/1 worker-02
Use a simple cluster view to build the habit of asking where a workload is and why it landed there.
Networking
Understand the difference between pod connectivity, service routing, and the fabric underneath both.
How interfaces, IPs, and policy enter the cluster.
How a stable virtual IP reaches an eligible endpoint.
When data plane work moves closer to the kernel.
How networks cross boundaries while preserving isolation.
GPU infrastructure
A GPU is not just a schedulable resource. Its locality to CPU, network, storage, and peer GPUs decides how useful it is.
Node topology CPU, memory, PCIe, and GPU placement.
GPU fabric NVLink, NVSwitch, and peer-to-peer bandwidth.
Cluster fabric InfiniBand and Ethernet paths between nodes.
AI platform
The platform is where scheduling policy, user workflows, observability, and data management meet.
Quotas and guardrails make demand visible before it becomes an incident.
Align resource requests with actual hardware topology and workload shape.
Surface health, cost, and capacity without turning every user into an operator.
Infrastructure
Servers, storage, and data center networking are part of the application experience. Treat them as first-class architecture.
Labs
Short, composable exercises are coming next: start local, inspect the system, deliberately break one thing, then repair it.