Distributed & Large-Scale Systems

I engineer systems that scale.

I study, measure, build, operate, and improve distributed and large-scale computing systems — from the behavior of a single Linux machine to systems spanning machines, networks, data, and failure.

Engineering Responsibility

Measure first. Understand deeply. Improve continuously.

My work is centered on one responsibility: understand why systems behave the way they do, measure that behavior, and improve the system.

Systems

From one machine to large-scale systems.

I approach systems from the bottom up: hardware and operating systems, processes and networks, services and databases, distributed components, and eventually large-scale platforms.

Linux Systems

Processes, memory, CPU, storage, scheduling and system calls.

Networking

Sockets, TCP/IP, latency, throughput, congestion and failure.

Concurrency

Threads, synchronization, contention and parallel execution.

Distributed Systems

Communication, replication, consistency, consensus and fault tolerance.

Data Systems

Storage, databases, partitioning, streaming and data movement.

Large-Scale Systems

Scalability, availability, reliability, observability and cost.

Cloud & Infrastructure

Containers, orchestration, deployment and resilient infrastructure.

Research Systems

Experimentation, measurement, modeling, implementation and analysis.

Method

The engineering loop.

Theory tells me what may happen. Code lets me create the conditions. Measurement tells me what actually happened.

OBSERVE
→
MEASURE
→
ASK WHY
→
MODEL
→
CODE
→
EXPERIMENT
→
EXPLAIN
→
IMPROVE
↺

Real-World Observations

Systems are understood through evidence.

I use real machines, real workloads, measurements, controlled experiments, and failure analysis to turn system behavior into understanding.

My Linux Laboratory

sys — an Ubuntu Linux machine used to investigate CPU, memory, processes, networking, storage, performance and system behavior.

CPUMemoryProcessesLatencyThroughput

My Android Laboratory

My Android smartphone is another Linux-based laboratory for studying real-world system behavior. I use it to observe hardware, operating system, networking, memory, processes, security, and runtime characteristics through direct measurement.

Engineering Stack

Tools for understanding and building systems.

CC++LinuxTCP/IPSocketsGitPythonSQLPostgreSQLMySQLDockerKubernetesKafkaSparkHadoopDistributed SystemsCloud Computing

Current Work

Building capability, one system at a time.

My current work combines systems experimentation, Linux and C, networking, distributed-systems fundamentals, implementation, measurement, and research-oriented learning.

Measure

Establish quantitative baselines for latency, throughput, utilization, concurrency, errors, resource consumption and reliability.

Build

Implement systems from the ground up, beginning with Linux and networking and progressing toward distributed components and data platforms.

Improve

Stress systems, expose bottlenecks and failure modes, explain observed behavior, redesign, and measure again.

Research

Move from implementation and measurement toward models, papers, experiments and new system designs.

Distributed & Large-Scale Systems

Understand the system.
Measure the system.
Improve the system.

This website documents that engineering practice — the systems I study, experiments I run, software I build, and what those systems teach me.

Site Statistics

Site Views —
Unique Visitors —