Agentic AI/ML · Quantitative Finance

Systems for models, markets, and scale.

I'm Zhanyl Abdybaeva, an ML infrastructure and quantitative finance engineer working at a quantitative trading firm. My work spans GPU clusters, inference systems, distributed training, and the schedulers that decide what runs when.

This site is where I publish long-form technical writing, shorter field notes, and project pages for the systems I'm building.

Role

Infrastructure engineering at a quantitative trading firm

Focus

GPU clusters · inference · Slurm at scale

Studying

MS CS — ML @ Georgia Tech
CQF

Writing

Deep dives & lab notes
most weeks

Four bands. Across the top, a measurement layer — slurm-rca-bench and
              k8s-gpu-scheduler-lab — scores both control planes beneath it. In the middle,
              two columns side by side: the Slurm control plane with slurm-scheduler-lab,
              cluster-sre-agent with slurm-mcp, and epilog-gpu-validator; the Kubernetes
              control plane with slinky-gitops and k8s-gpu-scheduler-lab covering Kueue,
              Volcano and NVIDIA KAI. Both allocate from a shared GPU fleet along the
              bottom, held by gpu-reaper and ib-slurm-exporter. The tools change with the
              substrate; the fleet and the measurement layer do not.

Last published July 26, 2026

ls ./writing --featured

All deep dives →

tail ./lab-notes

All lab notes →