Projects
Projects
Three portfolio projects at the intersection of HPC inference, distributed training, and quantitative finance. Permanent reference pages — updated as milestones ship.
- Epilog GPU Validator
Drain a node for a persistently faulty GPU — and never for a transient one.
- Slinky GitOps
Slurm on Kubernetes from nothing in one command — and the auth-key rotation nobody wants to test in production.
- IB Slurm Exporter
Correlate InfiniBand and RoCE counters with the Slurm job that owns them, so a slow collective can be traced to the fabric.
- GPU Reaper
Find wasted GPU allocations on a Slurm cluster and escalate them through alert, drain, and cancel.
- Slurm Scheduler Lab
Test Slurm priority and backfill policy against a job trace before it reaches a live controller.
- Agentic ML Inference Infrastructure
Production agentic AI infrastructure layer: autonomous systems that reason, decide, and act on real workloads using vLLM, LangGraph, and MCP.
- 128-GPU Distributed Training Platform
Distributed ML training infrastructure for large-scale models: FSDP, gradient checkpointing, mixed precision, and cost-optimized multi-node orchestration.
- GPU-Accelerated Volatility Surface Calibration
CUDA-accelerated local and stochastic volatility surface calibration — Dupire's model, Heston, and hybrid approaches for real-time derivatives pricing.