Hpc
Hpc
- Epilog GPU Validator
Drain a node for a persistently faulty GPU — and never for a transient one.
- GPU Reaper
Find wasted GPU allocations on a Slurm cluster and escalate them through alert, drain, and cancel.
- Slurm Scheduler Lab
Test Slurm priority and backfill policy against a job trace before it reaches a live controller.
- I Was Planning a Slurm 25.11 Upgrade. Then 26.05 Shipped.
Three versions, two config renames that break on restart, and one cgroup change that quietly breaks every job-attribution tool you own.
- Your Slurm Priority Weights Matter Less Than Your Users' Time Limits
I built a simulator to tune Slurm priority weights. It kept telling me the weights were not the problem.
- 128-GPU Distributed Training Platform
Distributed ML training infrastructure for large-scale models: FSDP, gradient checkpointing, mixed precision, and cost-optimized multi-node orchestration.
- GPU-Accelerated Volatility Surface Calibration
CUDA-accelerated local and stochastic volatility surface calibration — Dupire's model, Heston, and hybrid approaches for real-time derivatives pricing.