Leela Kumar
Senior DevOps Engineer
I bring 8+ years of DevOps and platform engineering experience building and operating Kubernetes and cloud infrastructure at enterprise scale — designing AWS cloud abstractions, running multi-environment Kubernetes clusters, and leading CI/CD and GitOps modernization that cuts deployment time. Comfortable owning on-call and incident response, with hands-on experience building GPU-enabled Kubernetes infrastructure (NVIDIA GPU Operator, EKS, Triton, NIM) for AI/ML workloads.
A little about me
Systems work rewards patience and pattern recognition — so does most of what I do outside of it. I've been playing chess since I was a kid, still play in local table tennis and volleyball leagues, and I'm working toward training five days a week. None of it pays the on-call pager, but it keeps the problem-solving muscle in shape.
Certifications
Kubernetes, cloud, and NVIDIA AI infrastructure credentials, kept current. View all badges on Credly →
Projects
Repositories from my GitLab — infrastructure-as-code, GPU/AI platforms, and a couple of side builds. Click any card to open the repo.
Terraform-defined AWS foundation — VPC, networking, and the account baseline shared across environments.
Kubernetes manifests and Helm charts for deploying and managing application workloads.
An AI-assisted Terraform workflow experiment — natural-language prompts driving infrastructure-as-code changes.
Early-stage desktop app scaffolding for running and interacting with AI models locally.
AWS-based Slurm HPC cluster — CPU/GPU partitions, Munge auth, and NVIDIA Base Command Manager.
CUDA-accelerated GPU compute workloads and benchmarking.
TensorFlow inference service with GPU scheduling, node taints, and tolerations.
PyTorch inference service with CUDA acceleration on GPU-scheduled nodes.
GPU utilization monitoring using nvidia-smi dmon alongside Prometheus and Grafana.
NVIDIA Triton Inference Server platform serving containerized Llama 3 and custom models.
End-to-end AI chat application on Kubernetes with Istio ingress, autoscaling, and observability.
NVIDIA NIM microservices platform for containerized model inference.
GPU-enabled AWS EKS infrastructure via Terraform, Helm, and the NVIDIA GPU Operator.
GPU discovery and scheduling — Device Plugin, GPU Feature Discovery, and Container Toolkit.
GPU-backed inference microservice with autoscaling and health checks.
Let's Connect
Open to Senior/Staff Platform, DevOps, and GPU/AI Infrastructure roles — remote or Denver-area. Pick whichever's easiest.
Email Best for detailed questions lkdevops27@gmail.com LinkedIn Let's connect professionally View profileSend a message
This opens your email client with the message pre-filled.