Experience: 4+ years in platform engineering, DevOps, or infrastructure engineering, with at least 1-2 years supporting AI/ML workloads.
Containerization & Orchestration: Advanced knowledge of Kubernetes (AKS, EKS, or GKE) and Docker for managing containerized AI services.
Infrastructure as Code (IaC): Deep expertise with tools like Terraform, OpenTofu, or Pulumi to automate environment setup.
CI/CD & MLOps Tools: Hands-on experience with GitLab CI, GitHub Actions, ArgoCD, and specialized platforms like Kubeflow or MLflow.
Cloud & GPU Provisioning: Experience managing cloud services (AWS, Azure, or GCP) with an emphasis on provisioning GPU instances and handling cluster autoscaling.
Data & Storage Systems: Experience setting up or maintaining vector databases (e.g., Pinecone, Milvus, Qdrant) and distributed storage solutions.