Job Type: Full Time
Job Category: IT

Job Description

Job Title: Platform Engineer - AI/ML Infrastructure
Job Type: Full Time
Location: Toronto, ON
Work Model: 4 Days/Week

Primary Skills: Kubernetes, GitOps (Flux CD/ArgoCD), Helm & HashiCorp Vault

Job Description

We are seeking a skilled Platform Engineer - AI/ML Infrastructure to deploy, manage, and support reliable, scalable, and secure AI/ML infrastructure across multiple environments. The ideal candidate will have strong hands-on experience with Kubernetes, GitOps practices, CI/CD pipelines, secrets management, and cloud-native infrastructure.

The successful candidate will work with AI/ML platforms such as LlamaIndex Cloud and KDB.AI, supporting deployments across Development, QA, and Production environments while ensuring platform reliability, security, and operational efficiency.

Roles & Responsibilities

  • Deploy and manage LlamaIndex Cloud, KDB.AI, and similar AI/ML applications across Dev, QA, and Production environments.
  • Build reliable, scalable, and secure deployment pipelines using modern GitOps practices.
  • Implement and maintain GitOps workflows using Flux CD for automated deployments.
  • Administer and maintain Kubernetes clusters across multiple environments.
  • Configure HashiCorp Vault for secrets management and integrate with ExternalSecrets.
  • Maintain CI/CD pipelines using GitHub Actions, Jenkins, or similar tools.
  • Work with the enterprise Identity Management team to configure OIDC authentication with Microsoft Entra ID.
  • Create, maintain, and optimize Helm charts and Kubernetes manifests.
  • Configure and manage Kubernetes networking, Ingress, and Gateway API.
  • Monitor application performance and troubleshoot production issues.
  • Support PostgreSQL, MongoDB, Redis, and RabbitMQ infrastructure as required.
  • Implement and maintain database high-availability and failover configurations.
  • Develop and maintain infrastructure documentation, runbooks, and operational procedures.
  • Collaborate with engineering, security, identity, and platform teams to improve infrastructure reliability and automation.

Required Skills & Qualifications

  • 3+ years of experience managing production Kubernetes clusters.
  • 2+ years of experience with Flux CD, ArgoCD, or similar GitOps tools.
  • Advanced experience developing and managing Helm charts.
  • Strong hands-on experience with HashiCorp Vault for secrets management.
  • Experience with Artifactory or similar container registries.
  • Strong experience with CI/CD tools such as GitHub Actions or Jenkins.
  • Administration experience with PostgreSQL, MongoDB, Redis, and RabbitMQ.
  • Experience with database HA and failover configurations, including PgBouncer and HAProxy.
  • Strong Linux/Unix administration skills and shell scripting using Bash or PowerShell.
  • Good understanding of Kubernetes networking, Ingress, and Gateway API.
  • Strong troubleshooting, monitoring, and problem-solving skills.

Nice-to-Have Skills

  • Experience with LlamaIndex, LangChain, or other AI/ML platforms.
  • Knowledge of vector databases or KDB.AI.
  • Experience with Temporal.io workflow orchestration.
  • Programming experience with Python or Go for infrastructure automation.
  • Experience supporting production AI/ML workloads.

Key Competencies

  • Kubernetes Administration
  • GitOps & Continuous Deployment
  • Flux CD / ArgoCD
  • Helm
  • HashiCorp Vault
  • CI/CD & DevOps
  • AI/ML Infrastructure
  • Linux/Unix
  • Infrastructure Automation
  • Production Support

Required Skills
Automation Anywhere Cloud Developer Data Governance DevOps Engineer Engineering Architect Recruitment Salesforce Commerce Cloud Consultant Terraform Workato

Fill below details & click “Apply”

Only add 10 digit number without prefix
Resume can be attached in PDF, JPG, Word , Txt format only

Share This Job