Job Title: Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)
Job Type: Full-Time
Location: Toronto, ON (Hybrid)
As an Intermediate Site Reliability Engineer, you will maintain, optimize, and ensure the production reliability of enterprise-level Azure and Databricks platforms. You will focus on platform availability, continuous monitoring, incident response, operational readiness, and platform support while collaborating closely with engineering, security, network, and data teams.
Platform Reliability & Support: Monitor and support production Azure and Databricks environments to ensure maximum availability, performance, and operational readiness.
Incident Response & On-Call: Respond to production incidents, participate in on-call support rotations, join incident bridges, and execute emergency changes when required.
Databricks & Data Services: Support Databricks workspaces, clusters, workflows, user access, and Unity Catalog governance (catalogs, schemas, storage credentials). Maintain integrations with ADLS Gen2, Azure Data Factory, Azure SQL, and Key Vault.
Infrastructure & Networking: Troubleshoot Azure infrastructure, including storage accounts, Blob Storage, VNets, NSGs, private endpoints, DNS, and hub-and-spoke connectivity.
Observability & Monitoring: Manage alerts, dashboards, and platform health using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.
Root Cause & Maintenance: Perform root cause analysis (RCA), problem management, system patching, upgrades, maintenance, and disaster recovery exercises.
Operations & Documentation: Maintain operational runbooks and knowledge articles while tracking tickets and tasks in JIRA and ServiceNow.
Experience: 3+ years supporting Azure production cloud infrastructure and 1+ years supporting Databricks environments.
OS Administration: 1+ years of Windows Server administration and 1+ years of Linux administration.
Data & Storage: Proven experience supporting Azure Storage services, including ADLS Gen2 and Blob Storage.
Networking & Security: Understanding of VNets, NSGs, private endpoints, DNS, routing, Entra ID (Azure AD), RBAC, managed identities, and Azure Key Vault.
Monitoring Tools: Hands-on experience with Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.
ITSM & Operations: Hands-on experience with incident escalation, change management, RCA, JIRA, ServiceNow, and operational runbooks.
Operational support experience with Azure SQL and Azure Data Factory (integration runtimes, linked services, orchestration).
Understanding of Unity Catalog governance and Disaster Recovery/Business Continuity (RTO/RPO).
Exposure to AI/GenAI platforms, Azure OpenAI, MLOps, model endpoints, or RAG services.
Experience in cost monitoring, capacity planning, and platform health reporting.
Strong analytical and calm problem-solving mindset during critical production incidents.
Excellent cross-team collaboration skills (working with network, security, and platform teams).
Strong documentation skills and customer-focused approach to platform reliability.