Job Type: Full Time
Job Category: IT

Job Description

Job Title: Senior Platform Engineer

Location:  Irvine, CA (Onsite)
Fulltime

 

Job Description

Senior Platform Engineer

 

We are seeking a Senior Platform Engineer to build, administer, automate, secure, and operate enterprise Databricks and cloud data-platform environments. The role is responsible for Databricks workspace administration, Unity Catalog governance, identity and access management, compute and cluster policies, infrastructure automation, CI/CD enablement, monitoring, reliability, cost management, and production support.

The ideal candidate combines hands-on Databricks administration with strong AWS or Azure platform engineering, Terraform, Python, Kubernetes, CI/CD, security, networking, observability, and Site Reliability Engineering practices.

 

Key Responsibilities

Databricks Administration and Platform Engineering

•              Administer Databricks accounts, workspaces, metastores, catalogs, schemas, external locations, storage credentials, connections, shares, recipients, and platform configurations.

•              Provision and manage development, test, staging, and production workspaces using standardized, repeatable patterns.

•              Configure workspace settings, repositories, jobs, notebooks, SQL warehouses, instance pools, job clusters, all-purpose compute, and serverless capabilities.

•              Define and enforce cluster policies, approved runtime versions, libraries, init scripts, autoscaling, tagging, and compute-usage standards.

•              Manage Databricks Runtime and platform upgrades, compatibility testing, release planning, maintenance windows, and rollback procedures.

•              Support Databricks Workflows, Delta Live Tables or Lakeflow Declarative Pipelines, Databricks SQL, MLflow, model registry, feature engineering, and structured-streaming services.

•              Troubleshoot workspace, permissions, connectivity, compute, storage, job execution, library, runtime, and performance issues.

•              Maintain administration standards, platform documentation, knowledge articles, support procedures, and operational runbooks.

 

Unity Catalog, Identity, Security, and Governance

•              Design and manage Unity Catalog metastores, catalogs, schemas, managed and external tables, volumes, external locations, storage credentials, and grants.

•              Implement user and group provisioning through single sign-on, SCIM, identity-provider integration, and enterprise directory services.

•              Administer account-level and workspace-level users, groups, service principals, permissions, entitlements, and access-control models.

•              Implement role-based and attribute-based access controls, least-privilege permissions, separation of duties, and privileged-access procedures.

•              Configure secure access patterns for secrets, tokens, credentials, service principals, private endpoints, storage, and external systems.

•              Enable audit logging, lineage, system tables, tagging, data classification, row-level security, column masking, and compliance reporting.

•              Partner with security and governance teams on encrypti on, key management, network controls, data loss prevention, retention, auditability, and regulatory requirements.

•              Review access, monitor privileged activities, remediate policy violations, and support internal and external audits.

 

Cloud Infrastructure and Networking

•              Build and operate Databricks on AWS or Azure, including secure integration with cloud storage, identity, networking, encryption, and monitoring services.

•              On AWS, work with S3, IAM, KMS, VPC, PrivateLink, security groups, Route 53, CloudWatch, Secrets Manager, and related services.

•              On Azure, work with ADLS Gen2, Microsoft Entra ID, managed identities, Key Vault, virtual networks, private endpoints, network security groups, Azure Monitor, and related services.

•              Configure control-plane and data-plane connectivity, private networking, DNS, routing, firewall, proxy, and egress controls.

•              Integrate Databricks with cloud data lakes, APIs, databases, message platforms, and enterprise applications.

•              Support Kubernetes, Docker, EKS or AKS, and container-based platform services where required.

•              Contribute to capacity planning, disaster recovery, high availability, backup, restoration, and business-continuity exercises.

 

Infrastructure as Code and Automation

•              Build and maintain reusable Terraform modules using cloud and Databricks providers.

•              Automate workspace, network, Unity Catalog, storage, identity, compute-policy, cluster, job, permission, and monitoring configurations.

•              Manage Terraform state, workspaces, variables, modules, versioning, policy checks, drift detection, and controlled promotion across environments.

•              Use Python, shell scripting, Databricks CLI, REST APIs, and SDKs to automate administrative and operational tasks.

•              Implement self-service workspace, catalog, schema, access, and compute vending with appropriate approval and governance controls.

•              Maintain configuration standards and reduce manual administration through repeatable automation.

 

DevOps, CI/CD, and Release Engineering

•              Design and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or Harness.

•              Automate deployment of notebooks, jobs, workflows, libraries, policies, infrastructure, and platform configuration.

•              Support Databricks Asset Bundles, Git integration, artifact management, environment promotion, testing, approvals, and rollback.

•              Integrate security scanning, policy validation, infrastructure testing, and release evidence into delivery pipelines.

•              Enable engineering teams through templates, reusable pipelines, documentation, and self-service platform capabilities.

•              Partner with application, data-engineering, and DevOps teams to ensure deploy ment standards are consistent and supportable.

 

Reliability, Monitoring, Operations, and FinOps

•              Establish monitoring, alerting, dashboards, logs, metrics, traces, and health checks for Databricks and connected cloud services.

•              Use CloudWatch, Azure Monitor, Datadog, Splunk, New Relic, or similar platforms to monitor availability, compute utilization, failures, security events, and cost.

•              Define platform service-level indicators, service-level objectives, operational metrics, and error budgets.

•              Lead incident response, problem management, root-cause analysis, corrective actions, and post-incident reviews.

•              Manage vulnerability remediation, runtime patching, dependency updates, security exceptions, and platform lifecycle activities.

•              Optimize cluster sizing, autoscaling, pools, SQL warehouses, job concurrency, serverless usage, storage, and workload scheduling.

•              Implement budget controls, chargeback or showback tagging, utilization reporting, anomaly detection, and cost-optimization recommendations.

•              Participate in operational support rotations and maintain escalation paths with Databricks and cloud providers.

•              Coordinate platform upgrades, disaster-recovery tests, security reviews, and production-readiness assessments.

 

Collaboration and Technical Leadership

•              Partner with architecture, data engineering, security, cloud, network, governance, FinOps, and service-management teams.

•              Advise engineering teams on Databricks platform standards, secure patterns, deployment models, performance, and cost.

•              Conduct technical reviews and ensure solutions meet enterprise architecture and operational-support requirements.

•              Mentor platform engineers and administrators and lead knowledge-transfer sessions.

•              Communicate platform health, risks, dependencies, incidents, and improvement roadmaps to technical and business stakeholders.

•              Drive continuous improvement in automation, reliability, security, developer experience, and operational efficiency.

 

Required Qualifications

•              Typically, 7–10 years of cloud, DevOps, Site Reliability Engineering, infrastructure, or platform-engineering experience.

•              At least 3 years of hands-on Databricks platform administration in an enterprise environment.

•              Strong experience administering Databricks workspaces, Unity Catalog, compute, cluster policies, jobs, SQL warehouses, permissions, and service principals.

•              Strong experience with AWS or Azure infrastructure, identity, storage, networking, encryption, monitoring, and private connectivity.

•              Strong proficiency with Terraform and infrastructure-as-code practices.

•              Experience with Python, shell scrip

Required Skills
DevOps Engineer

Fill below details & click “Apply”

Only add 10 digit number without prefix
Resume can be attached in PDF, JPG, Word , Txt format only

Share This Job