Job Type: Full Time
Job Category: IT
Job Description
- Job Title: Data Engineer – Kafka / PySpark / Hadoop
- Location: Toronto, ON
- Work Model: Onsite
- Job Type: Full Time (FTE)
Job Description
We are seeking an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop to design, develop, and support scalable batch and real-time data pipelines. The ideal candidate will have strong experience working with large-scale distributed data processing environments and enterprise data integration solutions.
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines.
- Develop data processing applications using Python and PySpark/Apache Spark.
- Build and support Kafka-based data ingestion and streaming pipelines.
- Work with Hadoop and related technologies to process large volumes of data.
- Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
- Perform data validation, reconciliation, and quality checks.
- Troubleshoot pipeline failures, data discrepancies, and performance issues.
- Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
- Monitor data pipelines and resolve production issues.
- Collaborate with data architects, developers, analysts, and business teams.
- Participate in Agile development, testing, deployment, and production support activities.
Required Skills
- Strong hands-on experience with Python for data engineering and automation.
- Strong expertise in PySpark / Apache Spark.
- Hands-on experience with Apache Kafka for real-time data ingestion and streaming.
- Strong experience with the Hadoop ecosystem and distributed data processing.
- Strong SQL skills and experience working with large datasets.
- Experience developing and maintaining ETL/ELT data pipelines.
- Strong understanding of distributed computing and data processing concepts.
- Experience with data ingestion, transformation, cleansing, and integration.
- Strong troubleshooting and performance optimization skills.
Good to Have
- Hive
- Databricks
- AWS, Azure, or GCP
- Git and CI/CD
- Unix/Linux
- Airflow or Autosys
- Relational and NoSQL databases
Required Skills
Cloud Developer Data / Python Engineer DevOps Engineer IT Business Continuity Analyst Net Back Engineer Python / DevOps Engineer