Job Type: Full Time
Job Category: IT

Job Description

  • Job Title: Data Engineer – Kafka / PySpark / Hadoop
  • Location: Toronto, ON
  • Work Model: Onsite
  • Job Type: Full Time (FTE)
  • Job Description

    We are seeking an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop to design, develop, and support scalable batch and real-time data pipelines. The ideal candidate will have strong experience working with large-scale distributed data processing environments and enterprise data integration solutions.

    Key Responsibilities

    • Design, develop, and maintain scalable batch and real-time data pipelines.
    • Develop data processing applications using Python and PySpark/Apache Spark.
    • Build and support Kafka-based data ingestion and streaming pipelines.
    • Work with Hadoop and related technologies to process large volumes of data.
    • Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
    • Perform data validation, reconciliation, and quality checks.
    • Troubleshoot pipeline failures, data discrepancies, and performance issues.
    • Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
    • Monitor data pipelines and resolve production issues.
    • Collaborate with data architects, developers, analysts, and business teams.
    • Participate in Agile development, testing, deployment, and production support activities.

    Required Skills

    • Strong hands-on experience with Python for data engineering and automation.
    • Strong expertise in PySpark / Apache Spark.
    • Hands-on experience with Apache Kafka for real-time data ingestion and streaming.
    • Strong experience with the Hadoop ecosystem and distributed data processing.
    • Strong SQL skills and experience working with large datasets.
    • Experience developing and maintaining ETL/ELT data pipelines.
    • Strong understanding of distributed computing and data processing concepts.
    • Experience with data ingestion, transformation, cleansing, and integration.
    • Strong troubleshooting and performance optimization skills.

    Good to Have

    • Hive
    • Databricks
    • AWS, Azure, or GCP
    • Git and CI/CD
    • Unix/Linux
    • Airflow or Autosys
    • Relational and NoSQL databases

Required Skills
Cloud Developer Data / Python Engineer DevOps Engineer IT Business Continuity Analyst Net Back Engineer Python / DevOps Engineer

Fill below details & click “Apply”

Only add 10 digit number without prefix
Resume can be attached in PDF, JPG, Word , Txt format only

Share This Job