Jobgether
Jobgether

Databricks Engineer - Senior/ Lead

engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Design, develop, and maintain scalable data pipelines and data engineering solutions using Azure Databricks.
    • Build robust ETL/ELT workflows using Python, PySpark, Spark SQL, and Apache Spark.
    • Develop complex SQL queries and transformations to process and prepare large datasets for analytics and machine learning.
    • Optimize Spark jobs and distributed processing workloads for performance, scalability, reliability, and cloud cost efficiency.
    • Work with structured and semi-structured data formats including Parquet, Delta, JSON, and CSV.
    • Design, build, and maintain Delta Lake tables, leveraging capabilities such as ACID transactions, time travel, and schema evolution.
    • Integrate Databricks workloads with Azure Data Lake Storage Gen2 and other Azure data services.
    • Apply distributed computing principles to efficiently process large-scale datasets.
    • Implement data quality, validation, monitoring, and error-handling processes across data pipelines.
    • Collaborate with data scientists, analysts, and business stakeholders to understand requirements and deliver reliable data solutions for analytics and ML use cases.
    • Apply Azure security, access control, governance, and RBAC best practices across data engineering environments.
    • Use Git-based version control and follow collaborative software development practices.
    • Contribute to CI/CD pipelines and automated deployment processes using tools such as Azure DevOps or GitHub Actions.
    • Support data modeling initiatives and help ensure data structures are optimized for downstream analytics and reporting.
    • Contribute to streaming data solutions using technologies such as Spark Structured Streaming, Azure Event Hubs, or Kafka.
    • Identify opportunities to improve pipeline architecture, engineering standards, performance, and operational efficiency.
    • Provide technical guidance and contribute to engineering best practices appropriate for a Senior or Lead-level position.
    • Requirements:

      • 6+ years of professional experience in Data Engineering or a closely related field.
      • Strong hands-on experience designing and developing solutions with Azure Databricks.
      • Advanced proficiency in Python for data processing, transformation, and pipeline development.
      • Strong SQL skills, including complex joins, window functions, data transformations, and query performance optimization.
      • Hands-on expertise with Apache Spark and PySpark, including distributed data processing and Spark job optimization.
      • Solid experience working with Delta Lake, including ACID transactions, time travel, and schema evolution.
      • Strong knowledge of Azure Data Lake Storage Gen2 and cloud-based data architecture.
      • Understanding of distributed computing concepts and the ability to design solutions for large-scale data workloads.
      • Experience using Git for version control and collaborative development.
      • Strong understanding of modern ETL/ELT patterns, data pipeline architecture, and data engineering best practices.
      • Experience with Azure Data Factory is preferred.
      • Exposure to CI/CD practices and tools such as Azure DevOps or GitHub Actions is an advantage.
      • Basic understanding of data modeling and its application to analytical workloads.
      • Familiarity with Azure security principles, RBAC, access control, and data governance.
      • Exposure to real-time and streaming data technologies such as Spark Structured Streaming, Event Hubs, or Kafka is a plus.
      • Strong analytical and problem-solving abilities, with a structured approach to debugging and performance optimization.
      • Strong communication and collaboration skills, with the ability to work effectively with technical teams, data scientists, analysts, and business stakeholders.
      • Ability to take ownership of complex engineering initiatives and provide technical direction at a senior or lead level.
      • Benefits:

        • Opportunity to work on large-scale data engineering and analytics initiatives within an AI and technology-focused environment.
        • Hands-on experience with Microsoft Azure, Azure Databricks, Apache Spark, Delta Lake, and modern cloud data platforms.
        • Opportunity to design and optimize production-grade data pipelines handling complex and large-scale datasets.
        • Exposure to advanced distributed data processing, cloud architecture, and modern ETL/ELT practices.
        • Opportunity to contribute to analytics and machine learning use cases in collaboration with data scientists and analysts.
        • Exposure to CI/CD, cloud security, data governance, and streaming data technologies.
        • Senior/Lead-level scope with opportunities to influence engineering standards and technical architecture.
        • Collaborative environment involving data engineering, analytics, data science, and business stakeholders.
        • Opportunities for continued technical growth across cloud data engineering, AI, and emerging data technologies.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now