We are looking for a talented and experienced Mid-Level Data Engineer with a passion for data and cloud technologies to join our team. In this role, you will be a key player in designing, developing, and maintaining our AWS and Databricks-based data platform. You will bridge the gap between robust Data Engineering, Database Administration (DBA), and high-performance Analytics Engineering.
The ideal candidate understands the Lakehouse architecture, has experience managing modern cloud data warehouses like Snowflake, and possesses the ability to build scalable ETL pipelines that drive data-driven business decisions.
Key Responsibilities
End-to-End Pipelines: Develop and maintain complex ETL/ELT pipelines using Python/PySpark on the Databricks platform, as well as loading and transforming data within Snowflake.
Data Architecture: Implement and manage data layers following the Medallion architecture (Bronze, Silver, Gold) using Delta Lake.
Optimization & Analytics: Set up and manage Databricks SQL Warehouses, perform query optimization, and utilize internal visualizations for rapid data exploration.
Cloud Infrastructure: Leverage AWS cloud services to manage data storage, compute resources, and secure data movement.
DBA & Performance Tuning: Act as a custodian for our data platform-managing indexing, clustering, vacuuming, and performance tuning across both Delta Lake and relational/warehouse environments to ensure cost-efficiency and high speed.
Data Modeling: Design data models (Star Schema / Snowflake) in the Gold layer to ensure optimal performance for BI tools (e.g., Power BI, Tableau).
Data Governance: Manage metadata, permissions, and lineage using Unity Catalog and platform-specific access controls.
Quality Control: Implement automated Data Quality tests as an integral part of the CI/CD data pipelines.
Requirements: 5+ years of experience as a Data Engineer or Data Engineer/DBA - Must.
Strong hands-on experience with the AWS ecosystem (S3, IAM, EC2, etc.) and modern cloud data warehouses, specifically Snowflake.
At least 1 year of intensive, hands-on experience with the Databricks platform (including Notebooks and Workflows).
High proficiency in PySpark (or Spark Scala).
Expertise in writing complex SQL (Window Functions, CTEs) alongside DBA-level performance tuning (query profiling, indexing, partition pruning).
Technical Knowledge: Practical experience with Delta Lake, file formats (Parquet), and working with Cloud Storage (AWS S3 / Azure ADLS).
Modeling: Proven experience in designing Fact and Dimension tables.
This position is open to all candidates.