Lead Data Engineer
Lead Data Engineer
Key Responsibilities
- Design and solution end-to-end data architecture for data hubs/data products, all the way from source systems to consumption.
- Oversee the design and management of data solutions to ensure data is stored, processed, curated, and utilized effectively.
- Own and build a reusable data pipeline utilizing Big Data and Azure/Databricks for Eligibility Program data products.
- Ingesting huge volumes of data from various platforms for Analytics needs and writing high-performance, reliable, and maintainable ETL code.
- Leadership: Lead and mentor a team of data engineers, ensuring the efficient flow of data within the organization with the defined processes and tools.
- Collect, store, process, and analyze large datasets to build and implement extract, transfer, load (ETL) processes.
- Develop reusable frameworks to reduce the development effort involved, thereby ensuring cost savings for the projects.
- Utilizing CI/CD Pipelines: Utilize and enhance CI/CD practices to automate the delivery of data solutions, ensuring reliability and scalability based on the defined tools.
- Utilize Cloud technologies (Azure Databricks) to enable data product solutions.
- Develop quality code through performance optimizations in place right at the development stage.
- Appetite to learn new technologies and be ready to work on new cutting-edge cloud technologies.
- Partner with Tech, Business, BI, and Data Science teams to create reusable data products.
- Work with a team spread across the globe in driving the delivery of projects and recommend development and performance improvements.
- Track and report on KPIs for solution delivery and data quality.
- Communicate and present use cases, solutions, and impact to business stakeholders and mid/senior management.
- Optimize reusable frameworks, Spark jobs for performance and cost efficiency in large-scale environments.
- Ability to interact with business stakeholders in getting the requirements and implementing solutions.
- Analyze the data in depth, using SQLs and other exploratory tools against various platforms such as Bigdata, Oracle, SQL Server, Databricks and others.
- Work with IT, business, and architects to develop and design requirements to formulate technical design.
Required Skills
Essential Business Experience and Technical Skills:
- 8+ years of Data Solutions, development, and delivery experience with 4+ years of recent experience in Azure/Databricks environments.
- Proficiency and extensive experience with SQL, Spark &/or Scala/Python and performance tuning.
- Hands-on expertise in: Big Data (ex: Hive and HBase), Azure Databricks, Azure Functions, Cosmos DB and/or Data Factory experience is a MUST.
- Strong experience in building/designing Data warehouses, data stores for analytics consumption on Cloud (real time as well as batch use cases).
- Design, build, and deploy robust data ingestion and curation pipelines utilizing cloud-based data platforms such as Azure Data Factory, Apache Spark (Scala or Python), Azure Databricks, and Delta Lake.
- Good scripting experience, primarily on shell/bash/ PowerShell.
- Strong SQL knowledge and data analysis skills for data anomaly detection and data quality assurance.
- Experience and familiarity implementing data governance and data quality using the enterprise toolset.
- Skilled in drafting functional, technical requirements, creating high-level design documents, and data-flow diagrams, etc.
- Expertise in writing validation scripts to validate the data, data integrations and ETL transformations.
- Very good problem solver and excellent communication skills - both written and verbal.