We are looking for a Data Architect to design, build, and operate scalable data pipelines and data products, enabling a modern data mesh architecture. The role focuses on integrating legacy systems with cloud-native platforms to support analytics, reporting, and domain-owned data products.
Primary Skills
- Databricks, Pyspark, Data Architecture (Data Modeling), Datalake Architecture, Good knowledge on real time and batch time data processing.
- Experience in integrating legacy data sources such as DB2 (mainframe) and SQL Server.
- Good to have β Kafka knowledge.
Key Responsibilities
- Build and maintain batch and streaming data pipelines from legacy mainframe DB2 and on prem SQL Server systems into an AWS-based data lake.
- Implement real-time ingestion using Confluent Kafka and land data in Amazon S3.
- Develop data transformations and analytics pipelines in Databricks using the medallion architecture (Bronze, Silver, Gold).
- Design and deliver high-quality, reusable data products aligned to data mesh principles.
- Ensure data quality, reliability, security, and governance across pipelines.
- Collaborate with domain teams, platform teams, and stakeholders to enable self-serve analytics and consumption.
Good to have
- Experience in banking or financial services data platforms.
- Familiarity with data governance, metadata management, and regulatory reporting
#J-18808-Ljbffr