About the Role
We're looking for a Data Engineer to build the data foundation our AI systems depend on — pipelines, feature stores, and lakehouse architecture that turn messy enterprise data into something models can actually be trained and served on.
What You'll Do
- Design and build batch and streaming data pipelines feeding AI/ML systems
- Build and maintain feature stores with correct point-in-time semantics for training and serving
- Own data quality — monitoring, validation, and alerting when upstream data drifts or breaks
- Work with ML engineers to understand feature requirements and translate them into reliable pipelines
- Optimize pipeline cost and performance as data volume scales
What We're Looking For
- 3+ years of data engineering experience with production batch and streaming pipelines
- Strong SQL and experience with Spark or an equivalent distributed processing framework
- Familiarity with streaming systems (Kafka or equivalent) and transformation tooling (dbt)
- Understanding of the specific data challenges ML systems introduce — training/serving skew, point-in-time correctness
Nice to Have
- Experience with a modern lakehouse platform (Databricks, Snowflake)
- Prior experience building or operating a feature store
- Exposure to BFSI, healthcare, or another regulated-data environment