Main Responsibilities
? As part of the AI & Data team, design and build the data structures, schemas, and models
that AI Engineers depend on for training, feature engineering, and inference.
? Develop and orchestrate scalable data pipelines on Databricks (Spark, Delta Lake) and load
curated, analytics-ready data into Snowflake.
? Own data ingestion, transformation (ELT/ETL), and storage across the Azure cloud (e.g.
ADLS, Data Factory, Event Hubs/Synapse), including structured, semi-structured, and
unstructured data such as text, images, and video.
? Dock machine-learning and computer-vision models into the data pipelines, and design the
data flow that feeds and consumes those AI services.
? Sync curated lakehouse data into Lakebase (managed Postgres) for low-latency serving,
manage change-data-capture back into Delta tables, and support online feature stores and
agent state for AI Engineers.
? Build and maintain API integrations and automated data ingestion from internal systems
and external third-party sources.
? Monitor pipeline performance, reliability, and cost; troubleshoot failed jobs and optimize
Snowflake and Databricks workloads.
? Implement data quality, validation, and lineage, and document the data dictionary and ETL
processes.
? Partner with AI Engineers and stakeholders to translate model and business requirements
into extensions of the data platform.
Skills, Qualifications & Education
? Bachelorβs degree in Computer Science, Data Engineering, or a related field.
? At least 4 years of work experience in data engineering or a similar data-focused role.
? Hands-on production experience with Databricks (Apache Spark, Delta Lake, notebooks,
workflows).
? Hands-on production experience with Snowflake (data modeling, performance tuning,
access control, cost management).
? Solid experience with Microsoft Azure data services (e.g. ADLS, Data Factory, Event Hubs /
Synapse).
? Experience with PostgreSQL and OLTP databases; familiarity with Lakebase (Databricks
managed Postgres) is a strong plus.
? Working knowledge of JavaScript / TypeScript, used for data APIs, microservices, or app
facing integrations.
? Strong expertise in SQL (will be tested during the recruitment process).
? Robust Python literacy, especially for data handling and pipeline development.
? Comfortable working with both structured and unstructured data; does not shy away from
troubleshooting failed ETL processes or API integrations.
? Outstanding data-structure and data-modeling design skills.
? Working knowledge of machine-learning, NLP, or computer-vision workflows is a plus.
? Experience with dbt, Airflow, or Databricks Workflows, and with CI/CD and infrastructure
as-code, is a plus.
? Strong ability to translate ideas between technical and non-technical audiences.
? Curious, collaborative, self-motivated, and organized; able to run multiple projects against
tight deadlines.