Senior Data Engineer
RainTech
Location
Onsite
Employment Type
Full-time
About This Role
- Design and build streaming pipelines on Kafka: topic design, partitioning, schema management, and processing services in Python (or Go)
- Build batch ingestion and transformation pipelines using an orchestrator (Airflow, Dagster, or equivalent) running on Kubernetes
- Design and optimize data models in ClickHouse, including time-series and location data: partitioning, sort keys, materialized views, and query tuning for high-volume dashboard queries
- Replicate operational data from PostgreSQL into the analytical layer using change data capture (CDC)
- Ingest results from AI models (such as object detections, OCR extractions, and match results) as structured, queryable data linked to their sources
- Define data contracts and query patterns with the backend and fullstack engineers who build on top of your data
- Implement data quality checks and pipeline monitoring (metrics, dashboards, alerting). You own the health of the pipelines you ship
- Implement access control at the data layer, aligned with the platform's Keycloak roles, together with the backend and fullstack engineers
- Containerize and deploy pipelines with the DevOps engineer, using Docker, Helm, GitLab CI/CD, and ArgoCD
What We're Looking For
- 5+ years in data engineering, with at least 2 years building streaming pipelines in production
- Production experience with Kafka as a builder, not just a consumer: topic design, schema evolution, and delivery guarantees
- Has designed and tuned at least one analytical database in production (ClickHouse, Trino/Presto, or equivalent)
- Strong Python and SQL
- Has run data pipelines on self-hosted or on-premise infrastructure, not only on managed cloud data services
- Has built and operated pipelines with an orchestrator (Airflow, Dagster, or equivalent) in production
Preferred:
- ClickHouse specifically
- CDC pipelines (Debezium or equivalent) from PostgreSQL
- Deploying services on Kubernetes with Helm and GitLab CI
- Experience with geospatial or time-series data
- Exposure to handling outputs from AI/ML models