خودرو45
خودرو45

Data Platform Engineer

Tehran/ Mirdamad
Full Time
Saturday to Wednesday
-
-
201 - 500 employees
Internet Provider / E-commerce / Online Services
Iranian company dealing only with Iranian entities
1397
Privately held
توضیحات بیشتر

key Requirements

4 years experience in similar position
GIT - Intermediate
Kubernetes - Intermediate

Job Description

About the Role

We’re building an internal data and AI platform on top of managed Kubernetes. You’ll own the data infrastructure layer from ingestion and storage to transformation and serving — with a strong focus on security, separation between production and analytical workloads, and building something the team can maintain long-term.

You’ll work closely with data engineers and ML engineers, providing the platform they run their work on.

What You’ll Do

  • Design and operate data pipelines: batch, micro-batch, and streaming flows
  • Build and maintain a data lakehouse: object storage, table formats (Iceberg/Delta Lake), and query engines
  • Set up and operate data orchestration: scheduling, dependency management, retries, and alerting on failures
  • Implement clear separation between production databases and BI/AI/analytical workloads (read replicas, CDC, staging layers)
  • Own data access control: who can read what, at what layer, with what credentials
  • Deploy and maintain MLOps tooling: experiment tracking, model registry, and notebook environments
  • Build observability into pipelines: data quality checks, pipeline SLAs, and failure alerting
  • Manage secrets and credentials across data systems securely
  • Write and maintain documentation: runbooks, data flow diagrams, and onboarding guides

What We’re Looking For

Must Have

  • Hands-on experience building and operating data pipelines (batch and/or streaming)
  • Experience with at least one orchestration tool: Airflow, Prefect, or Dagster
  • Familiarity with data lakehouse concepts: object storage (MinIO or S3-compatible), table formats (Iceberg or Delta Lake)
  • Experience with CDC or replication patterns to isolate production from analytical workloads
  • Understanding of data-level access control and RBAC across storage and query layers
  • Comfortable working on Kubernetes: deploying workloads, managing configs, basic troubleshooting
  • Secrets management experience (Vault, Sealed Secrets, or similar)

Good to Have

  • Experience with OLAP engines: ClickHouse, Trino, or Spark
  • Kafka or Redpanda for streaming ingestion
  • MLflow or similar for experiment tracking and model registry
  • JupyterHub or multi-tenant notebook environments
  • Data quality frameworks (Great Expectations, Soda, or similar)
  • Prometheus + Grafana for pipeline and infrastructure observability

Mindset We Value

  • You think about data flows end-to-end, not just individual components
  • You treat production data as sensitive by default and design access boundaries accordingly
  • You document pipelines and data flows so others can understand and maintain them
  • You’re comfortable owning trade-offs between simplicity and capability
  • You work with data and ML engineers as an enabler

Our Stack

  • Managed Kubernetes on Hamravesh and Sotoon
  • Self-hosted data tooling (we build what we can’t buy as managed services)
  • GitOps-driven deployments (handled by the team; you’ll consume, not own)

Why This Role

  • Real ownership over the data platform layer
  • Greenfield environment; you’ll design the architecture from the start
  • Small, focused team with high trust

Job Requirements

Age
20 - 40 Years Old
Gender
Men / Women
Software
Kubernetes| Intermediate GIT| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای خودرو45