Architecture
Plant devicesIngestion APIPostgreSQL 17 HAAirflow + PySparkDashboards & AI
The problem
Plant telemetry had outgrown a single database server. The platform needed high availability, tested backups and an analytics layer — and a planned move from Azure to AWS without losing data.
Approach
- Designed a PostgreSQL 17 HA cluster with Patroni, etcd and HAProxy spanning Azure and AWS.
- Set up pgBackRest backups to S3 with lifecycle tiering to Glacier for long-term retention.
- Planned and executed the full migration of the platform from Azure to AWS EC2.
- Built an analytics pipeline on Apache Airflow 2.9 and PySpark 3.5 for batch processing.
Outcome
- Automatic leader election and failover — no single database server to lose.
- Point-in-time recovery from object storage, with cold data tiered to cut storage cost.