Identity and access redesignAudited how 14 of our services handle tenants, then designed a new model for clients that run several brands: separate staff and client logins, a company, brand and sub-brand hierarchy, and row-level access in Trino through OPA. Now leading the first phase, moving Keycloak from v23 to v26.
A Postgres table that got too bigMoved a slow log table to monthly partitions with BRIN, GIN and partial indexes. When it reached 27M rows and 89 GB, I added a job that archives old months to Parquet on S3, which freed about 59 GB without a single DELETE.
Bigger file uploadsClients with sources we don't integrate upload their own reports and choose to upsert, delete or append. Moving uploads to presigned S3 URLs and background workers took the limit from 200 MB to 5 GB+.
Airbyte off EKSMoved Airbyte from EKS to a single EC2 instance. Same syncs, about $12,000 a year cheaper.
Replication lagFixed lag in our Postgres to Redshift replication by tuning the AWS DMS batch and LOB settings.
Failures nobody sawA secret rotation once broke Amazon token refresh for hours before anyone noticed, and empty Blinkit reports were being saved as if they were fine. Both now alert and retry.
Guard rails for the integrationsCircuit breakers with deduplicated Slack alerts, account quarantine after repeated OTP failures, and a Redis lock so two workers never open overlapping portal sessions.
Data that changes laterSyncs re-read a rolling window instead of only fetching new rows, because returns and tax adjustments keep changing orders after the fact.
Hosted dashboardsDashboards built from Lens are served from snapshots that refresh in the background, so nobody sits waiting on a cold query.
Deploys and accessNode services run in PM2 cluster mode with graceful shutdown, so a deploy never cuts off a running job. Internal access goes through Tailscale.