Platform

Site Reliability Engineer

Keep our clients' banking and aviation platforms highly available — building observability, automation, and incident response for mission-critical systems.

Full-timeEmployment type
Bengaluru / RemoteLocation

Responsibilities

  • Define and track SLIs/SLOs for production banking and aviation workloads
  • Build automation to reduce toil and speed up incident response
  • Lead or support on-call rotations and post-incident reviews (blameless postmortems)
  • Improve observability with metrics, logging, and tracing tooling (Prometheus, Grafana, ELK)
  • Design and test disaster-recovery and failover procedures for critical systems
  • Partner with engineering teams on capacity planning and performance tuning
  • Drive continuous improvement of runbooks, alerting thresholds, and escalation paths

What we're looking for

  • 4+ years in SRE/DevOps roles supporting high-availability systems
  • Strong scripting skills (Python/Go) and Kubernetes experience
  • Experience with incident management, on-call practices, and postmortem culture
  • Familiarity with observability stacks (Prometheus, Grafana, ELK, or similar)
  • Understanding of capacity planning and performance tuning for critical workloads
  • Remote-friendly; occasional travel to Bengaluru office

Nice to have

  • Experience with chaos engineering or resilience testing
  • Exposure to banking or aviation production environments
  • Certified Kubernetes Administrator (CKA) or equivalent

Don't see the right fit?

Write to us at info@esika.in — we're always looking for banking and aviation IT talent.