Job description

Site Reliability Engineer job description

A site reliability engineer applies software engineering to operations: defining SLOs, spending error budgets deliberately, running incident response, and automating away the repetitive work that keeps production upright. In India the role sits with product engineering rather than IT, owning on-call rotations, blameless postmortems, observability and capacity planning for services that must not go dark.

Reviewed 26 July 2026 · one of 119 role templates · how OnJob writes and checks these

Also known as: SRE, Reliability Engineer, Production Engineer.

Experience 3–12 yrs Typical pay typically ₹12L–₹45L/yr 9 core skills

What does a Site Reliability Engineer do?

A site reliability engineer applies software engineering to operations: defining SLOs, spending error budgets deliberately, running incident response, and automating away the repetitive work that keeps production upright. In India the role sits with product engineering rather than IT, owning on-call rotations, blameless postmortems, observability and capacity planning for services that must not go dark.

A site reliability engineer usually has around 3–12 yrs of experience and earns typically ₹12L–₹45L/yr in India. The day-to-day blends SLOs & error budgets, Incident response, Prometheus & Grafana and more — this page gives you a ready-to-use site reliability engineer job description template you can copy, plus the exact skills and salary employers expect.

Use this template

Site Reliability Engineer job description template

Copy the 8 responsibilities, 6 requirements and 9 skills below into your job post, then edit the parts that are specific to your company — pay band, location and reporting line.

Select the text above to copy it manually if the button is unavailable.

What are a Site Reliability Engineer's key responsibilities?

A site reliability engineer is typically responsible for the 8 duties below, which cover the day-to-day work most employers expect the role to own outright. Paste them into your job post as they are, or cut the ones another team already handles — a responsibility list that claims work the hire will not actually do is the fastest way to lose a candidate at offer stage:

  • Define SLIs and SLOs with product owners and track error-budget burn against them
  • Carry the pager on a rotation and act as incident commander during major outages
  • Run blameless postmortems and drive the action items to completion, not just to a document
  • Instrument services with metrics, traces and structured logs so failures stay diagnosable
  • Measure toil and automate the worst offenders — manual restarts, runbook steps, one-off fixes
  • Load-test and forecast capacity ahead of festive traffic peaks and product launches
  • Build and rehearse failover, backup-restore and disaster-recovery procedures
  • Review new services for production readiness before they take live traffic

What qualifications does a Site Reliability Engineer need?

Employers hiring a site reliability engineer in India usually ask for the 6 qualifications below, typically alongside 3–12 yrs of experience. Keep only the ones you will genuinely screen on: every extra must-have narrows the pool, and in the Indian market a long mandatory list filters out strong candidates whose background simply reads differently on paper:

  • Ability to write real code in Go, Python or Java, not only glue scripts
  • Deep Linux, networking and distributed-systems debugging fundamentals
  • Experience with an observability stack such as Prometheus, Grafana or OpenTelemetry
  • Practical Kubernetes and cloud infrastructure exposure under production load
  • Comfort making judgement calls during an outage with incomplete information
  • Understanding of reliability trade-offs against release velocity

What skills should a Site Reliability Engineer have?

These are the 9 skills employers name most often on live site reliability engineer listings in India. Treat the first few as the ones worth screening for directly and the rest as signals a candidate can pick up on the job — asking for all of them at once is what turns a reasonable role into an unfillable one:

SLOs & error budgetsIncident responsePrometheus & GrafanaKubernetesGo / PythonDistributed tracingCapacity planningChaos testingLinux internals

Listing the three or four skills you genuinely screen on — rather than all 9 — is what keeps a site reliability engineer posting from filtering out candidates who could do the job.

What does a Site Reliability Engineer earn in India?

Typical salary (India)

typically ₹12L–₹45L/yr

Experience range

3–12 yrs

These are typical ranges and vary by city, company and skills. For live, role-specific pay data, see the OnJob salary guide.

Site Reliability Engineer job description — FAQs

What is an error budget and how is it used?

An error budget is the unreliability an SLO permits — a 99.9% monthly target allows roughly 43 minutes of failure. Teams spend that budget on risky releases and migrations. Once it runs out, the standing agreement is that feature work pauses and reliability fixes take priority, turning an argument about caution into a number both sides accepted in advance.

How is reliability engineering different from DevOps?

DevOps names a culture and a set of practices for collaboration between development and operations. Reliability engineering is a specific implementation with measurable rules: service-level objectives, error budgets, a cap on toil, and formal incident and postmortem processes. DevOps says close the gap; this discipline supplies the yardsticks for whether it actually closed.

What happens during an on-call rotation?

On-call rotations usually run a week at a time with a primary and secondary responder, and paging tied to symptom-based alerts rather than to every metric. The responder acknowledges, stabilises and escalates as needed while communicating status, and someone else digs into root cause. Healthy teams track page volume and treat repeated night alerts as bugs.

What background leads into this role?

Reliability engineers arrive by two routes: backend developers who grew interested in how systems behave under real traffic, and infrastructure specialists who learned to write proper software. Both must close the other half. Indian interviews typically cover coding, Linux and networking debugging, system design under failure, and a scenario about running a live outage.

Hiring a Site Reliability Engineer?

Post this job description on OnJob and let matching surface the candidates who fit it — instead of reading every CV yourself.

Post a Site Reliability Engineer job

Looking for a Site Reliability Engineer role instead? Browse 20,000+ live jobs.

Explore the full cluster

Everything about Site Reliability Engineer on OnJob

Move across the whole Site Reliability Engineer topic — live openings, real salary data, the job description, interview prep, and early-career routes — all in one place.