Skit.ai logoS

Senior Site Reliability Engineer — Voice AI Platform

Skit.ai

Bangalore, KA, IndiaFull Time

Senior Site Reliability Engineer — Voice AI Platform at Skit.ai is a full time role based in Bangalore, KA, India. It was published on 5 July 2026 and was open at last check.

Senior Site Reliability Engineer — Voice AI Platform at Skit.ai — key details
RoleSenior Site Reliability Engineer — Voice AI Platform
CompanySkit.ai
LocationBangalore, KA, India
Employment typeFull Time
Published5 July 2026
StatusOpen at last check

About the Role

Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Senior Site Reliability Engineer — Voice AI Platform

Type: Full-time

Location: Bangalore

Why this role exists:

We run a voice AI platform that places and answers up to ~1 million calls per hour for regulated enterprises in banking, telecom, and collections. Unlike most SaaS, our workload is real-time and conversational: every call is a live media session where an extra few hundred milliseconds anywhere in the ASR → LLM → TTS loop is the difference between a natural exchange and a caller hanging up. Traffic is also bursty — outbound campaigns spin up huge concurrency inside narrow calling windows — and it runs across multiple clouds for resilience and data residency.

We are hiring a Senior SRE to own the reliability and performance of that system: the SLOs, the observability that makes problems visible, the capacity that absorbs campaign spikes, and the incident response that keeps regulated clients online. This is a systems-reliability role — latency, uptime, saturation, and the health of the telephony and serving path. (Model quality and evaluation live with a separate AI Observability role; you'll partner with them, not own their signals.)

If you want reliability problems that are genuinely hard — real-time media, sub-second budgets, six-figure concurrency, multi-cloud failover — this is that.

What you'll own:

  • SLOs and error budgets. Define and defend service-level objectives for availability and latency across the call path, and use error budgets to steer the balance between shipping and stability.
  • Observability. Own the metrics, tracing, and logging stack so failures surface fast and root cause is minutes not hours — distributed traces across signaling, ASR, LLM, TTS, and infra, with dashboards and alerting that page on real problems and stay quiet otherwise.
  • The real-time media path. Keep SIP signaling and RTP media healthy at scale — concurrency, jitter, packet loss, session setup — and the reliability of the components that carry them.
  • Capacity and autoscaling. Plan for peak (campaign windows that push toward the platform's concurrency ceiling), pre-warm capacity ahead of demand, and tune autoscaling so we neither drop calls nor burn money idling.
  • Incident response. Run a calm, structured on-call: triage, mitigation, clear comms to stakeholders on regulated accounts, and blameless postmortems that actually change the system.
  • Multi-cloud resilience. Design for failure across AWS, GCP, and Azure — redundancy, failover, disaster recovery, and the data-residency constraints that come with Indian banking and telecom clients.
  • Automation and toil reduction. Turn manual operations into infrastructure-as-code and self-healing systems. If you did it twice by hand, the third time is a script.

What the first your looks like:

  • First 90 days. Learn the call path end to end. Establish baseline SLIs for availability and latency, close the biggest gaps in alerting, and take a full turn in the on-call rotation.
  • By 6 months. Published SLOs with error budgets for the core services. A tracing/dashboards setup that makes the ASR→LLM→TTS latency budget visible per call. A repeatable pre-warm-and-scale playbook for campaign peaks.
  • By 12 months. Demonstrable reduction in incident frequency and time-to-mitigate. Tested multi-cloud failover for a critical path. On-call toil measurably down through automation.

What we're looking for

Must-have

  • Several years running high-availability, high-throughput production systems, including real on-call ownership.
  • Depth in at least one major cloud (AWS, GCP, or Azure) and with containers/Kubernetes.
  • Strong observability practice — metrics, distributed tracing, and logging (e.g. Prometheus/Grafana, OpenTelemetry, Tempo/Jaeger).
  • Fluency with SLIs/SLOs/error budgets and structured incident management.
  • Infrastructure-as-code (Terraform or similar) and CI/CD.
  • A programming language for automation and tooling (Python, Go, or similar) — beyond shell scripting.
  • Solid Linux systems and networking fundamentals; capacity planning and performance tuning.

Nice-to-have

  • Real-time media or VoIP experience — SIP/RTP, media servers, LiveKit, SBC/Kamailio.
  • Reliability of GPU/ML serving infrastructure.
  • Regulated-industry operations — uptime SLAs, DR, data residency.
  • Load testing and chaos engineering at scale.
  • PostgreSQL operations at scale.

Our stack:

Representative — you'll help shape it. Multi-cloud across AWS, GCP, and Azure; LiveKit/SIP for telephony; self-hosted and managed ASR (e.g. NVIDIA Parakeet / NeMo), LLMs, and TTS; Modal for ML deployment and pre-warming; PostgreSQL; Grafana/Tempo for metrics and traces; infrastructure-as-code and GitHub Actions CI/CD.

How you'll know you're succeeding:

Calls connect and stay fast even during the busiest campaign windows. Alerts mean something, and the ones that page you are worth waking up for. When something breaks, it's found and mitigated quickly and it doesn't break the same way twice. And the on-call rotation gets calmer over time, not busier, because the system increasingly heals itself.

We're an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]

Share:WhatsAppLinkedIn

Create your free OnJob profile to apply — we'll take you to Skit.ai's application after sign-up. · Posted 5 Jul 2026.

Senior Site Reliability Engineer — Voice AI Platform at Skit.ai — questions answered

What does the Senior Site Reliability Engineer — Voice AI Platform role at Skit.ai pay?

Skit.ai does not publish a salary on this Senior Site Reliability Engineer — Voice AI Platform listing, so OnJob shows no figure for it rather than an estimate. For what this role pays across the market, the OnJob salary guides aggregate the live listings that do disclose pay.

Where is the Senior Site Reliability Engineer — Voice AI Platform role at Skit.ai based?

Skit.ai lists this Senior Site Reliability Engineer — Voice AI Platform role in Bangalore, KA, India, advertised as full time work at that location. Larger employers sometimes cover several sites under one city name, so confirm the exact office with Skit.ai before you apply.

Is the Senior Site Reliability Engineer — Voice AI Platform role at Skit.ai still open?

The Senior Site Reliability Engineer — Voice AI Platform posting at Skit.ai was open at OnJob's last check of the employer's careers page, having been published on 5 July 2026. OnJob re-checks source listings on each build and marks a role closed once it disappears, but listings can close without notice, so the employer's own page is the final word.

How do you apply for the Senior Site Reliability Engineer — Voice AI Platform role at Skit.ai?

Apply to the Senior Site Reliability Engineer — Voice AI Platform at Skit.ai role through OnJob with a free profile: OnJob scores your fit against the listing, shows the skills lowering that score, and submits an ATS-ready profile to Skit.ai's own application page. Creating a profile is free and needs no card.

Related Data & AI jobs

Hand-picked roles that match this listing on skills, category and location — each scored to your profile inside OnJob.

Explore more on OnJob

Hiring for a role like this?

Post a job on OnJob and reach AI-matched candidates.

Post a Job