Skit.ai logoS

Lead DevOps Engineer

Skit.ai

Bangalore, KA, IndiaFull Time
Apply with OnJob Free profile · takes about a minute

Lead DevOps Engineer at Skit.ai is a full time role based in Bangalore, KA, India. It was published on 7 September 2026 and was open at last check.

Lead DevOps Engineer at Skit.ai — key details
RoleLead DevOps Engineer
CompanySkit.ai
LocationBangalore, KA, India
Employment typeFull Time
Published7 September 2026
StatusOpen at last check

Job Title: Lead DevOps Engineer

Location: Bengaluru (100% WFO)

Job Type: Full-time

The problem

Skit.ai runs autonomous voice agents for regulated enterprises — India's largest banks and telcos, and US collections operations. Every call is a live distributed system: PSTN/SIP → media server → ASR → LLM → TTS → back, spread across three clouds and multiple vendors, with a conversational response budget measured in hundreds of milliseconds.

The platform peaks at roughly **1 million calls per hour**. Billed minutes grew **5,000x+ in eight months**. At this scale, infrastructure is not a support function — latency, cost-per-minute, and auditability are product features. When infra degrades, a customer mid-sentence hears silence.

We're hiring a Lead DevOps Engineer to own this substrate and keep it ahead of the growth curve.

What you'll own:

  • Multi-cloud substrate : Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget — p95 per hop in tens of milliseconds, not "best effort."
  • Real-time media plane: Self-hosted LiveKit and SIP infrastructure at scale. Media servers are stateful; you'll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
  • Model-serving infrastructure : GPU fleets (A100/H100/B200-class) for self-hosted ASR and open-weight LLMs — inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets — alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
  • Reliability & observability : OTel-native tracing (Grafana/Tempo stack), per-turn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You'll act as incident commander for infrastructure and write the runbooks you'd want at 3 a.m.
  • Cost engineering : Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture — you'll own the next 10x.
  • Security & compliance : Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001; working command of ISO/IEC 42001:2023 (AI management systems) — hands-on preferred, rigorous theoretical grounding acceptable. You'll face bank and telecom auditors directly, including data-residency requirements.
  • Technical leadership : Terraform-first IaC standards, production-readiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team.

Problems on our plate right now

  • Scaling stateful, self-hosted media servers past current concurrency ceilings — HPA on CPU doesn't cut it
  • Migrating LLM inference from managed APIs to self-hosted open-weight models on GPUs without breaking TTFT budgets
  • ASR, LLM, and telephony living in different clouds: interconnect topology that keeps the packet path short and private
  • Multi-region DR that satisfies bank audits without doubling spend

If these read as interesting rather than terrifying, keep reading.

Must-have

  • 6+ years hands-on cloud infrastructure; 3+ years operating multiple clouds simultaneously in production; deep expertise in at least two of AWS/GCP/Azure
  • Real-time audio/video systems in production — WebRTC, SIP/PSTN, or streaming media; you've debugged jitter, not just read about it
  • Networking depth: VPC/VNet design, load balancing, DNS, NAT; private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect); transit gateways and cross-cloud mesh
  • Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling *stateful* workloads; service mesh familiarity (Istio/Linkerd)
  • Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates; drift is a bug
  • AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
  • Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR; ElevenLabs / Azure-class TTS)
  • Security fundamentals: IAM/RBAC, secrets management (Vault or cloud-native), encryption and key rotation
  • CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline

Strong signal (nice-to-have)

  • LiveKit, pipecat, Twilio, or comparable real-time platforms; SIP trunking and PSTN integration
  • KEDA or other event-driven autoscaling used in anger
  • MLOps: model versioning, canary and A/B rollout
  • FinOps discipline: reserved/spot strategy, unit-economics reporting
  • Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert
  • ISO/IEC 42001:2023 exposure

What we're NOT looking for

  • Single-cloud depth with documentation-level knowledge of the other two
  • Tool-checklist DevOps without production AI/ML serving scars
  • "Can learn quickly" as the primary qualification — this role needs day-one production credibility
  • Anyone who has never traced a packet across a cloud boundary
Share:WhatsAppLinkedIn

Create your free OnJob profile to apply — we'll take you to Skit.ai's application after sign-up. · Posted 7 Sept 2026.

Lead DevOps Engineer at Skit.ai — questions answered

What does the Lead DevOps Engineer role at Skit.ai pay?

Skit.ai does not publish a salary on this Lead DevOps Engineer listing, so OnJob shows no figure for it rather than an estimate. For what this role pays across the market, the OnJob salary guides aggregate the live listings that do disclose pay.

Where is the Lead DevOps Engineer role at Skit.ai based?

Skit.ai lists this Lead DevOps Engineer role in Bangalore, KA, India, advertised as full time work at that location. Larger employers sometimes cover several sites under one city name, so confirm the exact office with Skit.ai before you apply.

Is the Lead DevOps Engineer role at Skit.ai still open?

The Lead DevOps Engineer posting at Skit.ai was open at OnJob's last check of the employer's careers page, having been published on 7 September 2026. OnJob re-checks source listings on each build and marks a role closed once it disappears, but listings can close without notice, so the employer's own page is the final word.

How do you apply for the Lead DevOps Engineer role at Skit.ai?

Apply to the Lead DevOps Engineer at Skit.ai role through OnJob with a free profile: OnJob scores your fit against the listing, shows the skills lowering that score, and submits an ATS-ready profile to Skit.ai's own application page. Creating a profile is free and needs no card.

Explore more on OnJob

Hiring for a role like this?

Post a job on OnJob and reach AI-matched candidates.

Post a Job