NVIDIA logoN

Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA

Bengaluru, IndiaFull Time

Senior Platform and EngOps Engineer - Cluster Operations at NVIDIA is a full time role based in Bengaluru, India. It was published on 30 July 2026 and was open at last check.

Senior Platform and EngOps Engineer - Cluster Operations at NVIDIA — key details
RoleSenior Platform and EngOps Engineer - Cluster Operations
CompanyNVIDIA
LocationBengaluru, India
Employment typeFull Time
Published30 July 2026
StatusOpen at last check

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

Join our team of innovative engineers who develop and maintain software facilitating GPU communication, driving groundbreaking solutions in High Performance Computing and Deep Learning. We're looking for highly motivated EngOps and Platform Engineers to boost execution efficiency while managing and maintaining large GPU clusters interconnected via NVLink and InfiniBand.

What you will be doing:

  • Develop automated tools to efficiently deploy, provision, and maintain extensive GPU clusters interconnected via NVLink and InfiniBand
  • Implement modern DevOps tools to automate software updates, perform maintenance tasks, and monitor cluster availability, ensuring seamless operations.
  • Take ownership of daily cluster failures and issues, troubleshooting them promptly to maintain optimal cluster availability and performance.
  • Manage the rollout and rollback of cluster software and firmware updates, ensuring smooth transitions and minimal disruptions.
  • Collaborate effectively with dynamic Engineering and Product Teams across multiple time zones to align cluster operations with evolving project requirements.

What we need to see:

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • 5+ years of hands-on experience in deploying and administrating clusters, servers, switches, and related infrastructure.
  • Automation expert with hands on skills in Ansible, Python and Shell Scripting.
  • Deep understanding of operating systems, computer networks, and high-performance applications.
  • Proven ability to work effectively with developers and test engineers across different teams and time zones.
  • Proficient with Linux fundamentals.

Ways to stand out from the crowd:

  • Familiarity with resource scheduling managers, preferably Slurm.
  • Direct experience with industry standard alerting tools and emergency response practices.
  • Hands-on experience with GPU-focused hardware and software, such as DGX systems and Compute Clusters.
  • Proficiency in crafting and implementing a robust metrics collection and alerting infrastructure.
  • Proficiency in designing large scale networking technologies and the associated challenges.
Share:WhatsAppLinkedIn

Create your free OnJob profile to apply — we'll take you to NVIDIA's application after sign-up. · Posted 30 Jul 2026.

Senior Platform and EngOps Engineer - Cluster Operations at NVIDIA — questions answered

What does the Senior Platform and EngOps Engineer - Cluster Operations role at NVIDIA pay?

NVIDIA does not publish a salary on this Senior Platform and EngOps Engineer - Cluster Operations listing, so OnJob shows no figure for it rather than an estimate. For what this role pays across the market, the OnJob salary guides aggregate the live listings that do disclose pay.

Where is the Senior Platform and EngOps Engineer - Cluster Operations role at NVIDIA based?

NVIDIA lists this Senior Platform and EngOps Engineer - Cluster Operations role in Bengaluru, India, advertised as full time work at that location. Larger employers sometimes cover several sites under one city name, so confirm the exact office with NVIDIA before you apply.

Is the Senior Platform and EngOps Engineer - Cluster Operations role at NVIDIA still open?

The Senior Platform and EngOps Engineer - Cluster Operations posting at NVIDIA was open at OnJob's last check of the employer's careers page, having been published on 30 July 2026. OnJob re-checks source listings on each build and marks a role closed once it disappears, but listings can close without notice, so the employer's own page is the final word.

How do you apply for the Senior Platform and EngOps Engineer - Cluster Operations role at NVIDIA?

Apply to the Senior Platform and EngOps Engineer - Cluster Operations at NVIDIA role through OnJob with a free profile: OnJob scores your fit against the listing, shows the skills lowering that score, and submits an ATS-ready profile to NVIDIA's own application page. Creating a profile is free and needs no card.

Related Engineering jobs

Hand-picked roles that match this listing on skills, category and location — each scored to your profile inside OnJob.

Explore more on OnJob

Hiring for a role like this?

Post a job on OnJob and reach AI-matched candidates.

Post a Job