Site Reliability Engineer
CXM
Site Reliability Engineer at CXM is a full time role based in Remote · Worldwide (remote). It was published on 3 September 2026 and was open at last check.
| Role | Site Reliability Engineer |
|---|---|
| Company | CXM |
| Location | Remote · Worldwide (remote) |
| Employment type | Full Time |
| Published | 3 September 2026 |
| Status | Open at last check |
Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services.
This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience.
Position Details
TeamPlatform & Production Reliability
LocationRemote (Americas, LatAm preferred)
Working HoursAmericas time zones (UTC-3 to UTC-8)
On-callRotation aligned with the London trading day
Employment TypeFull-time, Permanent
Experience LevelMid-Level (3–5 years)
Technology Stack.NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform
About the Role
Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.
You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence.
What You'll Do
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
- Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
- Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Troubleshoot issues across:
- .NET/C# applications
- Windows Server
- Aurora PostgreSQL databases
- AWS infrastructure
- CI/CD pipelines and deployments
- Improve deployment safety, release automation, and rollback strategies.
- Partner with developers to improve application operability, resilience, and fault isolation.
- Automate operational tasks through scripting and infrastructure automation.
- Create and maintain runbooks, operational documentation, and incident response procedures.
- Continuously improve monitoring, alert quality, automation, and platform reliability.
Required Technical Skills
.NET & Windows
- Strong experience debugging and supporting .NET/C# applications in production.
- Hands-on experience with Windows Server environments.
Scripting & Automation
- Strong PowerShell scripting skills.
- Experience with Python or Bash.
Observability
- Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
- Solid understanding of metrics, logging, tracing, and alerting best practices.
CI/CD & DevOps
- Experience with modern CI/CD pipelines.
- Knowledge of deployment strategies, release automation, and rollback mechanisms.
Cloud & Infrastructure
- Experience working with AWS.
- Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
Databases
- Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
Reliability Engineering
- Practical experience with:
- SLIs & SLOs
- Error Budgets
- Incident Response
- Root Cause Analysis (RCA)
- Alert Design
- Production Operations
Preferred Qualifications
- Experience supporting high-availability or low-latency financial or trading systems.
- Familiarity with MetaTrader environments or financial technology platforms.
- Experience with distributed systems and microservices.
- Knowledge of OpenTelemetry or similar observability frameworks.
- Exposure to Docker, Kubernetes, or containerized environments.
Why Join Us?
- Work on mission-critical trading infrastructure that directly impacts customers.
- Solve challenging reliability and scalability problems in a real-time environment.
- Build world-class observability, automation, and deployment practices.
- Collaborate with experienced engineers in a modern engineering culture.
- Influence reliability strategy and engineering best practices across the platform.
If you're passionate about production engineering, automation, and building reliable systems at scale, we'd love to hear from you.
Originally posted on Himalayas
Create your free OnJob profile to apply — we'll take you to CXM's application after sign-up. · Posted 3 Sept 2026.
Site Reliability Engineer at CXM — questions answered
What does the Site Reliability Engineer role at CXM pay?
CXM does not publish a salary on this Site Reliability Engineer listing, so OnJob shows no figure for it rather than an estimate. For what this role pays across the market, the OnJob salary guides aggregate the live listings that do disclose pay.
Is the Site Reliability Engineer role at CXM remote?
Yes — CXM advertises this Site Reliability Engineer role as remote, tied to Remote · Worldwide, and it is listed as full time work. Remote terms come from the employer's own posting, so confirm the expected working hours, timezone and any on-site days with CXM before you apply.
Is the Site Reliability Engineer role at CXM still open?
The Site Reliability Engineer posting at CXM was open at OnJob's last check of the employer's careers page, having been published on 3 September 2026. OnJob re-checks source listings on each build and marks a role closed once it disappears, but listings can close without notice, so the employer's own page is the final word.
How do you apply for the Site Reliability Engineer role at CXM?
Apply to the Site Reliability Engineer at CXM role through OnJob with a free profile: OnJob scores your fit against the listing, shows the skills lowering that score, and submits an ATS-ready profile to CXM's own application page. Creating a profile is free and needs no card.
Explore more on OnJob
Hiring for a role like this?
Post a job on OnJob and reach AI-matched candidates.