Versant company logo

Versant is hiring a SRO Lead

Get the latest jobs to your inbox!

Job Description

The System Reliability Engineering (SRE) Lead is a hands-on technical leader responsible for improving the reliability, performance, and scalability of VERSANT’s software, production, and platform systems. 

Reporting to the VP of Infrastructure, this role works closely with Software Engineering, Production Engineering, Platform Engineering, and Infrastructure teams to implement reliability best practices, drive end-to-end testing, and ensure systems perform under real-world conditions. 

This is a player-coach role focused on execution—building testing frameworks, improving observability, and helping teams proactively identify and resolve system weaknesses before they impact production. 

Key Responsibilities

 
Reliability & Engineering Practices 

  • Partner with engineering teams to improve system reliability, availability, and performance. 
  • Help define and implement SLIs, SLOs, and basic reliability standards across services. 
  • Identify reliability gaps and work with teams to address risks in system design and operations. 
  • Contribute directly to code, tooling, and automation that improves system resilience. 

E2E & System Testing Execution 

  • Design and implement end-to-end (E2E) testing workflows across distributed systems. 
  • Build and maintain integration testing frameworks validating cross-service dependencies. 
  • Execute and scale load and performance testing to validate systems under peak conditions. 
  • Partner with teams to integrate automated testing into CI/CD pipelines. 
  • Help establish practical testing standards and ensure adoption across teams. 

Performance & Capacity 

  • Support performance benchmarking and system capacity planning efforts. 
  • Analyze system performance and identify bottlenecks across application and infrastructure layers. 
  • Partner with infrastructure and platform teams to optimize system throughput and latency. 

Observability & Operations 

  • Implement and improve monitoring, logging, and alerting across services. 
  • Help ensure systems are observable, debuggable, and well-instrumented. 
  • Participate in incident response and support root cause analysis efforts. 
  • Contribute to post-incident reviews and track follow-up actions to improve reliability. 

Cross-Team Collaboration 

  • Work closely with software, platform, enterprise and production engineering teams to embed reliability practices into day-to-day development. 
  • Provide guidance and hands-on support for testing, observability, and performance improvements. 
  • Help standardize tools, frameworks, and processes used across teams. 
  • Mentor engineers on reliability engineering fundamentals and testing best practices. 
Sponsored
⭐ Featured Partner

Explore Biotech Careers

Discover exciting opportunities in biotechnology. Join innovative companies that are advancing healthcare and life sciences through cutting-edge research and development.

Remote FriendlyCompetitive SalaryBiotech

Salary Information

Salary: $135,000 - $170,000

🤖 This salary estimate is calculated by AI based on the job title, location, company, and market data. Use this as a guide for salary expectations or negotiations. The actual salary may vary based on your experience, qualifications, and company policies.

Compare salaries in Englewood Cliffs

Create a Job Alert

Interested in building your career at Versant? Get future opportunities sent straight to your email.

Create Alert

Related Opportunities

Discover similar positions that might interest you