Live · read from company career sites daily at 06:00 UTC

Site Reliability Engineer II Verified today

Varda Space · El Segundo, CA · Production · first seen 2026-07-31

About Varda

Varda is accelerating commercial space infrastructure development, from in-orbit pharmaceutical processing to reentry capsules. The company designs, builds, and operates W-Series vehicles in-house, including pharmaceutical processing payloads, capsules, C-PICA heat shields, and satellite buses. Varda is headquartered in El Segundo, California, with additional offices in Washington, DC and Huntsville, AL.

About This Role

As a Site Reliability Engineer, you will build, scale, and maintain infrastructure supporting systems on Earth, in orbit, and in between. You are a hands-on operator with deep working knowledge of Kubernetes and containerized technologies, applying first-principles thinking to software delivery (DevOps) and production reliability (SRE) in mission-critical environments.

What You Will Do

  • Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems
  • Build and evolve Infrastructure as Code (IaC) frameworks using tools such as Terraform
  • Implement and operate observability systems (metrics, logging, tracing) and actionable alerting
  • Build and maintain CI/CD pipelines to enable safe, repeatable, and rapid deployments
  • Partner with software and hardware engineers to deliver highly operable, reliable, and scalable systems and pipelines
  • Identify, analyze, and resolve system bottlenecks and reliability risks; perform performance tuning and implement long-term stability improvements
  • Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems
  • Rotate through the team's on-call schedule to keep critical systems healthy and responsive
  • Occasionally travel to customer sites and other Varda locations to troubleshoot, deploy, or test critical infrastructure
  • Be willing to work extended hours and weekends as needed

Basic Qualifications

  • Bachelor's degree in computer science, engineering, or related STEM field with 3+ years of Site Reliability Engineering experience, OR 5+ years of progressive experience in DevOps, SRE, or Systems Engineering in lieu of a degree
  • Experience with Infrastructure as Code (IaC) using tools like Terraform to automate server provisioning
  • Experience operating Kubernetes or similar container orchestration platforms in production environments
  • Experience with Prometheus, Grafana, InfluxDB, or similar monitoring technologies
  • Knowledge of software-defined networking (VPC, Subnets, Firewalls, VPNs, etc.)
  • Python, Bash, PowerShell (or similar) scripting experience
  • Positive and strong communication skills, both written and oral
  • Must be physically able to regularly lift 25 lbs. for duties such as delivering computers, unpacking and rack-mounting equipment, etc.

Preferred Skills and Experience

  • Experience provisioning and managing scalable Azure cloud infrastructure using native tools and best practices
  • Experience implementing configuration management, provisioning, and workflow automation solutions via Infrastructure as Code, CI/CD, and GitOps (e.g., Ansible, Salt, Argo CD, etc.)
  • Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker)
  • Experience with GPU workloads or high-throughput computing
  • Hands-on experience operating and optimizing High-Performance Computing (HPC) environments, including workload schedulers such as Slurm
  • Experience with hybrid environments (cloud + on-prem or edge systems)
  • Experience debugging distributed systems at scale (network, storage, latency)
  • Experience with databases and data modeling

Compensation and Location

Base salary range: $133,000 to $170,000 per year. Leveling and base salary are determined by job-related skills, education level, experience level, and job performance. You will be eligible for long-term incentives in the form of stock options and/or long-term cash awards.

This role is on-site in El Segundo, CA.

Requirements

Varda hires only U.S. persons due to access to export-controlled items. U.S. person means: U.S. citizen, U.S. lawful permanent resident, or protected individual as defined by 8 U.S.C. 1324b(a)(3) (i.e., individual admitted to the U.S. as a refugee or granted asylum). Varda participates in E-Verify.

Site Reliability Engineer II at Varda Space | Flight Proven Space