Live · read from company career sites daily at 06:00 UTC

Senior Site Reliability Engineer Verified today

Umbra · Production · first seen 2026-08-04

Umbra is an American space technology company delivering advanced systems from sensors to spacecraft. Our Synthetic Aperture Radar (SAR) constellation transforms satellite data into actionable intelligence for national security, disaster response, and scientific discovery.

About the Role

We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems. You will leverage deep understanding of modern infrastructure and distributed systems to drive technical excellence, make thoughtful architectural decisions, and balance long-term scalability with operational reliability.

You'll partner closely with engineering teams to improve processes, champion best practices, evaluate emerging technologies, and implement solutions that enhance performance, resilience, and efficiency. This position is based on-site in Arlington VA, Reston VA, or Santa Barbara/Goleta CA.

Key Responsibilities

  • Ensure reliability and scalability of critical systems, meeting SLAs through proactive monitoring and effective incident response
  • Develop and promote new technologies and tools, conducting research and creating proofs of concept
  • Lead by example in fostering a culture of excellence and reliability
  • Continuously evaluate and improve team processes and workflows to increase efficiency
  • Collaborate with cross-functional teams, product managers, and stakeholders to align on technical strategy
  • Participate in on-call rotations, providing support and resolving complex technical issues

Required Qualifications

  • Bachelor's degree in Computer Science or a related technical field
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems
  • Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices
  • Proficiency running, optimizing, and scaling Kubernetes clusters in production environments
  • Experience using and writing Terraform to architect and manage production infrastructure
  • Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools
  • Proven success in leading teams or projects using Agile/Scrum methodologies
  • Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems
  • Experience developing and managing comprehensive infrastructure monitoring and alerting strategies

Desired Qualifications

  • 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform
  • Advanced understanding of cloud and application security, identity management, and compliance
  • Expertise in service mesh and service registration technologies, focusing on performance and reliability
  • Experience in the aerospace industry

Compensation and Benefits

Base salary range: $150,000 - $180,000 DOE

  • Flexible Time Off, Sick, Family and Medical Leave
  • Medical, Dental, Vision, Life, LTD, STD (employer funded)
  • 401k with 3% non-elective company contribution
  • Stock Options
  • Free Parking
  • Free lunch daily in office
Senior Site Reliability Engineer at Umbra | Flight Proven Space