Site Reliability Engineer Verified today
About Outpost
Outpost is building carrier-agnostic truck terminals across America. The platform combines AI-powered gate automation, computer vision, and operational software to help logistics operators run smarter, faster facilities. When the system goes down, trucks stop moving and customer yards stop running. The company is backed by $1B from Greenpoint Partners and is scaling rapidly.
The Role
You will own uptime and incident response for mission-critical infrastructure serving major logistics providers. Reliability is core to customer trust as the system scales 10X over the next 18 months. You will shift the team from reactive to proactive incident management.
Your responsibilities:
- Own reliability targets across backend/API, worker services, applications, and computer vision pipeline; track mean time to detection (MTD), mean time to mitigation (MTM), mean time to resolution (MTR), and drive root cause fixes
- Build out monitoring, alerting, and auto-remediation so on-call load scales with automation, not headcount
- Partner with the engineering team to build AI agents that triage alerts and handle routine remediation
- Harden and optimize GCP infrastructure (Cloud Run, Cloud SQL, Cloud Storage) for cost and performance as load scales
- Own database scale and performance: connection pooling, query optimization, indexing, read replicas, and capacity planning so Postgres doesn't become the bottleneck
- Improve reliability of ML training and monitoring infrastructure with the computer vision and ML team
- Run blameless postmortems and drive fixes for root causes, not just symptoms
- Participate in on-call rotation
Requirements
- 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership
- Deep experience with a major cloud provider (GCP preferred): compute, managed databases, object storage, networking
- Experience building monitoring, alerting, and observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar)
- Strong scripting and automation skills (Python, Bash, or similar)
- Comfortable with containerized workloads (Docker) and CI/CD pipelines
- Track record of reducing incident volume or improving reliability metrics, not just responding to incidents
- Strong communication skills; comfortable working with technical and non-technical stakeholders; clear written English and strong async communication
Preferred Qualifications
- Experience with ML/data infrastructure: training pipelines, model monitoring, feature stores
- Experience building or integrating AI agents for operational automation
- Infrastructure-as-code experience (Terraform or similar)
- PostgreSQL performance tuning at scale
- Background supporting physical or IoT systems (edge devices, cameras, on-site hardware)
- Experience with bare-metal infrastructure in colocation environments, hardware monitoring, redundancy, and failover/high-availability configuration
Working Arrangements
Remote, Latin America preferred. Full-time contract position. You will be embedded in sprints, standups, and Slack channels with the core engineering team. Employment is managed through an agency.