Somnia is looking for a Senior Infrastructure Engineer to define and maintain SLOs, build self-service infrastructure, and operate systems behind their nodes and validators. This is a remote, full-time position.
Somnia is looking for a Senior Infrastructure Engineer to define and maintain SLOs, build self-service infrastructure, and operate systems behind their nodes and validators. This is a remote, full-time position.
Key Responsibilities
- Define and maintain SLOs, SLIs, and error budgets, plus observability (metrics, logs, traces, alerts) that catches regressions before users do
- Build repeatable, self-service infrastructure through infrastructure-as-code, CI/CD and golden paths
- Own rollouts end-to-end: progressive delivery, canaries, safe migrations and clean rollbacks
- Operate the systems behind Somnia's nodes, validators, RPC and indexing, tuning for performance and cost across regions
- Lead incident response and on-call, run blameless postmortems, continuously harden the platform
- Partner with product and protocol teams to design and operate production-ready services
Requirements
- Strong experience operating production infrastructure at scale (cloud and/or bare metal), with deep Linux fundamentals
- Experience with infrastructure-as-code such as Terraform or Pulumi, alongside configuration management
- Experience running containers and orchestration platforms (Docker, Kubernetes) in production
- Strong programming skills, ideally in Go and/or TypeScript, for building automation and internal tooling
- Experience with observability stacks (Prometheus, Grafana, OpenTelemetry or equivalents)
- Experience operating and monitoring distributed systems, including capacity planning and performance tuning
- Genuine interest in crypto and on-chain systems
Nice to Have
- Experience operating blockchain node infrastructure (validators, RPC, archive nodes)
- High-performance networking or load balancing at scale
- Multi-region deployments with failover
- Security and key management (HSMs, secrets management)