Staff Software Engineer
OneLoginAliso Viejo, United States
APIsAWSCloud Computing
Responsibilities
Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source
- Own production health for services - from detection through resolution to prevention
- Improve mean time to detect (MTTD), mean time to resolve (MTTR), and recurrence rates for issues
- Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling
- Improve observability across services - logs, metrics, and alerting - for faster diagnosis and resolution
- Design and improve debugging workflows, runbooks, and internal tooling for engineers
- Reduce operational burden by making systems easier to understand, operate, and troubleshoot
- Partner closely with product teams to feed production learnings back into design and development
- Reduce support and incident load by addressing root causes and improving system design, not just resolving individual issues
Requirements
- Strong production debugging experience in distributed systems
- Experience troubleshooting complex, customer-impacting issues under real-world conditions
- Deep familiarity with observability tools (Datadog or similar - logs, metrics, APM, tracing)
- Experience improving operational workflows (runbooks, incident response, debugging tooling)
- Ability to identify patterns across incidents and drive systemic fixes, not just one-off resolutions
- Experience working across service boundaries (APIs, databases, infrastructure) to diagnose issues
- Experience using AI tools to analyze logs, incidents, and system behavior at scale to accelerate debugging and root cause identification
- Strong intuition for system behavior, failure modes, and performance bottlenecks
- 4+ years of software engineering experience with ownership of production systems, reliability, or operational improvements
- Strong backend development experience (Ruby, Node.js, or similar)
- Solid understanding of REST APIs, service contracts, and software design principles
- Experience working across backend services, APIs, and production systems
- Experience building and operating services in AWS or similar cloud environments
- Good understanding of distributed systems, cloud-native architecture, and CI/CD
- Experience with observability, production debugging, and incident response
- Willingness to participate in a mandatory 24/7 on-call rotation
- Experience responding to production incidents and contributing to reliability improvements
- Experience using, or strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor)
- Ability to evaluate and refine AI-generated code