Remote Frontend Jobs
OneLogin logo

Staff Software Engineer

OneLoginAliso Viejo, United States
APIsAWSCloud Computing

Responsibilities

Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source

  • Own production health for services - from detection through resolution to prevention
  • Improve mean time to detect (MTTD), mean time to resolve (MTTR), and recurrence rates for issues
  • Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling
  • Improve observability across services - logs, metrics, and alerting - for faster diagnosis and resolution
  • Design and improve debugging workflows, runbooks, and internal tooling for engineers
  • Reduce operational burden by making systems easier to understand, operate, and troubleshoot
  • Partner closely with product teams to feed production learnings back into design and development
  • Reduce support and incident load by addressing root causes and improving system design, not just resolving individual issues

Requirements

  • Strong production debugging experience in distributed systems
  • Experience troubleshooting complex, customer-impacting issues under real-world conditions
  • Deep familiarity with observability tools (Datadog or similar - logs, metrics, APM, tracing)
  • Experience improving operational workflows (runbooks, incident response, debugging tooling)
  • Ability to identify patterns across incidents and drive systemic fixes, not just one-off resolutions
  • Experience working across service boundaries (APIs, databases, infrastructure) to diagnose issues
  • Experience using AI tools to analyze logs, incidents, and system behavior at scale to accelerate debugging and root cause identification
  • Strong intuition for system behavior, failure modes, and performance bottlenecks
  • 4+ years of software engineering experience with ownership of production systems, reliability, or operational improvements
  • Strong backend development experience (Ruby, Node.js, or similar)
  • Solid understanding of REST APIs, service contracts, and software design principles
  • Experience working across backend services, APIs, and production systems
  • Experience building and operating services in AWS or similar cloud environments
  • Good understanding of distributed systems, cloud-native architecture, and CI/CD
  • Experience with observability, production debugging, and incident response
  • Willingness to participate in a mandatory 24/7 on-call rotation
  • Experience responding to production incidents and contributing to reliability improvements
  • Experience using, or strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor)
  • Ability to evaluate and refine AI-generated code