Gitlab
Site Reliability Engineer, Infrastructure Platforms — AMER (Intermediate to Senior Staff)
Remote, Canada; Remote, US · remote
Company's own board
First seen Aug 4 · seen live today · from Gitlab's own Greenhouse board
Skills mentioned
goawsgcpkubernetesterraformci/cd
The posting, as published
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.
The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.
* Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.
An overview of this role
Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure.
This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs. We hire Site Reliability Engineers from Intermediate through Senior Staff across multiple Infrastructure Platforms teams.
We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly. We'll support you in becoming successful with GitLab's tools, systems, and ways of working.
How our SRE hiring works
Because this is a single application for SRE roles across Infrastructure Platforms, our process is built to evaluate you once and match you well, rather than interviewing separately for every team.
Recruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction.
Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering system architecture and incident review.
Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab.
Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on.
Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams.
After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time.
We’ll calibrate your level throughout the interview process based on the scope and impact of your experience.
Intermediate: You independently deliver meaningful reliability improvements within a defined area.
Senior: You own complex reliability work end to end and raise the effectiveness of your team.
Staff: You shape reliability across multiple teams, solving systemic problems and creating approaches others can reuse.
Senior Staff: You set technical direction across a broader Infrastructure area and influence reliability strategy at organizational scale.
What you'll do
Keep user-facing services and production systems reliable, scalable, and efficient
Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows
Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps
Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately
Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
Document runbooks, architecture decisions, and reviews so your findings become repeatable practices
What you'll bring
Experience keeping production systems reliable, combining an operations mindset with real software engineering practice
Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes
Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level
Hands-on experience with at least one major cloud provider (GCP or AWS)
Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions
Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure
Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment
A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work
Alignment with GitLab's values and a commitment to working in accordance with them
About the team
Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab’s user-facing services, most notably GitLab.com . The organization spans teams across Production Engineering , GitLab Dedicated , GitLab Delivery , and Developer Experience , covering everything from the production fleet and networking platform to observability, incident response, deployment infrastructure, tenant scale, and our single-tenant Dedicated offering.
We are a globally distributed, remote-first organization that works asynchronously, favors automation over toil, and uses monitoring, metrics, and clear ownership to continuously improve the reliability of GitLab at scale. For more on how we work, see the Infrastructure Handbook Page .
The base salary range for this role’s listed level is currently for residents of the United States only. This range is intended to reflect the role's base salary rate in locations throughout the US. Grade level and salary ranges are determined through interviews and a review of education, experience, knowledge, skills, abilities of the applicant, equity with other team members, alignment with market data, and geographic location. The base salary range does not include any bonuses, equity, or benefits. See more information on our benefits and equity . Sales roles are also eligible for incentive pay targeted at up to 100% of the offered base salary.
United States Salary Range
$126,400 — $314,400 USD
How GitLab Supports Full-Time Employees