On-Call Processes and Policies - Tier 1
Tier 1 Rotations refer to on-call rotations that respond to pages from automated systems.
Active Tier 1 Rotations
SRE EOC GitLab.com
- Rotation Leader: Sarah Walker
- Coverage: 24x7
- Schedule: schedule
- Slack: #eoc-general
Responsibilities
In addition to incident management responsibilities, the EOC also is responsible for time sensitive interrupt work required to support the production environment that is not owned by another team. This includes:
- Reviewing and handling certain change requests (CRs). This includes:
- Reviewing CRs to ensure they do not conflict with any ongoing incidents or investigations
- Executing the CR directly if the author does not have the required permissions to make the change themselves (such as admin-level changes)
- Support during C1 CRs, such as database upgrades, that may occur on weekends
- Handling incident related teleport access requests
- Approving an exception for running ChatOps commands when they fail their safety checks
- Investigating and fixing buggy/flapping alerts
- Removing alerts that are no longer relevant
- Collecting production information when requested
- Responding to
@sre-oncallSlack mentions - Assisting Release Managers with deployment problems
GitLab Dedicated Platform
- Rotation Leader: Florbela Viegas
- Coverage: 24x7
- Schedule: schedule
GitLab Dedicated PubSec
- Rotation Leader: Florbela Viegas
- Coverage: 24x7
- Schedule: schedule
Incident Managers (aka IMOC)
- Rotation Leader: Devin Sylva
- Coverage: 24x7
- Schedule: schedule
Tier-1 for Service teams (STEOC)
Service teams are paged for alerts for their owned service reducing MTTR and ensuring code ownership.
The teams are paged for alerts explicitly configured to page and carry the Service label that the team owns. EOC is part of the escalation chain as the first fallback for missed pages.
Responsibilities
Service teams on Tier-1 rotations are responsible for:
- Responding to pages for alerts configured to page their owned service
- Investigating and resolving incidents related to their owned services
- Escalating to EOC or other Tier 2 and Tier 1 teams when the issue is outside the scope of their service or requires additional support
- Updating alert configurations to reduce noise and improve signal quality
- Documenting findings and resolutions in the post incident review
- Participating in post-incident reviews & creating corrective actions to improve processes and prevent recurrence
Service teams on an active Tier-1 rotation:
Runners Platform
- Rotation Leader: Kam Kyrala
- Coverage: 24x5
- Schedule: Runners Platform on-call schedule
Further details
- Incident manager rotation is staffed by certain team members in the Engineering Group.
- More information regarding the Incident Manager role, including shift schedules, responsibilities can be found in the Incident Manager on-boarding page.
Last modified August 18, 2026: Add documentation for Tier-1 oncall for Service teams (
b81a25c9)
