About the job Remote | Site Reliability Engineer (SRE & Incident Management) — $70–$110/hour
We are sharing a specialised part-time consulting opportunity for experienced Site Reliability Engineering and incident management professionals with strong expertise in production reliability, incident response, operational resilience, root-cause analysis, and service performance.
This role focuses on reviewing professional documents, spreadsheets, and presentation materials related to SRE, reliability engineering, and incident management. Selected experts will assess outputs for technical accuracy, operational rigour, reliability best practices, presentation quality, and overall professional credibility.
Key Responsibilities
Site Reliability Engineering
- Evaluate work products involving production reliability, service availability, and operational resilience
- Assess whether recommendations reflect sound SRE principles and realistic production environments
- Review approaches to reliability, scalability, performance, and service health
- Identify technically weak assumptions, operational gaps, and impractical recommendations
- Apply professional judgement grounded in real-world Site Reliability Engineering experience
Incident Management & Response
- Review incident-response plans, escalation workflows, and operational procedures
- Assess incident classification, prioritisation, ownership, and coordination
- Evaluate whether proposed response actions are appropriate for severity and business impact
- Identify gaps in communication, escalation, containment, or recovery
- Review incident-management approaches for speed, clarity, and operational effectiveness
Root Cause & Post-Incident Analysis
- Evaluate root-cause analyses and post-incident reviews
- Assess whether conclusions are supported by technical and operational evidence
- Identify shallow causal analysis, unsupported assumptions, or missed contributing factors
- Review corrective and preventive actions for practicality and effectiveness
- Evaluate whether lessons learned translate into meaningful reliability improvements
Reliability Metrics & Service Health
- Review analyses involving availability, latency, reliability, and service-performance metrics
- Assess use of SLIs, SLOs, error budgets, and related reliability measures where relevant
- Evaluate whether metrics appropriately reflect service health and user impact
- Identify inconsistencies between underlying data and reported conclusions
- Review whether reliability targets and operational recommendations are realistic
Monitoring & Operational Readiness
- Evaluate monitoring, alerting, observability, and operational-readiness approaches
- Assess whether alerts are actionable and aligned with meaningful service conditions
- Review escalation paths, runbooks, and response procedures
- Identify gaps in detection, diagnosis, or operational preparedness
- Evaluate whether proposed controls support reliable production operations
Resilience & Failure Management
- Review scenarios involving outages, degraded performance, capacity constraints, and system failures
- Assess mitigation, recovery, and resilience strategies
- Evaluate trade-offs between reliability, performance, complexity, and operational cost
- Identify single points of failure or poorly addressed dependencies
- Review recommendations for reducing recurrence and improving service resilience
Documents & Presentation Review
- Evaluate incident reports, reliability analyses, operational documents, spreadsheets, and slide decks for accuracy and completeness
- Review stakeholder and executive presentations for clarity and decision usefulness
- Identify factual, technical, analytical, aesthetic, and formatting issues
- Assess whether charts, tables, and visuals accurately represent underlying operational information
- Ensure conclusions and recommendations are clearly supported by evidence
Structured Evaluation & Feedback
- Assess assigned outputs against domain-specific quality criteria
- Identify technical, operational, analytical, and presentation weaknesses
- Distinguish substantive reliability issues from minor editorial concerns
- Provide clear, structured written feedback explaining identified strengths and weaknesses
- Apply evaluation standards consistently across different SRE and incident-management work products
Ideal Profile
- 5+ years of relevant professional experience in Site Reliability Engineering, incident management, reliability engineering, DevOps, production engineering, systems engineering, or a closely related field
- Strong practical understanding of SRE and production reliability
- Hands-on experience with incident response, escalation, post-incident review, and root-cause analysis
- Experience with monitoring, observability, service health, and operational readiness
- Strong understanding of availability, performance, resilience, and service-level objectives
- Ability to assess technical recommendations for operational feasibility and reliability impact
- Highly proficient with Microsoft Office and Google Workspace
- Advanced proficiency with PowerPoint / Google Slides
- Strong spreadsheet and analytical skills
- Native or professional fluency in English
- Excellent written communication and ability to provide precise, structured feedback
- Strong attention to technical, operational, analytical, and presentation detail
- Master's degree or higher from a recognised institution is advantageous
Engagement Details
- Part-time independent contractor engagement
- Fully remote
- Flexible scheduling based on project requirements
- Compensation: $70–$110/hour
- Work includes evaluation of incident-management plans, reliability analyses, operational documentation, spreadsheets, and presentation materials
- Projects may be extended, shortened, or concluded based on project needs and performance
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
- H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.