Job Openings M01 - Cloud Logging & Data Platform Engineer

About the job M01 - Cloud Logging & Data Platform Engineer

Overview 

We are looking for a Logging & Data Platform Engineer to design, build, and operate the logging and operational data-platform capabilities supporting the future SSOE platform.

You will build the platform that collects, transports, stores, indexes, searches, and serves operational data across MOE's technology environment, spanning on-premise infrastructure, networks, applications, GCC, AWS, Azure, and hybrid environments.

What You Will Be Working On

  • As a Logging & Data Platform Engineer, you will build the shared platform capabilities that enable engineering and operations teams to reliably collect and use operational data at scale.
  • You will work across logging, telemetry ingestion, data movement, storage, search, retention, and platform integration.
  • The role requires an engineer who understands both traditional enterprise infrastructure and modern cloud-native architectures and can design solutions that work securely and reliably across environment boundaries.
  • You will work closely with the Observability Engineer on telemetry requirements and with Data Engineering & Analytics on shared data-platform capabilities and integration patterns.

Key Responsibilities

Logging Platform Engineering

  • Design, build, and operate logging capabilities across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments
  • Collect logs from servers, network devices, applications, containers, databases, security appliances, cloud services, and other infrastructure sources
  • Build scalable log ingestion, routing, enrichment, storage, indexing, search, and retrieval capabilities
  • Define structured logging standards, schemas, metadata, tagging, and correlation conventions across services
  • Design appropriate retention, archival, lifecycle, and deletion policies for different classes of operational data
  • Support correlation between logs, metrics, events, and traces using common identifiers and telemetry standards
  • Work with the Observability Engineer to implement platform capabilities supporting end-to-end service observability

Data Collection & Integration

  • Design secure and resilient data movement between on-premise environments, GCC, and approved external services
  • Implement collection and forwarding patterns appropriate to different infrastructure, application, network, and security environments
  • Design for intermittent connectivity, network constraints, buffering, retry, back-pressure, and recovery between environments
  • Build event-driven and streaming patterns for moving operational data between producers and consumers
  • Integrate legacy and enterprise systems with modern cloud-native platform capabilities
  • Define clear interfaces and integration patterns between logging, observability, data engineering, and application platforms

Data Platform Engineering

  • Build shared platform capabilities for ingesting, storing, processing, querying, and serving operational data
  • Design scalable storage and query architectures appropriate to data volume, access patterns, retention requirements, and cost
  • Build ingestion, filtering, enrichment, and transformation pipelines for operational data
  • Provide APIs, query interfaces, or other serving mechanisms for authorised downstream consumers
  • Support operational datasets consumed by the User Portal, dashboards, reporting, automation, and Data Engineering & Analytics
  • Define schemas and data contracts for shared platform interfaces
  • Ensure platform changes remain backwards compatible or are coordinated with downstream consumers

Cloud & Platform Engineering

  • Design solutions using cloud-native logging, streaming, storage, search, and data capabilities
  • Build infrastructure and platform configuration using Infrastructure as Code
  • Automate build, test, deployment, configuration, and platform changes through CI/CD
  • Design for scalability, resilience, high availability, recoverability, and operational simplicity
  • Monitor platform capacity, performance, reliability, and cost

Security & Governance

  • Enforce MOE and Government data-classification requirements
  • Design secure routing and storage of operational data across security zones and environment boundaries
  • Apply appropriate encryption, access controls, authentication, and authorisation
  • Ensure logging pipelines do not unnecessarily expose credentials, secrets, or sensitive information
  • Implement audit-trail preservation and appropriate retention controls
  • Ensure data-residency requirements are considered when routing operational data between on-premise, GCC, cloud, and SaaS environments
  • Participate in security, architecture, and operational-readiness reviews

Reliability & Operations

  • Define SLOs and operational health indicators for logging and data-platform services
  • Build monitoring, alerting, failure detection, retry, and recovery into platform components
  • Monitor ingestion health, processing latency, data loss, storage utilisation, search performance, and platform availability
  • Participate in incident investigation, root-cause analysis, and post-incident reviews
  • Participate in operational support and on-call responsibilities for owned services
  • Maintain architecture documentation, operational procedures, and runbooks

What We Are Looking For

Experience

  • Minimum 3–5 years of experience in cloud engineering, platform engineering, DevOps, SRE, logging engineering, data platform engineering, or a related discipline
  • At least 2 years of hands-on experience building or operating production logging, telemetry, or data-platform capabilities
  • Demonstrated experience working with AWS and/or Azure cloud-native services
  • Experience integrating on-premise and cloud environments, or operating systems in a hybrid environment
  • Experience with production data ingestion, streaming, routing, storage, indexing, or search platforms
  • Experience implementing Infrastructure as Code and CI/CD for production environments
  • Experience designing systems for scalability, resilience, security, and operational support