About the job Observability Engineer (SRE)
Job Description:
Collaborate with engineering, operations, and other stakeholders to understand enterprise architecture, monitoring requirements & performance goals.
Identify and define key performance indicators (KPIs) metrics, diagnose issues, and proactively identify areas for optimization.
Develop and implement observability frameworks, tools, and processes to enable comprehensive monitoring, logging, and tracing of systems and applications.
Ensure the availability, scalability, and reliability of infrastructure and deployment environments.
Implement and manage monitoring and observability tools(AppDynamics/DataDog/Splunk/ELK/Sentry etc) to gain insights into system performance and health.
Provide timely and accurate reports on application performance, highlighting key insights and trends.
Collaborate with digital squads to implement performance improvements, including code optimizations and infrastructure adjustments.
Offer guidance and training to end-users and internal teams on best practices for APM and optimizing application performance.
Required Skills and Experience:
Bachelor's or master's degree in Information Technology, Computer Science, or a related quantitative discipline
Overall, around 8+ years of experience with IT Infrastructure, Applications
3-5 years of hands-on experience in Observability and continuous integration.
2 years of programming background in Java or relevant technologies
Knowledge of cloud infrastructure (Azure) and cluster management tools like Kubernetes
Strong communication skills with ability to align the organization on complex technical decisions