Posted 2026-06-23

Senior Observability Engineer

Description

FanDuel is looking for a Senior Observability Engineer to design, build, and mature the observability ecosystem that underpins our platform and services. You will deliver deep visibility into system behaviour by combining system telemetry with user signals to provide a holistic view of performance, reliability, and user experience. You will also explore how AI and machine learning can enhance observability, from intelligent alerting and anomaly detection to accelerating root cause analysis.

This is a hands-on role where you will partner closely with engineering and product teams to deliver scalable observability capabilities. You will serve as a subject matter expert in monitoring, alerting, and incident management, and equip teams with self-service insights and tooling. By connecting system behaviour to real user impact and leveraging AI-assisted workflows to surface issues faster, you will drive improvements in reliability, performance, and data-informed decision-making across the organisation.

Responsibilities

Contribute to the observability strategy and roadmap, partnering with multiple teams to align with business priorities and engineering goals.
Design and enhance scalable observability solutions that provide actionable insights into system health, performance, and user experience.
Help establish and promote best practices for monitoring, alerting, incident management, and postmortems across teams.
Support operational excellence by improving incident response processes, on-call practices, and post-incident reviews, focusing on continuous improvement.
Collaborate on cross-team initiatives to improve system reliability, identifying risks and contributing to their resolution.
Apply automation and AI-assisted workflows to improve root cause analysis and reduce operational toil.
Work with engineering and product stakeholders to surface observability insights that inform technical decisions and prioritisation.
Analyze system and user signals to help detect, prevent, and mitigate reliability issues.
Contribute to optimizing observability platforms for performance, scalability, and cost-efficiency.
Mentor peers and contribute to raising observability and reliability standards within the team.

Requirements

Solid hands-on experience in observability engineering, SRE, platform engineering, or related roles, with impact across team-level systems. (required)
Strong expertise in monitoring and observability practices, with hands-on experience using tools such as Datadog. (required)
Experience contributing to observability or reliability initiatives across teams or services. (required)
Proficiency with Kubernetes, cloud infrastructure (e.g. AWS), and infrastructure-as-code tools such as Terraform. (required)
Ability to influence technical decisions within and across teams, collaborating effectively with a range of stakeholders. (required)
Good understanding of distributed systems principles (e.g. consistency, availability, partition tolerance) and practical trade-offs. (required)
Experience defining and implementing SLOs, SLIs, and alerting strategies, including an understanding of user-impacting metrics. (required)
Strong software engineering fundamentals, with proficiency in at least one modern programming language (e.g. Go, Java, Python, or TypeScript), and experience building tooling, automation, and scalable systems. (required)
Experience improving systems through automation, helping reduce operational toil and recurring issues. (required)
Strong analytical and problem-solving skills, with the ability to interpret technical signals and relate them to system performance and reliability. (required)
Good communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders. (required)
A sense of ownership and accountability, with a focus on delivering reliable, scalable solutions and continuous improvement. (required)

Benefits

Health plans including programs for fertility and family planning, mental health support, and fitness benefits.
Generous paid time off (PTO & sick leave).
Annual bonus and long-term incentive opportunities.
401k with up to a 5% match.
Commuter benefits.
Pet insurance.
Medical, vision, and dental insurance.
Life insurance.
Disability insurance.
14 paid company holidays.

About Flutter (FanDuel)

FanDuel Group is an innovative sports-tech entertainment company that is changing the way consumers engage with their favorite sports, teams, and leagues. The premier gaming destination in the North America, FanDuel Group consists of a portfolio of leading brands across gaming, sports betting, daily fantasy sports, advance-deposit wagering, and TV/media, including FanDuel, Stardust Casino and TVG. The company is based in New York with US offices in Los Angeles, Atlanta, and Jersey City, as well as global offices in Canada and Scotland. The company’s affiliates have offices worldwide, including in Ireland, Portugal, Romania, and Australia. FanDuel Group is a subsidiary of Flutter Entertainment, the world's largest sports betting and gaming operator with a portfolio of globally recognized brands and traded on the New York Stock Exchange (NYSE: FLUT).

Senior Observability Engineer

Want to see more roles like this?

Sign in

Job Alerts