JPMC Candidate Experience page

Senior Lead Site Reliability / DevOps Engineer

📍 United Kingdom United KingdomUnited Kingdom 🕐 20h ago
Full-time Senior Engineering
View all jobs at JPMC Candidate Experience page

Job description

Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability / DevOps Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. Drive significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business and its technical teams, contractors, and vendorsDevelops secure and high-quality production code for reliability tooling and telemetry pipelines, and reviews and debugs code written by othersDrives decisions that influence reliability design, observability architecture, application functionality, and technical operations and processesServes as a function-wide subject matter expert in one or more areas of site reliability, observability, or telemetry engineeringLeads resiliency design reviews and breaks up complex reliability problems into digestible work for other engineers, acting as a technical lead for large-sized productsActs as the main point of contact during major incidents, demonstrating the skills to identify and solve issues quickly to avoid financial losses, and champions blameless postmortem cultureCollaborates with team members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgetsDesigns, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearchDrives the assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stabilityActively contributes to the engineering community as an advocate of firmwide frameworks, tools, and practices, and influences peers and project decision-makers to consider the use and application of leading-edge observability and reliability technologiesAdds to the team culture of diversity, opportunity, inclusion, and respectRequired qualifications, capabilities, and skills Formal training or certification on software engineering concepts and advanced applied experience delivering system design, application development, testing, and operational stabilityAdvanced knowledge of reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with considerable in-depth knowledge in one or more technical disciplines (e.g., cloud, observability, distributed systems, etc.)Advanced proficiency in one or more programming languages (e.g., Java, Python, Go, etc.)Advanced proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc.Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)Hands-on experience with the design, deployment, and operation of OpenTelemetry collectors in production environments, focusing on technical aspects such as configuring, optimizing, and troubleshooting OTLP endpoints and receiversAbility to tackle reliability design and functionality problems independently with little to no oversightPractical cloud native experienceAbility to expand and collaborate across different levels and stakeholder groups Preferred qualifications, capabilities, and skills Knowledge of distributed tracing, metrics, and logging best practicesCertification in AWS, Kubernetes, or relevant technologiesProven track record in system health monitoring, capacity management, and blameless postmortems for high-availability servicesDeep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), and Linux internalsContributions to open-source observability or telemetry projectsExperience working with agent control planes and management protocols; hands-on knowledge of OpAMP is highly desirable
Apply for this job
Similar roles you may be interested in

AWS Devops (Active SC)

Scrumconnect Consulting
📍 United KingdomUnited KingdomUnited Kingdom🕐 2w ago

About Scrumconnect Consulting: Scrumconnect Consulting is a multi-award-winning digital consultancy, recognised for delivering impactful technology solutions across UK government departments…

NewFull-timeEngineering

Senior DevOps Engineer: Cloud, CI/CD & Kubernetes (Hybrid)

Artlist
📍 NorwichUnited KingdomUnited Kingdom🕐 20h ago

A leading creative technology company in the United Kingdom is seeking a skilled DevOps Engineer to maintain CI pipelines and…

New💰 £48k–£72k a yearFull-timeEngineering

Lead DevOps Engineer

Moneypenny
📍 WrexhamUnited KingdomUnited Kingdom🕐 2w ago

We’re the leaders in outsourced calls, live chat and more, delivering brilliant conversations on behalf of businesses of all sizes…

NewInternshipEngineering

Senior DevOps Engineer / Cloud Infrastructure Engineer (Linux & AWS)

Avensure Ltd
📍 ManchesterUnited KingdomUnited Kingdom🕐 20h ago

Senior DevOps Engineer / Cloud Infrastructure Engineer (Linux & AWS) Location: Manchester (3 days onsite) / Home (2 days remote)…

New💰 £60k a yearFull-timeEngineering

DevOps Engineer

PSK Metonym Systems Ltd
📍 RainhamUnited KingdomUnited Kingdom🕐 2w ago

PSK Metonym Systems Ltd is seeking highly qualified candidates to join our team as DevOps Engineer with the experience of…

New💰 £45k–£55k a yearFull-timeEngineering

Senior DevOps Platform Engineer

Leap29 Digital
📍 LondonUnited KingdomUnited Kingdom🕐 20h ago

Senior DevOps Platform Engineer London We are currently supporting a leading organisation looking to hire an experienced Senior DevOps Platform…

NewFull-timeEngineering
All jobs at JPMC Candidate Experience page

Reading for this role

All articles