This job is no longer available
The position may have been filled or the posting has expired. Browse similar opportunities below.
Link copied to clipboard!
Back to Jobs
Platform Reliability Engineer - Principal Engineer at Wells Fargo
Wells Fargo
No longer available
Information Technology
Posted 13 hours ago
JOB DESCRIPTION
Wells Fargo is seeking a Principal Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one of the following technology domains: Enterprise Tools / Applications (API Management, Scheduling, Observability, Telemetry, CICD, Testing) and have demonstrated experience collaborating across at least one additional infrastructure domains (Network, Database or Storage). The expectation is that this engineer will elevate themselves in looking for trends and patterns that are not limited to these streams and investigate systemic issues that span multiple streams or need a deeper troubleshooting.As part of our Platform Reliability Engineering (PRE) team, you will apply modern Site Reliability Engineering (SRE) practices to improve the availability, resiliency, observability, scalability, and operational excellence of critical enterprise platforms. You will leverage your domain expertise to identify systemic issues, drive automation, and deliver engineering solutions that strengthen platform stability at scale.In this role, you will:Serve as the reliability engineering expert for your primary domain (Enterprise Tools /Applications, Network, Database, Storage) while partnering across adjacent technology disciplinesLead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actionsApply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform healthLead capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact occursPerform deep performance analysis across infrastructure layers, identifying bottlenecks, contention points, latency drivers, and resource inefficienciesIdentify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliabilityDrive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineeringDesign and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilitiesDefine and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoringPartner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availabilityLead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvementsIdentify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholdersMentor engineers and technical teams on reliability engineering, operational excellence, automation, and platform best practicesAct as an advisor to leadership to develop or influence applications, network, information security, database, operating systems, or web technologies for highly complex business and technical needs across multiple groupsLead the strategy and resolution of highly complex and unique challenges requiring in-depth evaluation across multiple areas or the enterprise, delivering solutions that are long-term, large-scale and require vision, creativity, innovation, advanced analytical and inductive thinkingMaintain knowledge of industry best practices and new technologies and recommends innovations that enhance operations or provide a competitive advantage to the organizationStrategically engage with all levels of professionals and managers across the enterprise and serve as an expert advisor to leadershipRequired Qualifications:7+ years of Engineering experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education5+ years supporting and engineering enterprise-scale production environments5+ years of experience with hands-on expertise in one of the following technology domains: Enterprise Tools / Applications (API Management, Scheduling, Observability, Telemetry, CICD, Testing) and secondary skills in:Network Engineering (routing, switching, load balancing, DNS, network observability, performance analysis)Database Engineering (Oracle, SQL Server, PostgreSQL, MongoDB, database performance, replication, HA/DR)Storage Engineering (SAN/NAS technologies, storage virtualization, backup/recovery, performance and capacity management)Desired QualificationsStrong experience applying SRE principles, including SLI/SLO development, error budgets, incident analysis, and reliability measurementExperience supporting highly available, mission-critical production environmentsProven success troubleshooting complex issues spanning multiple technology domains in large-scale distributed environmentsExperience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, and performance optimizationHands-on experience with observability and monitoring platforms such as Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, or similar technologiesExperience building dashboards, alerts, service health indicators, and operational reportingStrong automation and scripting experience using Python, Bash, PowerShell, or similar technologiesExperience developing operational tooling, API integrations, self-healing capabilities, and automated remediation solutionsFamiliarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or TerraformExperience diagnosing and resolving issues that span multiple infrastructure layersAbility to influence technical direction across infrastructure and engineering organizationsExperience leading major incident reviews and driving sustainable operational improvementsDemonstrated success mentoring engineers and promoting reliability engineering best practicesStrong communication skills with the ability to translate technical concepts into business-focused outcomesJob Expectations:This position offers a hybrid scheduleThis position does not offer Visa sponsorshipPay RangeReflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.$159,000.00 - $305,000.00Benefits Wells Fargo provides eligible employees with a comprehensive set of benefits, many of which are listed below. Visit Benefits - Wells Fargo Jobs for an overview of the following benefit plans and programs offered to employees.Health benefits401(k) PlanPaid time offDisability benefitsLife insurance, critical illness insurance, and accident insuranceParental leaveCritical caregiving leaveDiscounts and savingsCommuter benefitsTuition reimbursementScholarships for dependent childrenAdoption reimbursementPosting End Date:6 Aug 2026*Job posting may come down early due to volume of applicants.We Value Equal OpportunityWells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit’s risk appetite and all risk and compliance program requirements.Applicants with DisabilitiesTo request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo.Drug and Alcohol PolicyWells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.Wells Fargo Recruitment and Hiring Requirements:a. Third-Party recordings are prohibited unless authorized by Wells Fargo.b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.#DNP-INDSummaryLocation: ISELIN, NJ; IRVING, TX; CHARLOTTE, NCType: Full time