Overview As a Site Reliability Engineer (SRE) Level II, you will play a key role in maintaining the availability, scalability, and performance of critical infrastructure and services. You will be responsible for building and automating solutions that enhance system reliability and support continuous delivery. In this role, you will handle more complex operational tasks and incidents, provide mentorship to junior SREs, and collaborate with development teams to ensure systems are designed for reliability from the ground up. Responsibilities Lead troubleshooting efforts for high-impact production issues, providing detailed root cause analysis (RCA) and preventative measures. Participate in on-call rotations, acting as an escalation point for Level 1 SREs during major incidents. Develop and maintain automation scripts and infrastructure using tools like Terraform, Ansible, or CloudFormation. Implement automation solutions to eliminate manual tasks and improve system reliability, scalability, and performance. Analyze system performance and recommend optimizations for scalability and reliability. Support capacity planning by monitoring system metrics, traffic patterns, and usage trends to predict future resource needs. Collaborate with software engineering teams to influence the design of new services and applications, ensuring they are scalable, reliable, and resilient from the start. Contribute to architectural decisions, ensuring alignment with best practices in fault tolerance, redundancy, and recovery. Build and maintain robust monitoring, alerting, and observability solutions to proactively detect and resolve issues before they impact end users. Optimize existing monitoring tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) and build custom dashboards for better visibility into system health. Ensure systems and infrastructure are secure, compliant, and aligned with organizational policies and industry best practices. Assist with vulnerability management, system patching, and implementing security measures to protect the integrity and availability of services. Lead efforts to continuously improve operational processes, tools, and workflows. Implement and enforce best practices in deployment, monitoring, and incident management to improve overall system reliability and reduce downtime. Qualifications Bachelor’s degree in computer science, Information Technology, or a related field, or equivalent work experience. 3 years of experience in site reliability engineering, DevOps, systems administration, or related roles. Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance. Preferred Qualifications Strong experience with Linux/Unix administration and proficiency in scripting (e.g., Python, Bash, Go). 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles. Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (EC2, S3, Lambda, Kubernetes, etc.). Experience with containerization and orchestration technologies like Docker and Kubernetes. Proficiency with monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Datadog, ELK Stack, or similar platforms. Strong understanding of networking fundamentals (DNS, TCP/IP), load balancing, and CDNs. Experience with CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure automation (Terraform, Ansible, Puppet). Familiarity with distributed systems and microservices architecture. Excellent problem-solving and troubleshooting skills, especially in diagnosing production issues in high-scale environments. Microsoft Office experience. Experience working in multi-platform environments. Ability to balance both development and support roles. Experience in working on projects that involve business segments. Strong analytical and troubleshooting skills and excellent communication skills. Strong interpersonal skills, focus on customer service, and the ability to work well with IT, vendor, and business groups. Exempt Status: Yes (not eligible for overtime pay) / No (eligible for overtime pay). Huntington is an Equal Opportunity Employer. Tobacco-Free Hiring Practice: Visit Huntington's Career Web Site for more details. Note to Agency Recruiters: Huntington Bank will not pay a fee for any placement resulting from the receipt of an unsolicited resume. All unsolicited resumes sent to any Huntington Bank colleagues, directly or indirectly, will be considered Huntington Bank property. Recruiting agencies must have a valid, written and fully executed Master Service Agreement and Statement of Work for consideration. #J-18808-Ljbffr Huntington National Bank
Job Description Who We ARE: When you work at the Best. Gym. Ever, you join the Best. Team. Ever. Youll walk into our clean and spacious gyms with a smile on your face and a pep in your step because you know you are about to change lives! High-five your team and get...
...Description Now offering Daily Pay for select positions! Arcadia Home Care & Staffing, part of the Addus Homecare family of companies... ...to our employees. Arcadia has immediate need for Home Health Aides (HHA) / Caregiver throughout Michigan! We are offering...
Overview Employer Industry: Higher Education and Athletics Why consider this job opportunity: Salary up to $81,500 Generous 75% tuition discount for employees, spouses, and children Comprehensive benefits package including medical, dental, and vision ...
...other. Interested in joining our Team?DescriptionTeam Fishel,the Best Choice in Utility Construction since 1936, is hiringFiber Optic Cable Blowing/ Cable Pulling Crew Leadersin the South Bend, Indianaarea.The purpose of aCrew Leaderat Team Fishel is to manage the...
...Electric is seeking to hire Summer 2026 Engineering Interns within our Energy Engineering team. This position will be a rotational program internship, giving an opportunity to work in Energy Engineering and 1 or more of the specified areas below. What do you get...