What are the responsibilities and job description for the Senior Site Reliability Engineer position at Care IT Services Inc?
Job Details
Title: Senior Site Reliability Engineer
Location: 2 Day hybrid Owings Mills MD Mon & Tue, must be onsite from day 1
Duration: 6 months base contract w/ extensions (long-term project)
Work Authorization: EAD
Offer Rate: $55-$60/HR W2
The Technology Engineering team is looking for an experienced Site Reliability Engineer to join us as we are reimagining the production application and infrastructure management. The team is responsible for engineering scalable and resilient hybrid cloud solutions (both AWS and On-prem). You will be responsible for creating tooling and software that monitors and improves the reliability of our systems. In this role, you will research problems, evaluate modern technologies, create prototypes, develop (integrated process, automation, define standards) observability tooling, and provide SRE consulting on complex projects.
Requires specialized in-depth knowledge and expertise in your own job discipline, Amazon Web Services (AWS) platform and/or other cloud-based platforms and deep experience in integrating related disciplinary knowledge
Works independently, receives minimal guidance
Accountable for work of yourself and others; sets standards around which others will operate
Proactively identifies problems and can present and implement solutions to these problems
Role summary and job responsibilities
Design and implement highly automated systems/services that ensure the availability, reliability, and scalability of infrastructure and applications.
Build and maintain monitoring and alerting to provide timely feedback on the performance and health of systems, network, and applications.
Design and implement automation tools to reduce manual toil, streamline repetitive tasks, and enhance overall operational efficiency.
Design and build Service Level Indicator (SLIs) metrics, including but not limited to Service Level Objectives (SLOs), Error Budget, Burn Rate Alerts
Work closely with development teams to embed reliability best practices into the software development process. Provide mentorship and training to cross-functional teams on SRE principles, encouraging a shared responsibility for the reliability of our services.
Collaborating with our support, operations and engineering teams to investigate and troubleshoot complex problems
Observe and monitor systems to make sure you have the insight into system performance, health, availability and what is happening internally in the system.
Understands what to monitor based on the system(s) you are managing, how the monitoring data is stored, and how to look at the data to make determinations about future actions.
Participates in continuous improvement efforts that span multiple multi-functional domains and informs the generation of new standards
Be a part of an on-call rotation, continuously enhance automation & documentation, and mentor others on the standard methodologies of infrastructure automation to encourage adoption.
Able to overcome differences of opinion and drive team alignment around a specific goal or solution
Holds associates and teams accountable for adhering to practices and policies
Business knowledge
Demonstrates deep knowledge of products/flows within supported businesses
Decomposes the most complex problems into discrete work units.
Identifies non-obvious relationships and anomalies often overlooked by others.
Balances strategic and pragmatic concerns when solving problems.
Makes sound decisions with limited facts or resources.
Makes decisions that are cognizant of the firm s broader business strategy.
Demonstrates deep knowledge of products/flows within the businesses they support.
Articulates broader business concerns and/or regulatory landscape, including key risks and controls (e.g., GDPR, MIFID, SOX).
SRE Skills:
AWS
Monitoring - Grafana Prometheus
Automation -Terraform, Python
Kubernetes
Observability EKS, ECS
CloudWatch
Salary : $55 - $60