
Job Title: Site Reliability Engineer (SRE)
Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CD
Experience: +6 YOE.
Location: Costa Rica, Peru, Colombia, and Bolivia.
Mode: Remote.
We at Coforge are hiring Site Reliability Engineer (SRE) (#22323) with the following skill set.
Key Responsibilities
· Design, build, and operate scalable and highly available cloud platforms.
· Ensure reliability, performance, and stability of distributed production systems.
· Implement and maintain Infrastructure as Code using Terraform or similar tools.
· Manage Kubernetes-based and containerized environments.
· Define and operate SLOs, SLIs, error budgets, dashboards, runbooks, and alerting standards.
· Implement observability, monitoring, and incident response practices.
· Participate in on-call rotations and respond to production incidents.
· Collaborate with engineering teams to improve automation, scalability, and platform resilience.
· Conduct postmortem reviews and drive continuous reliability improvements.
Required Skills & Qualifications
· Bachelor’s degree in Computer Science, Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience.
· 6+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps Engineering, Backend Engineering, or Production Engineering.
· Strong software engineering skills in at least one language such as Python, Go, Java, TypeScript, or C#.
· Strong understanding of distributed systems, microservices, APIs, asynchronous processing, queues, databases, caching, retries, idempotency, and failure modes.
· Experience with cloud infrastructure on AWS, Azure, or GCP.
· Experience with Kubernetes, containers, Terraform or similar IaC tooling, CI/CD pipelines, and Linux-based systems.
· Experience with observability tools such as Datadog, Prometheus, Grafana, OpenTelemetry, CloudWatch, New Relic, Splunk, or Sentry.
· Experience defining and operating SLOs, SLIs, error budgets, alerting standards, dashboards, runbooks, and incident response practices.
· Strong communication skills and experience working across cross-functional teams.
Preferred Skills
· Cloud, Kubernetes, Infrastructure, Reliability Engineering, Security, or DevOps certifications.
· Experience in logistics, transportation, final-mile delivery, field-service software, routing, dispatch, or fleet operations.
· Experience working with operational SaaS or marketplace platforms.
· Experience driving automation, platform reliability, and operational excellence initiatives.
Posted On: 14-08-2026
At Coforge, we hire professionals based solely on their skills and qualifications and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.
You will be redirected to the company website to complete your application.




