At Datadope, we provide intelligent operational analytics and monitoring solutions that help our clients enhance visibility into their services and business processes โ driving productivity, operational efficiency, and ultimately reducing maintenance efforts and costs.
Weโre looking for fearless professionals ready to reach the impossible, because on the path to the impossible we often discover the unexpected.
We are seeking a Senior SRE focused on supporting, implementing, optimising, and automating monitoring and observability solutions across both on-premise and cloud environments.
The ideal candidate will possess strong skills in infrastructure, automation, and observability tool integration, along with problem-solving abilities, teamwork, and effective client communication.
Key Responsibilities
- Implement and optimise monitoring and observability solutions across hybrid infrastructures (on-premise and cloud).
- Deploy and configure tools such as Zabbix, Prometheus, Grafana, Elastic Stack, Loki and others.
- Automate monitoring and infrastructure configuration processes using Ansible, Terraform or similar technologies.
- Ensure stability and performance of the monitoring environment and the monitored systems.
- Implement and manage dynamic and predictive alerting systems through data analysis.
- Develop technical, functional, and executive dashboards integrating multiple data sources in Grafana.
- Manage infrastructure on AWS, Azure, and GCP, including Kubernetes and Docker environments.
- Administer relational and time-series databases such as PostgreSQL, MySQL, InfluxDB and Prometheus.
- Promote SRE best practices to increase automation and streamline routine tasks through scripting and tool development.
- Handle daily support tasks such as troubleshooting incidents, processing change requests, and performing preventive maintenance.
- Develop, maintain, and evaluate metrics to ensure 24/7 support readiness.
Minimum Requirements
- 4+ years of experience in support, automation, and optimisation of critical deployments on large-scale infrastructures.
- Experience in Linux/Unix system administration.
- Proficiency in scripting and automation (Python, Bash, Shell scripting).
- Cloud administration experience: AWS, Azure, Google Cloud, Oracle Cloud.
- Hands-on experience with CI/CD (Jenkins, ArgoCD, GitOps).
- Advanced monitoring and logging with tools such as Prometheus, Thanos, Zabbix, Elastic Stack, Loki, Redis, Metricbeat, and Filebeat.
Desirable Skills
- Infrastructure as Code (Terraform, Ansible, Kustomize).
- Kubernetes administration and container orchestration (Docker, OpenShift, Fleet/OSquery).
- Machine Learning foundations.
- University degree in Computer Science, Telecommunications or related fields.