My client is a growing cloud service company. Having a passionate and goal-oriented team in Taiwan, looking for proactive talents.
Responsibilities
Design and implement end-to-end observability solutions, including data collection, monitoring, visualization, and alerting.
Develop and integrate monitoring solutions using platforms such as Prometheus, Grafana, Datadog, and OpenTelemetry.
Design monitoring dashboards and alerting strategies that provide actionable insights into system performance and reliability.
Architect solutions across VMs, Kubernetes, and serverless environments.
Work with major cloud platforms including AWS, GCP, and Azure.
Apply knowledge of cloud networking, CDN, WAF, and security best practices to ensure reliable and secure deployments.
Develop automation and integration tools using Python and Bash.
Work with Docker and CI/CD to support efficient deployment and integration workflows.
Collaborate directly with international clients and internal engineering teams to understand technical requirements and deliver practical, scalable solutions.
Requirements
5+ years of experience in DevOps, SRE and Solution Architecture with AWS, GCP or Azure.
Strong hands-on experience in DevOps/SRE, monitoring, alerting, and system reliability.
In-depth experience with at least two of the following: Prometheus, Grafana, Datadog andOpenTelemetry.
Solid experience with K8s, cloud infrastructure, VMs and serverless environments.
Knowledge of cloud networking, CDN, WAF, and security practices.
Proficiency in Python and Bash scripting for automation.
Strong English communication skills, with the ability to work directly with international clients and cross-functional teams.