Job Responsibilities: Ensure the stability, reliability, and high availability of the company’s overseas production environment, continuously improving system availability and service quality. Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability. Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency. Participate in on-call rotations, responding promptly to production incidents to safeguard business operations. Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency. Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness. Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency. Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development. Requirements: Bachelor’s degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer. Proficiency in at least one programming language (e.g., Python, Go, Java, or C++), with strong software development and automation skills. Familiarity with cloud computing services; experience with multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS, GCP) is a plus. Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures. Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning. Experience with system monitoring and observability tools (e.g., Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred. Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus. Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.