About this role
About the Team The DCS team supports the company's fast growth by building and operating hyperscale datacenters. The team manages the end to end lifecycle of server fleet, providing cloud solutions and various infrastructure services ensuring that they are scalable and are reliable. - Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds. - Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure. - Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards. - Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability. - Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement. Minimum Qualifications: - Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field. - 2+ years of experience in Linux operations, SRE, or DevOps; - Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation. - Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills. - Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes. - Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset. Preferred Qualifications: - Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms. - Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU. - Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and contained, with a solid understanding of isolation mechanisms like cgroups and namespaces. - Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines. - Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization. - Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred. - Experience operating large-scale production environments is a strong plus.
Frequently asked questions
What does a Cloud Site Relibility Engineer - DCS at ByteDance do?
About the Team The DCS team supports the company's fast growth by building and operating hyperscale datacenters. The team manages the end to end lifecycle of server fleet, providing cloud solutions and various infrastructure services ensuring that they are scalable and are reliable. - Design, build,…
How much does a Cloud Site Relibility Engineer - DCS at ByteDance pay?
The employer did not list a salary for this role. Most similar Singapore roles publish their band on the job page.
Is this Cloud Site Relibility Engineer - DCS role remote, hybrid, or on-site?
The listing is based in Singapore. Check the posting for remote or hybrid options.
How do I apply for this Cloud Site Relibility Engineer - DCS role?
You can apply directly on ByteDance's careers page. ApplyLah can tailor your résumé and cover letter to this exact role in seconds first.