About this role
About the Team The Recommendation Architecture team is responsible for building up and optimizing our recommendation system's architecture to provide the most stable and best experience for our users. You’ll join a high-impact team focused on optimizing Large Language Models (LLMs) and large-scale recommender models optimization on GPU platforms. You’ll build and scale AI infrastructure that powers state-of-the-art models in production. We are looking for talented individuals to join our team in 2026. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at our Company. Successful candidates must be able to commit to an onboarding date by end of year 2026. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early. Responsibilities - Optimize model performance and memory efficiency on GPU-based systems. - Collaborate with research and infra teams to deploy high-throughput training and inference pipelines. - Develop tools and libraries to accelerate deep learning workloads at scale. - Analyze system performance (e.g., GPU profiling, kernel analysis, throughput tuning). Minimum Qualifications: - Individuals who are completing or have recently completed a Bachelor’s/ Master’s degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline. - Solid programming skills in C++/CUDA/Trition/Python. - Familiarity with GPU architecture and distributed training is highly desirable. Preferred Qualifications: - Experience building production-grade training and inference systems for large-scale models. - Hands-on experience optimizing Large Language Models (LLMs), including memory efficiency, latency, and throughput improvements. - Knowledge of distributed training frameworks (e.g., NCCL, Horovod, DeepSpeed, FSDP) is a plus. - Familiarity with deep learning compiler frameworks such as TVM or LLVM, and understanding of their underlying principles. - Contributions to open-source projects or relevant research publications. By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://jobs.bytedance.com/en/legal/privacy If you have any questions, please reach out to us at apac-earlycareers@bytedance.com
Frequently asked questions
What does a Model Infrastructure Engineer Graduate (Recommendation Architecture) - 2026 Start (BS/MS) at ByteDance do?
About the Team The Recommendation Architecture team is responsible for building up and optimizing our recommendation system's architecture to provide the most stable and best experience for our users. You’ll join a high-impact team focused on optimizing Large Language Models (LLMs) and large-scale r…
How much does a Model Infrastructure Engineer Graduate (Recommendation Architecture) - 2026 Start (BS/MS) at ByteDance pay?
The employer did not list a salary for this role. Most similar Singapore roles publish their band on the job page.
Is this Model Infrastructure Engineer Graduate (Recommendation Architecture) - 2026 Start (BS/MS) role remote, hybrid, or on-site?
The listing is based in Singapore. Check the posting for remote or hybrid options.
How do I apply for this Model Infrastructure Engineer Graduate (Recommendation Architecture) - 2026 Start (BS/MS) role?
You can apply directly on ByteDance's careers page. ApplyLah can tailor your résumé and cover letter to this exact role in seconds first.