Team Introduction: Our Live Recommendation Architecture Team is responsible for building up and optimizing the architecture for live broadcast recommendation system to provide the most stable and best experience for our users. The team is responsible for system stability and high availability, online services and offline data flow performance optimization, solving system bottlenecks, reducing cost overhead, building data and service mid-platform, realizing flexible and scalable high-performance storage and computing systems. We work closely with applied machine learning engineers and build scalable systems to support all kinds of innovative algorithms and techniques.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities
- Optimize model performance and memory efficiency on GPU-based systems for TikTok Live recommendation scenarios
- Deploy and scale high-throughput training and inference pipelines for live broadcast recommendation models
- Develop tools, libraries, and compiler optimizations to accelerate deep learning workloads for live streaming use cases
- Analyze system performance through GPU profiling, kernel analysis, and throughput tuning for real-time workloads
- Build and maintain high-performance computing frameworks and storage systems for live recommendation model training and inference
Minimum Qualifications:
- Individuals who are completing or have recently completed a Bachelor's or Master's degree in Artificial Intelligence, Software Development, Computer Science, Computer Engineering or a related discipline.
- Solid programming skills in C++/CUDA/Trition/Python.
- Familiarity with GPU architecture and distributed training is highly desirable.
- Experience programming in at least one of the following programming languages: C, C++, Java or Golang.
- Effective communication skills and a sense of ownership and drive.
Preferred Qualifications:
- Experience building production-grade training and inference systems for large-scale models, especially for recommendation or live streaming scenarios.
- Hands-on experience optimizing recommendation models, including memory efficiency, latency, and throughput improvements.
- Knowledge of distributed training frameworks (e.g., NCCL, Horovod, DeepSpeed, FSDP) is a plus.
- Familiarity with deep learning compiler frameworks such as TVM or LLVM, and understanding of their underlying principles.
- Experience in livestream or real-time system related areas is a plus.
- Contributions to open-source projects or relevant research publications.
- Agile, quick self learner, highly self-motivated with strong sense of product ownership and creative problem solver.
- Good collaborator and team player, comfortable working in a fast moving, culturally diverse and globally distributed team environment.