• FAQ
  • Docs
  • Jobs
  • Contact
  • FAQ
  • Docs
  • Jobs
  • Contact
Get Started

Copyright © 2026 SpeedyApply. All rights reserved.

GitHubLinkedInDiscordDiscordTikTok

Resources

ContactFAQBlogAffiliate

Legal

Privacy PolicyTerms of Service

Chrome Web StoreEdge Add-onsAdd-ons for Firefox
  • FAQ
  • Docs
  • Jobs
  • Contact
  • FAQ
  • Docs
  • Jobs
  • Contact
Get Started

Research Scientist Graduate - Conversational AI- 2027 Start - PhD

TikTok • Seattle, WA

T
TikTok

Research Scientist Graduate - Conversational AI- 2027 Start - PhD

AI / MLNew GradSeattle, WA45 days ago

We build the next-generation unified Agent system for TikTok's global e-commerce customer service — running in 30+ languages across one of the largest e-commerce surfaces on the internet.

Our north star is a self-evolving Agent: post-training, harness, memory / context engineering, tools, and evaluation form one closed loop, and every served conversation becomes the next iteration's training / eval / retrieval / skill-induction signal. This loop is already running in production — cases are mined, root-caused, turned into constrained candidates, replayed against frozen regression sets, and shipped behind guardrails.

Two things make this team different from most "LLM application" work:

  • We build the agent runtime itself — Codex / Claude-Code-class — not prompts on top of a vendor API.
  • Evaluation and experimentation are first-class systems, not an afterthought. A self-improving loop optimizes whatever signal you give it, so the hardest and most valuable engineering here is making the judgment trustworthy — not just making the model change. As a new grad, you'll own a real end-to-end piece from day one and ship it to production.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.

Responsibilities:

  • Agent runtime (harness / agent loop). Orchestrate skills, tools, and context; implement loop control & intervention, progressive disclosure, and behavior-level guardrails. Build the production safety layer — pre-flight budgets and timeout truncation, serve-time gates, shadow / swap-in answer delivery, and safe fallback paths.
  • Context & memory for long multi-turn agents. Agentic memory (structured note-taking), context compaction / summarization, context editing / observation masking, and just-in-time (retrieve-then-load) retrieval. Treat context as an evolving, itemized playbook — with structured diffs and a deterministic curator — rather than an ever-growing prompt.
  • Post-training & the data flywheel. SFT / DPO / RL to internalize rules into weights (so the prompt gets shorter, not longer), plus distillation to smaller serving models. Turn served conversations into training / eval / retrieval signals.
  • Tools, Skills, and MCP. Tools-as-APIs, connectors, skill / tool search for large inventories, and skill-library governance — description conflicts, trigger evals, cross-skill mis-fire matrices, and on-demand loading instead of dumping every definition into context.
  • Evaluation you can bet a launch on. LLM-as-judge with human-agreement calibration; statistical rigor — paired comparison, confidence intervals, repeated sampling, pass^k; held-out and time-rolling eval splits with overfitting alarms; cascaded scoring and cross-family judge panels to make evaluation affordable at scale.
  • The self-evolving loop. Case mining → automatic root-cause → constrained candidate generation → replay verification against frozen regression sets → canary → flywheel. Build the plumbing that makes it auditable: candidate registry with exact runtime read-back, change lineage, and an archive of rejected candidates you can sample from next round.
  • Online experimentation & causal readout. Shadow / canary / A-B, non-inferiority gates, traffic-split health, metric definitions that survive scrutiny, and off-policy counterfactual evaluation where live A/B isn't possible.
  • Safety & anti-gaming. Keep the evaluator and the release gate outside the loop that edits the system; maintain never-optimized anchor sets; monitor full execution traces rather than final answers alone; pair every quality objective with a cost-side constraint.
  • Own one high-leverage end-to-end surface and ship it to production across 30+ languages, measured on real business metrics (CSAT, resolution / containment rate).

Minimum Qualifications:

  • Individuals who are completing or have recently completed a PhD degree in CS/AI/Math/Quantitative degree or a related discipline.
  • Strong Python plus one of C++ / Go / Rust / Java
  • Solid ML / DL / NLP fundamentals, with genuine hands-on experience with LLMs or agents (coursework, research, internship, competition, open-source, or a serious side project)
  • Basic statistical literacy — you can compute a confidence interval, explain what a p-value does and doesn't mean, and tell the difference between "the number went up" and "the system got better"
  • Able to read a paper or an engineering blog and turn it into working code

Preferred Qualifications

  • Have built the runtime, not just called an API — even at research / hobby / competition scale: your own agent loop / harness, a memory / context-management system, a RAG or tool-use agent, or a fine-tuned / post-trained model
  • Post-training: SFT / DPO / RLHF / RLAIF / RLVR, reward modeling, reward hacking and how to defend against it
  • Agent systems: harness, context engineering, MCP / Skills, sub-agents, tool search
  • Evaluation & experimentation: LLM-as-judge and judge calibration, pass^k, regression suites, A/B and non-inferiority testing, off-policy evaluation
  • Self-improving / evolutionary systems: evolutionary program search, candidate archives and parent sampling, automatic prompt / context optimization, multi-objective (Pareto) selection and credit assignment
  • Inference & serving: vLLM / TensorRT-LLM, MoE, KV / prompt caching, and the cost engineering that comes with it
  • Publications (for PhD), strong competition results (ACM-ICPC / Kaggle / ML competitions), or notable open-source contributions

469 results

  • T
    TikTokMachine Learning Engineer Graduate - E-Commerce Recommendation Video - 2027 Start
    San Jose, CA
    AI / MLNew Grad
    1 day ago
  • T
    TikTokMachine Learning Engineer Intern - E-Commerce Recommendation Mall - 2027 Start - PhD
    San Jose, CA
    AI / MLInternship
    3 days ago
  • T
    TikTokAI Model and Agent Operation Project Intern - Business Integrity - 2027 Start
    Singapore
    AI / MLInternship
    5 days ago
  • T
    TikTokMachine Learning Engineer Graduate - E-Commerce Content Recommendation - Generative & Large Recommendation Model - 2027 Start - PhD
    Seattle, WA
    AI / MLNew Grad
    5 days ago
  • T
    TikTokMachine Learning Engineer Graduate - E-Commerce Content Recommendation - Generative & Large Recommendation Model - 2027 Start - PhD
    San Jose, CA
    AI / MLNew Grad
    5 days ago
  • T
    TikTokBackend Software Engineer Project Intern - Trust & Safety - 2026 Start
    Singapore
    Software EngineeringInternship
    11 days ago
  • T
    TikTokSoftware Development Engineer in Test - Multiple Positions
    San Jose, CA• $187.1K–$216K/yr
    Software EngineeringNew Grad
    11 days ago
  • T
    TikTokAI Model Data Operation Graduate - TikTok ADSO - 2027 Start
    Ho Chi Minh City, Vietnam
    AI / MLNew Grad
    12 days ago
  • T
    TikTokAI Model Data Operation Graduate - TikTok ADSO - 2027 Start
    Bangkok, Thailand
    AI / MLNew Grad
    12 days ago
  • T
    TikTokAI Model Data Operation Graduate - TikTok ADSO - 2027 Start
    Jakarta, Indonesia
    AI / MLNew Grad
    12 days ago
  • T
    TikTokAI Model Data Operation Graduate - TikTok ADSO - 2027 Start
    Kuala Lumpur, Malaysia
    AI / MLNew Grad
    12 days ago
  • T
    TikTokBackend Software Engineer Project Intern - Data Trust and Safety - 2026 Start
    Sydney, Australia
    Software EngineeringInternship
    13 days ago
  • T
    TikTokAI Data Project Intern - Eco & Social Creation - 2027 Start
    Singapore
    AI / MLInternship
    13 days ago
  • T
    TikTokData Engineer Graduate - Data Platfrom TikTok BP - 2027 Start
    San Jose, CA
    AI / MLNew Grad
    15 days ago
  • T
    TikTokAI Model Evaluation & Quality Strategy Specialist Graduate - AI Data Service Operations - 2027 Start
    Jakarta, Indonesia
    AI / MLNew Grad
    21 days ago
  • T
    TikTokAI Model Evaluation & Quality Strategy Specialist Graduate - AI Data Service Operations - 2027 Start
    Kuala Lumpur, Malaysia
    AI / MLNew Grad
    21 days ago
  • T
    TikTokAI Model Evaluation & Quality Strategy Specialist Graduate - AI Data Service Operations - 2027 Start
    Bangkok, Thailand
    AI / MLNew Grad
    21 days ago
  • T
    TikTokAI Model Evaluation & Quality Strategy Specialist Graduate - AI Data Service Operations - 2027 Start
    Ho Chi Minh City, Vietnam
    AI / MLNew Grad
    21 days ago
  • T
    TikTokMachine Learning Engineer Graduate - E-Commerce Knowledge Graph - 2027 Start
    San Jose, CA
    AI / MLNew Grad
    22 days ago
  • T
    TikTokAI Data Strategy & Model Operations Specialist Graduate - AI Data Service Operations - 2027 Start
    Singapore
    AI / MLNew Grad
    23 days ago
1 / 24Next

Logos provided by Logo.dev