招聘公告 · 职位检索 · 央国企/事业单位/名企

Research Engineer - (Seed Model - Infra - Reinforcement Learning (RL) Systems & Infrastructure)

单位:字节跳动类别:研发类型:社招地点:圣何塞更新:2026-09-24

岗位信息

招聘单位字节跳动
工作地点圣何塞
官方更新时间2026-02-26 07:34:52

职位描述

The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

Responsibilities

- Design and build end-to-end reinforcement learning (RL) systems for large-scale models, covering rollout, training, evaluation, and deployment pipelines.

- Develop scalable and fault-tolerant RL infrastructure that operates efficiently under dynamic workloads and heterogeneous compute environments.

- Optimize distributed training performance across GPU clusters, improving throughput, resource utilization, and system stability.

- Collaborate with cross-team researchers on targeted system–algorithm co-design to translate research ideas into robust, production-grade implementations.

- Build tooling, monitoring, and debugging frameworks to ensure reliability and observability of large-scale RL training systems.

任职要求

Minimum Qualifications:

- Strong background in distributed systems, large-scale ML systems, or deep learning infrastructure

- Experience building or optimizing large-scale training systems (e.g., RL, LLM, multimodal models)

- Solid engineering skills in Python/C++ and familiarity with modern ML stacks (PyTorch, distributed training frameworks, etc.)

- Experience with GPU optimization, parallelism strategies, and system-level performance tuning

- Understanding of reinforcement learning workflows (rollout, policy update, evaluation loops)

Preferred Qualifications:

- Experience with large-scale agent systems

- Familiarity with system design under heterogeneous or dynamic workloads

- Exposure to RL + LLM training or post-training pipelines

前往官方投递

提示:投递请认准招聘单位官方招聘官网,谨防中介收费。