招聘公告 · 职位检索 · 央国企/事业单位/名企

Research Engineer - [Seed Model - Infra - LLM/VLM Inference Optimization (Kernel & Compiler)]

单位:字节跳动类别:研发类型:社招地点:圣何塞更新:2026-09-24

岗位信息

招聘单位字节跳动
工作地点圣何塞
官方更新时间2026-04-09 04:47:35

职位描述

The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

Responsibilities

- Design, implement, and optimize high-performance GPU kernels for large-scale LLM/VLM inference workloads, including attention, GEMM, and other compute- and memory-intensive operators.

- Develop and tune inference kernels in CUDA and Triton, and drive end-to-end performance optimization of production inference systems at scale.

- Conduct in-depth performance analysis and profiling to identify bottlenecks across the inference stack, from kernel level to serving level.

- Collaborate with research and infrastructure teams to land kernel- and compiler-level optimizations in production inference systems.

任职要求

Minimum Qualifications:

- Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.

- Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.

- Hands-on experience in LLM/VLM inference optimization with demonstrated impact on latency, throughput, or serving cost.

- Hands-on experience writing and optimizing GPU kernels in CUDA and/or Triton.

- Deep understanding of GPU architecture (memory hierarchy, occupancy, instruction throughput) with solid optimization experience.

Preferred Qualifications:

- Experience with ML compiler internals (e.g., Triton, MLIR, LLVM).

- Contributions to related open-source projects (e.g., Triton, vLLM, SGLang, FlashAttention, CUTLASS).

- Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).

前往官方投递

提示:投递请认准招聘单位官方招聘官网,谨防中介收费。