招聘公告 · 职位检索 · 央国企/事业单位/名企

Production System Engineer

单位:字节跳动类别:研发类型:社招地点:伦敦更新:2026-09-24

岗位信息

招聘单位字节跳动
工作地点伦敦
官方更新时间2025-10-17 21:08:17

职位描述

About the Team

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across ByteDance’s self-built data centers in the United States and Europe.

Our scope covers new hardware introduction, data center delivery, production operations, hardware maintenance, configuration and firmware changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

- Hardware New Product Introduction (NPI)

- Server and data center operations

- Field maintenance and infrastructure management

- Hardware vendors and service providers

- Supply chain and asset management

- Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

Role Overview

We are looking for a hands-on Production Systems Engineer with a strong foundation in Linux systems, server infrastructure, automation, and production operations. This role is open to engineers across a range of experience levels, from early-career engineers with strong technical fundamentals to experienced infrastructure engineers who can take ownership of complex systems and large-scale initiatives.

The scope and level of ownership will grow with experience, ranging from hands-on infrastructure engineering and automation development to leading complex global infrastructure initiatives across organizational boundaries

Responsibilities

- Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.

- Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.

- Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.

- GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements. Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.

- Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.

- Technical Documentation: Create and improve technical documentation, standard operating procedures, troubleshooting guides, and internal knowledge bases.

- Cross-functional Collaboration: Work with infrastructure engineers, hardware teams, data center operations, platform developers, supply chain teams, and other stakeholders on global infrastructure projects.

任职要求

Minimum Qualification(s)

- Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.

- 2 years of experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands-on project experience.

- Strong foundation in Linux system administration and troubleshooting, with an understanding of basic server architecture, operating systems, storage, networking, and hardware management concepts.

- Programming or scripting experience in Python, Bash, Go, or another modern programming language, with the ability to develop tools or automation for infrastructure or operational tasks.

- Hands-on experience troubleshooting system, hardware, storage, networking, or performance-related issues in Linux-based environments.

- Strong analytical and problem-solving skills, with the ability to learn unfamiliar technologies quickly and investigate complex technical issues in a structured manner.

- Good communication and collaboration skills, with the ability to work effectively with engineers and cross-functional stakeholders across different technical domains and regions.

Preferred Qualification(s)

- Familiarity with technologies such as BIOS/UEFI, BMC, firmware, PCIe, NVMe, NICs, or hardware telemetry.

- Proficiency in Python, Go, Bash, or another programming language for production-grade infrastructure automation, including experience designing, building, or maintaining tools and platforms used in large-scale production environments.

- Deep knowledge of Linux administration and troubleshooting, preferably Debian or Ubuntu, combined with strong understanding of server architecture and management technologies such as kernels, drivers, BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, hardware telemetry, and failure diagnostics.

- Proven hands-on experience introducing and productionizing large-scale GPU infrastructure, including ownership of hardware NPI or fleet onboarding across qualification, system integration, deployment, production validation, operational handoff, and post-launch reliability.

- Strong understanding of distributed AI workload behavior and performance analysis, including collective communication, multi-node training, inference serving, GPU scheduling, checkpointing, workload-related bottlenecks, DCGM, NCCL testing, CUDA profiling, and network-fabric telemetry.

- Experience building and operating monitoring, telemetry, hardware management, or automated remediation platforms at substantial scale, with measurable improvements in fleet availability, deployment efficiency, incident reduction, operational efficiency, or reliability.

- Experience working directly with OEMs, ODMs, component suppliers, or GPU platform vendors throughout qualification, technical escalation, root-cause analysis, and corrective-action processes.

- Experience with one or more advanced infrastructure technologies or engineering areas, such as containerisation and orchestration (e.g., Docker, Kubernetes), infrastructure automation frameworks (e.g., Ansible), AI-powered automation, AI agents, Large Language Models, Retrieval-Augmented Generation (RAG), open-source infrastructure projects, technical publications, patents, or relevant industry standards.

前往官方投递

提示:投递请认准招聘单位官方招聘官网,谨防中介收费。