招聘公告 · 职位检索 · 央国企/事业单位/名企

大模型算法工程师(行程规划方向)

单位:携程集团类别:AI & BI类型:社招地点:上海更新:2026-09-21

岗位信息

招聘单位携程集团
部门/事业群Content
工作地点上海
官方更新时间2025-10-28

任职要求

<p>【岗位职责】</p><p>1、参与构建旅游领域AI Agent系统,主导旅游垂类模型的训练与调优,解决多城市、多约束条件下的动态路线优化问题。</p><p>2、设计强化学习(RL)框架,结合PPO、GRPO等算法实现,在行程规划领域达到优于顶尖LLM或Agent能力。</p><p>3、构建端到端模型评测体系:设计多维度评估指标,研发大模型与传统运筹算法的融合架构。</p><p>4、开发个性化推荐能力,整合文本、POI、用户行为数据,推动生成式AI在旅游场景的落地应用。</p><p>【任职要求】</p><p>基础要求:</p><p>1、计算机科学、应用数学或运筹学硕士及以上学历,1年以上大模型全流程实战项目经验(数据构建→训练→评测→部署)。</p><p>2、深度掌握模型训练与评测技术:</p><p>3、精通大模型微调技术及分布式训练框架(DeepSpeed、Megatron)</p><p>4、具备强化学习实战经验,熟悉PPO、GRPO等算法在决策优化场景的应用</p><p>5、掌握模型评测方法论(自动指标+人工评估)及A/B测试设计</p><p>优先条件:</p><p>1、RL项目经验:有基于强化学习的决策优化系统(如路径规划)落地经验者优先</p><p>2、大模型训练经验:主导过模型的训练/微调/评测全流程者优先</p><p><br></p><p>Job Responsibilities</p><ol><li>Participate in building the AI Agent system in the tourism field, lead the training and tuning of tourism vertical models, and solve dynamic route optimization problems under multi-city and multi-constraint conditions.</li><li>Design reinforcement learning (RL) frameworks, implement them with algorithms such as PPO and GRPO, and achieve performance superior to top-tier LLMs or Agents in the itinerary planning field.</li><li>Construct an end-to-end model evaluation system: design multi-dimensional evaluation metrics and develop a fusion architecture of large models and traditional operations research algorithms.</li><li>Develop personalized recommendation capabilities, integrate text, POI, and user behavior data, and promote the practical application of generative AI in tourism scenarios.</li></ol><p>Qualifications</p><p>Basic Requirements:</p><ol><li>Master's degree or above in Computer Science, Applied Mathematics, Operations Research or related majors, with more than 1 year of practical experience in the full process of large model projects (data construction → training → evaluation → deployment).</li><li>Have a deep grasp of model training and evaluation technologies.</li><li>Proficient in large model fine-tuning technologies and distributed training frameworks (DeepSpeed, Megatron).</li><li>Possess practical experience in reinforcement learning, and be familiar with the application of algorithms such as PPO and GRPO in decision optimization scenarios.</li><li>Master model evaluation methodologies (automatic metrics + manual evaluation) and A/B test design.</li></ol><p>Preferred Qualifications:</p><ol><li>RL project experience: Priority is given to candidates with practical experience in deploying decision optimization systems (such as path planning) based on reinforcement learning.</li><li>Large model training experience: Priority is given to candidates who have led the full process of model training/fine-tuning/evaluation.</li></ol><p><br></p><p><br></p>

前往官方投递

提示:投递请认准招聘单位官方招聘官网,谨防中介收费。