Senior Infrastructure Engineer (GPU Platform)
- Грейд
- Senior
- Категория
- DevOps
- Источник
- IT Outstaff Projects
Контакт HR — бесплатно после входа
Войдите через Telegram — покажем, кому писать напрямую. Без резюме и анкет.
Описание
We are looking for a Senior Infrastructure Engineer (GPU Platform)
Requirements:
• Production experience with multi-node GPU training infrastructure
• Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge
• Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
• Deep experience with Kubernetes or Slurm
• Experience with infrastructure automation and observability
• Experience diagnosing issues across training workloads, networking, storage, and GPU hosts
• Strong incident leadership and provider-facing communication skills
• English – Upper-Intermediate or higher
Would be a plus:
• Experience in an AI lab, HPC environment, or specialist GPU cloud
• PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks
• Experience with parallel storage and checkpoint optimization
• Experience working with multi-provider GPU platforms
📩 Send your CV:
cv@talentstoday.com
Requirements:
• Production experience with multi-node GPU training infrastructure
• Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge
• Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
• Deep experience with Kubernetes or Slurm
• Experience with infrastructure automation and observability
• Experience diagnosing issues across training workloads, networking, storage, and GPU hosts
• Strong incident leadership and provider-facing communication skills
• English – Upper-Intermediate or higher
Would be a plus:
• Experience in an AI lab, HPC environment, or specialist GPU cloud
• PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks
• Experience with parallel storage and checkpoint optimization
• Experience working with multi-provider GPU platforms
📩 Send your CV:
cv@talentstoday.com
- linux
- containers
- cuda
- nvidia
- gpu
- nccl
- infiniband
- roce
- rdma
- kubernetes
- slurm
- pytorch
- megatron
- deepspeed
Оценка вакансии
38/100 · минимум информации
- Описание полное
- Зарплата не указана
- Компания не названа
- Контакт без валидации
- Стек описан подробно
- Формат не указан
Похожие вакансии
Lead DevOps Engineer / Head of Infrastructure
AI-балл 89/100Scalable Solutions · Грузия
до 4 000 $
LeadУдалённоlinuxkubernetesdocker
DevOps / Infrastructure Engineer
AI-балл 54/100Flex Databases
Удалённоlinuxwindowsci_cd
Senior Engineer, IT System Operations
AI-балл 64/100Aprio · Makati City, Philippines
SeniorГибридagilebpmndevops
Senior DevOps Engineer – FedRAMP
AI-балл 84/100G2i Inc. · United States
170 000 – 200 000 $
SeniorУдалённоansibleawsbash
TechOps SAAS Engineer
AI-балл 52/100MiddleУдалённоpythongrafanadatadog