Senior ML/MLOps Engineer
- Грейд
- Senior
- Формат
- Удалённо
- Занятость
- Полная
- Категория
- Программист
- Источник
- 💻☕️ Jobs_IT
Прямого контакта в посте нет
Откройте вакансию в Telegram-боте: там весь исходный пост и способ отклика.
Открыть в TelegramОписание
We are looking for a Senior ML/MLOps Engineer 🚀
#vacancy #job #ml #mlops #python #remote #ai #sre #platform
ℹ️ About the Company & Project: A global leader in AI video generation, trusted by over 90% of Fortune 100 companies to transform how teams communicate and create content. As a high-growth Series E unicorn (valued at $4B+ with over $500M raised from premier investors including Accel, Kleiner Perkins, and Nvidia's VC arm), are pushing the boundaries of generative AI.
💼 Responsibilities:
— Design and improve platform systems for model training, evaluation, and production serving
— Build robust infrastructure and tooling to make ML workloads scalable, reliable, and cost-efficient
— Architect the deployment and serving of ML models across research and production environments
— Improve scheduling, monitoring, and debugging for GPU and cloud-based workloads
— Develop internal abstractions, developer tools, and agentic systems to reduce operational overhead
— Drive continuous improvements across observability, automation, reliability, and developer experience (DX)
— Collaborate closely with ML researchers and product engineers to turn pain points into robust platform capabilities
— Contribute to technical direction and make pragmatic architectural trade-offs
🧠 Requirements:
— Strong experience building and operating complex, high-load production systems
— Deep systems mindset: ability to analyze bottlenecks, failure modes, and resource usage
— Solid hands-on experience with Linux, cloud infrastructure, and infrastructure automation
— Extensive experience with Kubernetes (K8s) and operating distributed workloads in production
— Strong coding skills in Python (or similar) for backend systems and tooling
— Proven experience building internal platforms, infrastructure abstractions, or developer tools
— Pragmatic approach to problem-solving with a focus on reliability without over-engineering
— Strong ownership and comfort working in ambiguous environments
— English — B2+
💫 Nice to have:
— Direct experience operating ML infrastructure, GPU clusters, or model serving systems in production
— Familiarity with workflow orchestration systems (e.g., Temporal)
— Experience building LLM-powered or agentic internal tools
— Strong background in observability and debugging distributed systems (Datadog, Prometheus, etc.)
— Hands-on experience with Terraform, GitHub Actions, and CI/CD pipelines
— Experience bridging the gap between research and production engineering
CVs to @vladiskashh
#vacancy #job #ml #mlops #python #remote #ai #sre #platform
ℹ️ About the Company & Project: A global leader in AI video generation, trusted by over 90% of Fortune 100 companies to transform how teams communicate and create content. As a high-growth Series E unicorn (valued at $4B+ with over $500M raised from premier investors including Accel, Kleiner Perkins, and Nvidia's VC arm), are pushing the boundaries of generative AI.
💼 Responsibilities:
— Design and improve platform systems for model training, evaluation, and production serving
— Build robust infrastructure and tooling to make ML workloads scalable, reliable, and cost-efficient
— Architect the deployment and serving of ML models across research and production environments
— Improve scheduling, monitoring, and debugging for GPU and cloud-based workloads
— Develop internal abstractions, developer tools, and agentic systems to reduce operational overhead
— Drive continuous improvements across observability, automation, reliability, and developer experience (DX)
— Collaborate closely with ML researchers and product engineers to turn pain points into robust platform capabilities
— Contribute to technical direction and make pragmatic architectural trade-offs
🧠 Requirements:
— Strong experience building and operating complex, high-load production systems
— Deep systems mindset: ability to analyze bottlenecks, failure modes, and resource usage
— Solid hands-on experience with Linux, cloud infrastructure, and infrastructure automation
— Extensive experience with Kubernetes (K8s) and operating distributed workloads in production
— Strong coding skills in Python (or similar) for backend systems and tooling
— Proven experience building internal platforms, infrastructure abstractions, or developer tools
— Pragmatic approach to problem-solving with a focus on reliability without over-engineering
— Strong ownership and comfort working in ambiguous environments
— English — B2+
💫 Nice to have:
— Direct experience operating ML infrastructure, GPU clusters, or model serving systems in production
— Familiarity with workflow orchestration systems (e.g., Temporal)
— Experience building LLM-powered or agentic internal tools
— Strong background in observability and debugging distributed systems (Datadog, Prometheus, etc.)
— Hands-on experience with Terraform, GitHub Actions, and CI/CD pipelines
— Experience bridging the gap between research and production engineering
CVs to @vladiskashh
- python
- ai
- mlops
- kubernetes
- linux
- cloud
- infrastructure
- terraform
- github
- ci/cd
- datadog
- prometheus
Оценка вакансии
43/100 · минимум информации
- Оценка описания недоступна
- Зарплата не указана
- Компания не названа
- Контакт HR подтверждён
- Стек описан подробно
- Формат работы понятен
Похожие вакансии
Junior+ Middle Python разработчики
AI-балл 70/100JoSpace · РБ
Зарплата: Обсуждается на собеседовании
Удалённоpythonfastapidjango
Senior Golang Developer
AI-балл 54/100Instinctools · EU, Poland, Kazakhstan, Serbia, Georgia
SeniorУдалённоgolangsqlredis
Founding CTO / Technical Lead
AI-балл 56/100Ihsan
Leadgotypescriptpostgresql
Senior/TechLead Flutter разработчик
AI-балл 51/100Ставка - до 2200 (Senior), до 2400 (TechLead), включая НДС
Leadflutterdartjava
Backend-разработчик
AI-балл 40/100backendhttpauthorization