Notifications

Loading notifications...
Elevenlabs Cover

HPC Infrastructure Engineer - GPU Clusters

Elevenlabs
Remote Full Time Negotiable 6 days ago

About the Job

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and ...

We have expanded from voice into three main platforms:

ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.

Key Responsibilities

Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.
This is a builder-operator role with real breadth: one week you're writing automation that eliminates a whole class of manual work, the next you're benchmarking a new provider's InfiniBand fabric or on-site bringing new hardware online. You'll have unusual scope and autonomy - we're a lean team where decisions are made by the people closest to the problem.
Operate and improve our GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, capacity planning
Build automation that keeps the fleet healthy without human intervention — node health checks, automated draining and remediation, burn-in pipelines for new capacity
Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE)
Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast
Build and maintain high-performance storage for datasets and checkpoints
Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs — and fix the class of problem, not just the instance
Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs
Hands-on hardware work when it's needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors
Keep clusters secure by default: access control, network isolation, secrets

Required Skills & Abilities

Have run large-scale Linux server or GPU environments in production and enjoy both building and operating
Know the NVIDIA stack well — drivers, CUDA, NCCL, DCGM — or have deep systems experience and learn hardware stacks fast
Are comfortable with bare-metal environments, server hardware, and high-speed networking
Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform
Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong
Like owning real scope end to end and being the person others rely on
Don't consider any task above or beneath you — datacenter trips included
Experience supporting ML training workloads from the infra side (distributed training failure modes, checkpointing patterns)
Experience evaluating and working with GPU cloud providers
Parallel filesystems (WEKA, VAST, etc) or large-scale object storage
BMC/IPMI/Redfish automation, PXE provisioning at scale
Power and cooling awareness for dense GPU deployments

Apply now

Please let Elevenlabs know you found this job on Job Vista. This helps us grow!