Notifications

Loading notifications...
Elevenlabs Cover

Research Engineer - Inference

Elevenlabs
Remote Full Time Negotiable 12 days ago

About the Job

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and ...

We have expanded from voice into three main platforms:

ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.

Key Responsibilities

We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy:
Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.
Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.
Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.

Required Skills & Abilities

We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:
Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.

Apply now

Please let Elevenlabs know you found this job on Job Vista. This helps us grow!