About the Job
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. Fo...
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
Key Responsibilities
Design and build unified job submission APIs, CLI, and web UI for all AI workload types (training, inference, fine-tuning) on Kubernetes and Slurm with Firmus AI Factory context (tenant isolation, resource requests, metadata tagging, observability hooks).
Implement comprehensive job metadata models and schemas: track job ID, job type, tenant, user, resource requirements, priority class, timestamps, lineage, execution status.
Integrate authentication/authorization (RBAC) and resource quotas; enforce multi-tenant isolation at submission time across all job types.
Build AI job scheduling and orchestration layer: priority classes, preemption policies, fairness algorithms, resource quota enforcement, and intelligent job routing.
Build the AI Factory template catalog: discovery, parameter validation, and manifest generation for training templates, inference serving templates, and fine-tuning recipes.
Wire job submissions to observability pipeline: inject labels/annotations (job_id, tenant, user, model_name, job_type) so metrics are tagged per-job.
Expose job-level telemetry APIs (GPU metrics, cost accrual, MFU progression for training; latency, throughput, tokenomics for inferencing) for platform telemetry and monitoring.
Extend job submission to handle inference workloads: design inference job specifications (model, batch size, latency SLA, cost constraints); integrate with inference serving APIs.
Coordinate with platform team on observability dashboard integration, with LLM engineers on template design, and with ModelOps on reliability standards.
Required Skills & Abilities
5–7 years of backend engineering experience building production APIs and distributed systems (Python, Go, or Java).
Deep Kubernetes expertise: understand Job controllers, Pod specs, resource requests/limits, RBAC, network policies, debugging.
Hands-on Slurm experience: job submission, resource allocation, job queues, sbatch scripting.
Strong distributed systems knowledge: understand scheduling algorithms, fairness, preemption, resource management.
Strong data modelling: can design clear schemas for job metadata, handle versioning and migrations, ensure backward compatibility.
DevOps mindset: comfortable with observability, logging, tracing, and production troubleshooting.
Experience with streaming APIs and real-time webhooks, and system-level integration patterns.
Job Orchestration & Scheduling: shipped job scheduling or workflow systems at scale; understands job lifecycle, failure modes, and scheduling policies.
Multi-Tenancy Design: can architect fair resource allocation, quota enforcement, pre-emption, and data isolation across job types.
API Design: RESTful or gRPC APIs that are intuitive and extensible; handles versioning gracefully.
Systems Architecture: understands how job submission connects to training, inferencing, observability, cost tracking, and incident response.
Cross-Domain Partnership: works closely with infra team, platform team, LLM engineers; clear handoff points and API contracts.
Unified orchestration adoption increases: teams use the standard job interface rather than bespoke/manual pathways.
Scheduling effectiveness & fairness improves: predictable scheduling under contention with reduced noisy-neighbor impact.
Orchestration reliability stays high: jobs reliably start, run, and complete across K8s/Slurm/inference integrations.
End-to-end workflow automation increases: higher share of workflows complete without human intervention (e.g., train→register→serve).
Interface stability & compatibility remains strong: the orchestration API evolves without breaking users.