Notifications

Loading notifications...
Firmus Cover

Senior Software Engineer, AI Job Orchestration

Firmus
Sydney, AU Full Time Negotiable Aug 2, 2026

About the Job

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. Fo...

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

Key Responsibilities

Design and build unified job submission APIs, CLI, and web UI for all AI workload types (training, inference, fine-tuning) on Kubernetes and Slurm with Firmus AI Factory context (tenant isolation, resource requests, metadata tagging, observability hooks).
Implement comprehensive job metadata models and schemas: track job ID, job type, tenant, user, resource requirements, priority class, timestamps, lineage, execution status.
Integrate authentication/authorization (RBAC) and resource quotas; enforce multi-tenant isolation at submission time across all job types.
Build AI job scheduling and orchestration layer: priority classes, preemption policies, fairness algorithms, resource quota enforcement, and intelligent job routing.
Build the AI Factory template catalog: discovery, parameter validation, and manifest generation for training templates, inference serving templates, and fine-tuning recipes.
Wire job submissions to observability pipeline: inject labels/annotations (job_id, tenant, user, model_name, job_type) so metrics are tagged per-job.
Expose job-level telemetry APIs (GPU metrics, cost accrual, MFU progression for training; latency, throughput, tokenomics for inferencing) for platform telemetry and monitoring.
Extend job submission to handle inference workloads: design inference job specifications (model, batch size, latency SLA, cost constraints); integrate with inference serving APIs.
Coordinate with platform team on observability dashboard integration, with LLM engineers on template design, and with ModelOps on reliability standards. 

Required Skills & Abilities

5–7 years of backend engineering experience building production APIs and distributed systems (Python, Go, or Java).
Deep Kubernetes expertise: understand Job controllers, Pod specs, resource requests/limits, RBAC, network policies, debugging.
Hands-on Slurm experience: job submission, resource allocation, job queues, sbatch scripting.
Strong distributed systems knowledge: understand scheduling algorithms, fairness, preemption, resource management.
Strong data modelling: can design clear schemas for job metadata, handle versioning and migrations, ensure backward compatibility.
DevOps mindset: comfortable with observability, logging, tracing, and production troubleshooting.
Experience with streaming APIs and real-time webhooks, and system-level integration patterns. 
Job Orchestration & Scheduling: shipped job scheduling or workflow systems at scale; understands job lifecycle, failure modes, and scheduling policies.
Multi-Tenancy Design: can architect fair resource allocation, quota enforcement, pre-emption, and data isolation across job types.
API Design: RESTful or gRPC APIs that are intuitive and extensible; handles versioning gracefully.
Systems Architecture: understands how job submission connects to training, inferencing, observability, cost tracking, and incident response.
Cross-Domain Partnership: works closely with infra team, platform team, LLM engineers; clear handoff points and API contracts. 
Unified orchestration adoption increases: teams use the standard job interface rather than bespoke/manual pathways.
Scheduling effectiveness & fairness improves: predictable scheduling under contention with reduced noisy-neighbor impact.
Orchestration reliability stays high: jobs reliably start, run, and complete across K8s/Slurm/inference integrations.
End-to-end workflow automation increases: higher share of workflows complete without human intervention (e.g., train→register→serve).
Interface stability & compatibility remains strong: the orchestration API evolves without breaking users. 

Apply now

Please let Firmus know you found this job on Job Vista. This helps us grow!