Senior AI Engineer (Kubernetes & Customised Scheduler)
✨ AI Summary
Firmus Technologies, a global leader in efficient AI infrastructure, is hiring a Senior AI Engineer to build a proprietary Kubernetes-based job scheduler for massive-scale AI factories. The role focuses on developing topology-aware and resource-aware scheduling policies for NVIDIA NVL72 GB300-scale GPU deployments. Key responsibilities include designing custom controllers, admission webhooks, and scheduling plugins to optimize workload placement and grid integration. The position is based in Singapore and requires strong expertise in Kubernetes, platform engineering, and AI workload orchestration.
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why Firmus?
As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.
We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.
Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.
What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.
Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
ROLE SUMMARY
The Senior AI Engineer will be a core builder of the AI & Applications team’s Model-to-Grid product, designing, developing, and operating the Kubernetes-based workload orchestration layer for massive-scale AI factories. The role has a strong software and platform engineering focus: building a proprietary job scheduler that enables efficient, reliable, and policy-driven execution of training, fine-tuning, inference, batch, benchmarking, and agentic workloads across cutting-edge GPU systems, including NVIDIA NVL72 GB300-scale deployments and future-generation platforms such as VR200.
The proprietary scheduler is a central product differentiator. It must be network-topology aware and AI-factory-resource aware: making placement, queueing, prioritization, admission, and execution decisions based not only on nominal GPU availability, but also on GPU and NVLink/NVSwitch topology, node and fault-domain boundaries, RDMA and fabric health, storage locality and throughput, workload characteristics, capacity, power, thermal state, maintenance activity, and other operational constraints.
The resulting capability improves outcomes in two directions. For AI users, it provides better workload placement, lower queue times, higher GPU utilization, stronger job-success rates, improved end-to-end throughput, and faster time-to-results. For AI-factory and grid operators, it serves as a technical shock absorber by making demand more observable, controllable, schedulable, and responsive to infrastructure availability, system health, power, thermal, capacity, and operational conditions.
The role will work closely with other AI engineers, inference and optimization engineers, the Model-to-Grid product and program lead, Platform, Infrastructure, Security, and operations teams. It will turn scheduler product requirements into robust production capabilities, integrating Kubernetes, custom controllers, APIs, observability, automation, workload recipes, benchmark signals, and supporting platform services. The work will align with the architectural direction of NVIDIA DSX OS and AI Factory Blueprint concepts - co-designed, resilient, multi-tenant AI-factory operations, while delivering the organization’s proprietary scheduler intelligence and differentiated Model-to-Grid capabilities.
KEY RESPONSIBILITIES
- Design, build, operate, and continuously improve the proprietary Kubernetes-native job scheduling platform for AI workloads.
- Develop scheduler architecture, custom resource definitions, Kubernetes controllers, admission webhooks, scheduling plugins, APIs, CLI tools, and automation required to support workload submission, placement, execution, monitoring, and recovery.
- Define and implement scheduling policies for training, fine-tuning, inference, benchmarking, batch, data-processing, and agentic workloads.
- Build topology-aware placement mechanisms that account for GPU locality, NVLink/NVSwitch domains, node topology, NUMA affinity, NIC placement, RDMA paths, network-fabric topology, storage locality, and fault-domain boundaries.
- Build AI-factory resource-aware scheduling mechanisms that account for cluster capacity, node health, fabric condition, storage performance, GPU availability, maintenance windows, software compatibility, capacity reservations, power limits, thermal conditions, and operational constraints.
- Implement workload-control capabilities including admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, workload aging, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery.
- Develop workload demand-shaping and grid-aware controls that allow eligible workloads to be delayed, paced, prioritized, rescheduled, right-sized, or placed differently in response to capacity, health, power, thermal, maintenance, or other AI-factory operating signals.
- Integrate the scheduler with Kubernetes control-plane capabilities, GPU device plugins, node-feature discovery, network and storage services, observability systems, and policy engines.
- Integrate, where appropriate, with Slurm, Slinky, vCluster, Kueue, KAI, Volcano, YuniKorn, or related systems to deliver a coherent workload-management experience across AI and HPC use cases.
- Build a stable developer and user experience for workload submission and management, including APIs, SDKs, CLI workflows, templates, workload definitions, job-status visibility, event streams, scheduling explanations, and self-service troubleshooting capabilities.
- Work with AI and inference engineers to encode validated model and workload recipes into scheduler-aware templates, including requirements for model size, precision, distributed parallelism, GPU count, topology, network, storage, runtime, benchmark target, and expected resource profile.
- Enable Model-to-Grid benchmarking by exposing scheduler, placement, resource, and workload-lifecycle data for end-to-end analysis across models, runtimes, GPUs, network, storage, power, thermals, and application performance.
- Build observability into the scheduling platform, including queue depth, scheduling latency, admission outcomes, placement decisions, resource fragmentation, topology quality, GPU utilization, job lifecycle, failure reasons, preemption events, retry behavior, power and thermal signals, and workload performance.
- Develop actionable scheduler explanations and operator views that show why a workload was queued, admitted, placed, deferred, preempted, or rescheduled and what changes could improve execution outcomes.
- Establish performance testing, load testing, scale testing, chaos testing, fault recovery testing, and regression testing for scheduler components and Kubernetes platform integrations.
- Operate production Kubernetes clusters and associated scheduling services with strong reliability, security, change-management, backup, recovery, and incident-response practices.
- Build CI/CD and GitOps workflows for scheduler code, configuration, CRDs, controllers, policy definitions, workload templates, and cluster-service releases.
- Implement secure multi-tenancy and resource-governance controls, including namespace and tenant isolation, RBAC, quotas, workload identity, policy enforcement, audit trails, and controlled access to GPU, data, and infrastructure resources.
- Partner with platform, infrastructure, networking, storage, and operations teams to ensure scheduler decisions reflect actual infrastructure capabilities, constraints, health, performance, maintenance, and capacity conditions.
- Partner with the Model-to-Grid product and program lead to define scheduler feature scope, milestone plans, acceptance criteria, release gates, benchmark targets, documentation, and adoption measures.
- Contribute to technical roadmaps, architecture decisions, design reviews, release readiness, incident postmortems, and continuous improvement of Model-to-Grid and AI-factory operations.
SKILLS AND EXPERIENCE
- 5+ years of experience in DevOps, site reliability engineering, platform engineering, distributed systems, cloud infrastructure, or comparable software and systems-engineering roles.
- Deep practical expertise in Kubernetes architecture and operations, including control planes, scheduling, controllers, operators, CRDs, admission controllers, APIs, networking, storage, security, and multi-tenancy.
- Demonstrated experience building, extending, or deeply integrating workload schedulers, resource managers, or orchestration systems, such as Kubernetes Scheduler Framework, Kueue, KAI, Volcano, YuniKorn, Slurm, Slinky, Run:ai, or proprietary implementations.
- Strong programming skills in Go, with Python proficiency for automation, integration, tooling, and performance analysis.
- Experience building production Kubernetes controllers, operators, scheduling plugins, admission webhooks, REST or gRPC APIs, CLI tools, and event-driven distributed services.
- Strong understanding of GPU-accelerated AI/ML workloads, including distributed training, fine-tuning, model serving, inference, benchmarking, and batch processing.
- Practical experience with GPU scheduling and resource allocation, including GPU partitioning, MIG where applicable, GPU affinity, multi-GPU workloads, gang scheduling, and topology-aware placement.
- Understanding of GPU system topology and high-performance AI networking, including NVLink, NVSwitch, PCIe, NUMA, RDMA, RoCEv2, NIC affinity, collective communication, and network contention.
- Familiarity with large-scale NVIDIA GPU architectures, including NVL72 GB300-class systems and future-generation high-density GPU platforms, and their implications for cluster scheduling, workload placement, network topology, operations, and fault handling.
- Experience working with storage and data-path considerations for AI workloads, including high-throughput shared file systems, object storage, caching, data locality, checkpointing, and storage-performance constraints.
- Understanding of AI-factory or data-center operations, including capacity planning, workload demand forecasting, power-aware scheduling, energy efficiency, thermal constraints, maintenance coordination, system health, resiliency, and infrastructure telemetry.
- Experience with observability and performance analysis using tools and concepts such as Prometheus, Grafana, OpenTelemetry, logs, traces, metrics, DCGM, GPU telemetry, network telemetry, workload profiling, and SLOs.
- Experience with CI/CD, GitOps, infrastructure as code, policy as code, and safe production rollout practices using tools such as GitHub Actions, GitLab CI, Argo CD, Flux, Terraform, Helm, Kustomize, OPA, or Kyverno.
- Familiarity with cloud-native security practices, including container security, image signing, software supply-chain controls, secrets management, RBAC, workload identity, network policies, runtime security, and vulnerability remediation.
- Strong troubleshooting skills across Kubernetes, distributed systems, GPU workloads, networking, storage, operating systems, and application-level job execution.
KEY COMPETENCIES
- Kubernetes scheduler and controller engineering.
- Distributed systems architecture, reliability, and performance engineering.
- GPU workload orchestration for training, inference, benchmarking, and AI applications.
- Network-topology-aware and AI-factory-resource-aware workload placement.
- Deep understanding of the relationship between workload intent, model configuration, runtime behavior, GPU topology, network, storage, and end-to-end performance.
- Model-to-Grid engineering: connecting AI demand, scheduler decisions, benchmark data, infrastructure telemetry, and AI-factory operating constraints.
- Multi-objective scheduling design across user throughput, latency, time-to-results, utilization, fairness, cost, power, energy, resilience, and grid or operator constraints.
- Developer experience for AI workload submission, templates, APIs, observability, and troubleshooting.
- Production ownership, incident response, change control, testing discipline, and operational excellence.
- Security-conscious multi-tenant platform engineering.
- Pragmatic architecture and technical decision-making under evolving hardware, software, and operational constraints.
- Cross-functional collaboration with AI, applications, platform, infrastructure, networking, storage, security, operations, and product teams.
- Delivery and adoption of a reliable proprietary Kubernetes-native scheduler that supports the AI & Applications team’s Model-to-Grid product roadmap.
- Scheduler availability, reliability, recovery performance, and successful completion of scheduled workload operations.
- Improvement in end-user workload outcomes, including lower queue time, reduced time-to-start, improved placement quality, increased job completion rate, faster time-to-results, higher training scaling efficiency, and improved inference throughput and latency.
- Improvement in AI-factory resource efficiency, including GPU utilization, reduced resource fragmentation, improved multi-node placement, increased cluster throughput, more effective capacity use, and reduced avoidable network or storage contention.
- Percentage of eligible workloads scheduled using topology-aware placement policies that account for GPU, NVLink/NVSwitch, network, RDMA, node, storage, and fault-domain characteristics.
- Percentage of eligible workloads scheduled using AI-factory-resource-aware policies incorporating relevant capacity, health, maintenance, power, thermal, storage, or fabric signals.
- Measurable accuracy and effectiveness of scheduler placement and admission decisions, based on workload performance relative to expected benchmarks, queue-time outcomes, job-success rate, resource utilization, user overrides, and operator intervention rates.
- Reduction in manual scheduling, workload-placement, capacity-allocation, and operational-triage effort for users and AI-factory operators.
- Quality and adoption of scheduler APIs, CLI tools, workload templates, model recipes, self-service workflows, scheduling explanations, and developer documentation.
- Coverage and quality of scheduler observability, including end-to-end visibility of workload demand, queue state, scheduling decisions, placement rationale, resource use, failure modes, topology signals, and operational constraints.
- Successful integration of scheduler telemetry and policy outcomes into Model-to-Grid benchmarking, capacity planning, performance optimization, and operator-facing insights.
- Demonstrated support for grid and operator objectives, including improved workload-demand visibility, more predictable capacity use, controlled workload behavior, support for maintenance coordination, and reduced exposure to avoidable power, thermal, capacity, or infrastructure hotspots.
- Release quality for scheduler and platform components, measured through automated-test coverage, performance-regression prevention, scale-test results, fault-recovery validation, security posture, deployment reliability, rollback readiness, and incident trends.
Effective collaboration with the Model-to-Grid product and program lead, Platform, Infrastructure, Security, and operations teams, demonstrated by predictable delivery against agreed feature milestones, early dependency identification, and timely resolution of critical blockers
LOCATION
Singapore or Australia (Launceston, Hobart, Syndey, Melbourne)
Employment Basis
Permanent full-time
Diversity
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionise the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.
About the Company
No detailed information available about this company.
More jobs at Firmus Technologies
-
Senior AI Engineer (Agents & Applications)
Singapore · Full-time · Sep 8, 2026
-
AI Engineer - Inference
Sydney, New South Wales, Australia · Full-time · Sep 8, 2026
-
Principal Data Engineer
Singapore · Contract · Sep 7, 2026
-
Security Engineer
Sydney, New South Wales, Australia · Contract · Sep 6, 2026
-
Senior Security Engineer, Platform Engineering
Sydney, New South Wales, Australia · Full-time · Sep 4, 2026