Principal Infrastructure Engineer - Platform
About Radiant
Radiant is redefining how AI infrastructure is built. We design and operate AI-native infrastructure platforms engineered for sovereignty, performance, and scale — powering GPU-native workloads, multi-tenant control planes, and high-performance AI systems for the most demanding environments. We are building purpose-built AI infrastructure from powered land, to compute, to software.
As we scale our operations and deploy capital into the next generation of AI infrastructure, we are looking to expand our finance team with leaders who can combine technical strength with execution excellence and are driven to build.
Radiant was established by Brookfield, a leading global alternative asset manager with over US$1 trillion of assets under management across real estate, infrastructure, renewable power and transition, private equity and credit. Brookfield's global relationships, investment expertise and access to long-term institutional capital provide Radiant with a differentiated platform from which to develop, finance and operate AI infrastructure assets. This combination of entrepreneurial execution and institutional sponsorship enables Radiant to pursue large-scale GPU and AI infrastructure opportunities globally.
About the Role
Location: UK (occasional office travel) | Employment Type: Full-time, Permanent | Reports to: Engineering Manager
We’re looking for a Principal Infrastructure Engineer in our Platform team to drive the architecture, design, and implementation of our next-generation AI-native cloud platform. You will analyse cutting-edge technologies, design scalable and resilient infrastructure solutions, and lead their implementation across a globally distributed environment. You’ll combine deep technical expertise with solution-focused leadership, you will enable our infrastructure to power HPC and AI/ML workloads, delivering value through innovation and collaboration.
What You'll Do
Infrastructure Design & Virtualisation:
Architect, design and implement virtualisation solutions optimised for AI-native workloads, focusing on HPC environments and hypervisor tuning.
Design infrastructure that dynamically adjusts to meet customer demands for storage and networking resources, ensuring resilience and scalability.
Bare-Metal and Operating System Management:
Lead the lifecycle management of bare-metal hardware, ensuring efficient provisioning, orchestration, and optimisation.
Maintain deep expertise in Unix/Linux systems, delivering secure, performant configurations at scale.
Networking and High-Performance Storage:
Design and implement high-performance, cloud-native storage and networking solutions for demanding workloads.
Apply advanced knowledge of networking protocols (TCP, UDP, DNS, encryption) and software-defined networking (SDN) technologies.
Kubernetes and Cloud-Native Platforms:
Deploy and manage Kubernetes clusters across hybrid and multi-cloud environments, leveraging CNIs and service meshes.
Architect scalable CI/CD pipelines to automate infrastructure delivery and ensure reliable global operations.
Security and Audit:
Kubernetes security hardening (RBAC, PSA/PSP successors, network policies, admission controllers - OPA/Gatekeeper, Kyverino)
Secrets management and rotation across clusters (Vault/OpenBao, sealed-secrets, external-secrets patterns)
Supply chain security — image scanning, signed artifacts (cosign/sigstore), provenance attestation
IAM and least-privilege design across bare-metal, cluster, and cloud control planes
Observability and Automation:
Build observability pipelines to monitor and troubleshoot global systems, integrating tools for logs,metrics, and tracing.
Create automation frameworks to reduce operational toil, streamline deployments, and enhance scalability using tools like Terraform, Ansible, Go, and Python.
Architecture & Solution Design:
Analyse emerging technologies and assess their potential fit within our platform, considering scalability, security, and performance.
Develop high- and low-level architectural proposals, balancing technical and business requirements.
Collaborate with cross-functional teams to ensure successful implementation of infrastructure solutions.
Provide thought leadership on architecture best practices, driving alignment across engineering and product teams.
Collaboration and Leadership:
Foster a 'how do we achieve this?' mindset, championing a positive, solution-oriented approach to challenges.
Mentor and guide team members in architectural principles, emerging technologies, and best practices.
Partner with engineering leaders and stakeholders to align infrastructure projects with broader organisational goals.
What You'll Bring
Proven experience in infrastructure architecture and solution design, including the ability to produce detailed technical proposals.
Demonstrated success in evaluating and adopting new toolsets and technologies.
Extensive expertise in large-scale global infrastructure deployments, particularly in AI-native, HPC, and cloud-native environments.
Advanced knowledge of Kubernetes fundamentals, container networking (CNI), and service mesh technologies.
Deep understanding of virtualisation and hypervisor technologies for high-performance workloads.
Strong proficiency in networking protocols (TCP, UDP, DNS, BGP) and SDN solutions.
Skilled in Infrastructure-as-Code (IaC) tools like Terraform, Ansible, and cloud orchestration frameworks.
Hands-on coding experience in Go, Python, or similar languages for automation.
Solid grasp of observability tools (Prometheus, Grafana) and distributed tracing systems.
Strong communication and collaboration skills, with the ability to convey complex technical concepts to cross-functional teams.
Demonstrate ability to create clear and concise documentation for infrastructure designs, processes and systems.
Nice to Have
Expertise in AI/ML workloads, GPU-accelerated systems, or HPC infrastructures, including GPU partitioning with MIG and the NVIDIA GPU Operator.
Experience with RDMA fabrics such as InfiniBand or RoCEv2, and GPUDirect data paths.
High-throughput distributed storage at scale, such as Ceph or a parallel filesystem.
Familiarity with hybrid cloud and edge computing principles.
Experience leading architecture initiatives in agile environments, working iteratively to deliver high-quality results.
Why Join Radiant
Work on a rapidly scaling AI infrastructure business at the forefront of the industry
Get broad exposure to financial operations, AR/AP, close processes, entity setup, audit - developing well-rounded accounting skills
Be part of a team where you're genuinely valued and your work directly enables business decisions
Develop your career in a Brookfield-backed company with resources, mentorship, and genuine growth opportunities
About the Company
More jobs at radiant
-
Engineering Manager
London · full_time · Aug 18, 2026
-
DevOps Engineer
London · full_time · Aug 18, 2026
-
Senior SDET
London · full_time · Aug 18, 2026
-
HPC Infrastructure Site Reliability Engineer
Gloucestershire · full_time · Aug 18, 2026
-
Platform Site Reliability Engineer
Gloucestershire · full_time · Aug 18, 2026