Senior / Staff Software Engineer, Cloud & Real Time Infrastructure
Antora Energy delivers affordable, reliable energy to industry, data centers, and the grid. Our thermal batteries turn low-cost electricity into always-on heat and power, without supply-constrained critical minerals or multi-year construction timelines.
We are growing our company with people who put team and mission first, value connection through laughter and joy, and work with humility and openness. We are committed to building a diverse, passionate, and creative team dedicated to a future powered by abundant, clean, low-cost energy.
About The Role
Antora Energy is seeking an engineer to own the real-time cloud infrastructure connecting our software systems to physical hardware.
Every five minutes, our software determines how our thermal batteries should charge and discharge. Those decisions must travel reliably through our messaging infrastructure and reach physical assets in the field. In the other direction, plant telemetry flows from the edge into our cloud data warehouse, where it must be complete, timely, and trustworthy enough to support operational decisions and financial trading.
You’ll also own the corporate cloud platform that the rest of the company builds on. While these systems have very different service-level requirements, they share a need for thoughtful architecture, strong operational practices, and reliable tooling.
We’re hiring one person for a scope that many companies divide across multiple teams. This is a high-impact opportunity to establish platform-wide standards, unify observability, and shape our infrastructure and operational strategy from end to end.
What You'll Do
Own the Real-Time Telemetry and Dispatch Path
Own the full data path from the edge, through streaming ingestion, into the warehouse, and back out to asset dispatch.
You’ll define service-level objectives, build freshness and gap detection, and ensure silent data loss becomes an actionable alert—not something discovered in a report a week later.
Manage Our AWS Infrastructure as Code
Own and evolve our AWS footprint across multiple accounts using Terraform.
You’ll operate containerized services on ECS and Lambda and manage orchestration, networking, secrets, access controls, and supporting infrastructure. You’ll design systems with failure modes, blast radius, observability, capacity, and recovery in mind from the beginning.
Advance an AI-First Software Development Lifecycle
Our engineers already work alongside AI coding agents. The constraint is no longer how quickly code can be written—it is how quickly that code can be verified and deployed safely.
You’ll build the systems that make agent-authored changes safe at scale, including fast and trustworthy CI, hermetic test infrastructure, preview environments, progressive delivery, and automated guardrails.
You’ll also establish the cost controls and verification workflows that allow autonomous development systems to operate without requiring a person to supervise every step.
Own Reliability for Systems Connected to Physical Assets
Participate in the on-call rotation for systems that dispatch instructions to a physical plant.
You’ll build the observability, alerting, and response processes needed to identify problems before they affect field operations. You’ll also define clear escalation paths across cloud infrastructure, site controls, and our market operations desk.
Build a Platform Engineers Want to Use
Treat infrastructure as an internal product. Create self-service tools, paved paths, and platform capabilities that help engineers ship safely without unnecessary friction.
You’ll set standards where consistency matters while avoiding process for process’s sake.
What We're Looking For
- Deep production infrastructure experience. You have several years of experience building and operating cloud infrastructure at scale, with strong expertise in AWS, Terraform, containers, and CI/CD.
- Hands-on site reliability experience. You have carried primary on-call responsibility for a production system that mattered. You have also planned and completed at least one zero-downtime migration involving a stateful, always-on service.
- Experience with real-time or physical systems. You have worked in industrial technology, energy, robotics, automotive, or another environment where software failures have tangible operational consequences.
- An internal-product mindset. You have built self-service tooling that engineers chose to adopt—not simply tooling they were required to use. You understand how to balance standardization, usability, and developer autonomy.
- Strong programming skills. You have strong coding skills in Python, Go, or both. You do more than provision infrastructure: you write and operate the software that keeps systems reliable.
- Practical experience with AI development tools. You use AI coding tools regularly and have a thoughtful perspective on where they work well, where they fail, and how their output should be reviewed and verified.
- Sound operational judgment. You can distinguish between a safeguard that prevents a meaningful failure and a process that only creates friction. You design controls proportionate to the risks involved.
- A willingness to learn across the stack. Deep experience with industrial protocols is valuable but not required. A strong systems engineer who is excited to learn the plant and operational technology side of the business can thrive in this role.
Bonus Qualifications
- Experience with industrial and operational technology protocols, including MQTT or Sparkplug B
- Familiarity with platforms such as HiveMQ, NATS, Ignition, or industrial historians
- Experience operating across strict OT/IT boundaries
- Experience with air-gapped or intermittently connected environments
- Expertise in time-series or high-cardinality data at scale
- Experience with ClickHouse, Dagster, or streaming ingestion systems
- Experience building developer tooling specifically for AI coding agents
- Familiarity with sandboxing, AI cost governance, or agent-readable test output
- Experience with energy markets, dispatch systems, or grid-connected assets
Work Location: Remote / US based
Salary Range: $183-240k USD
Salary Basis:
Please note that the salary range listed above reflects Antora Energy's estimated pay for this position. The actual salary offered will be within the posted range and determined based on several factors including but not limited to a candidate's experiences, credentials and expertise, as they pertain to the position's requirements.
In addition to a competitive base salary, Antora Energy’s Total Rewards program includes equity compensation in the form of stock options, a premium health benefits package with life and disability insurance, a 401K plan with employer contributions, flexible spending accounts, and an industry leading paid-time-off policy that features flexible and inclusive holiday observance, as well as paid volunteer time off.
About the Company
More jobs at Antora Energy
-
Senior / Staff Software Engineer, Full Stack Software Products
Remote · Remote · Aug 19, 2026
-
CNC Machinist/Programmer (Contract)
San Jose, CA · Contract · Aug 19, 2026
-
Senior / Staff Turbomachinery Systems Engineer
San Jose, CA · Contract · Aug 19, 2026
-
Electrical Engineering Manager
San Jose, CA · · Aug 19, 2026