Stellar - Director of SRE
✨ AI Summary
Stellar Development Foundation, a blockchain organization expanding access to the global financial system, is hiring a Director of Site Reliability Engineering in New York. The role involves leading SRE strategy and infrastructure across AWS, Kubernetes, CI/CD, and observability tools. Candidates require 10+ years of SRE or infrastructure experience and 5+ years of leadership experience.
About Stellar Development Foundation
The Stellar Development Foundation (SDF) is a mission-driven organization supporting the development and growth of the Stellar blockchain network, an open-source platform designed to expand access to the global financial system.
Since 2014, Stellar has grown into a global blockchain ecosystem used by developers and companies building financial applications and infrastructure around the world.
SDF is now looking for a Director of Site Reliability Engineering to lead its SRE function and shape how engineering teams own, operate and improve production services.
The Role
This is a senior engineering leadership position reporting directly to the CTO.
You’ll lead a small, high-leverage SRE team while defining the broader SRE vision, operating model and reliability culture across engineering.
Rather than SRE acting as the operational owner of every production system, engineering teams at SDF own the services they build. Your role will be to create the infrastructure, frameworks, tooling, standards and observability practices that allow those teams to operate their services reliably and independently.
You’ll combine hands-on technical judgment with organizational leadership, helping SDF improve reliability, infrastructure maturity and developer productivity without introducing unnecessary process or complexity.
What You’ll Work On
You’ll lead, coach and develop a distributed SRE team while establishing its charter, priorities, operating model and measures of success.
A major focus will be defining and rolling out a Service Ownership & Maturity Framework, establishing appropriate reliability and operational standards based on the criticality of individual services.
You’ll own and evolve core engineering infrastructure across:
Cloud infrastructure and foundations
Kubernetes and containerized compute
CI/CD and deployment infrastructure
Observability, monitoring and alerting
Secrets and access management
GitHub workflows
Infrastructure-as-code and automation
You’ll help engineering teams become stronger owners of their production services through better dashboards, runbooks, alerting, escalation paths, deployment practices and operational readiness.
You’ll also improve deployment automation, resilience, self-healing systems, disaster recovery and service reliability, prioritizing improvements based on real operational risk and impact.
Another important part of the role will be evolving incident response, postmortems, escalation and on-call practices across a geographically distributed engineering organization.
You’ll build paved paths and self-service infrastructure that reduce engineering toil and cognitive load while allowing teams to ship faster without compromising reliability.
The role also works closely with Security, Compliance, Legal, Finance, Procurement and Corporate IT wherever cloud infrastructure, access management, vendors or operational controls intersect with engineering.
SDF is also interested in pragmatically exploring AI-assisted and agentic workflows where they can improve infrastructure operations, observability, developer productivity and service ownership.
What We’re Looking For
You bring 10+ years of experience across Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, cloud infrastructure, production operations or closely related areas.
You also have 5+ years of leadership experience, managing or formally developing SRE, infrastructure, platform or reliability engineers.
You have strong experience defining:
Team charters and operating models
Infrastructure and reliability roadmaps
Engineering standards and practices
Success metrics and operational maturity frameworks
You bring deep technical judgment across distributed systems, cloud infrastructure, production operations, automation, reliability engineering and operational risk.
You should also have strong practical experience with:
AWS, GCP or comparable cloud platforms
Kubernetes and container orchestration
Infrastructure-as-code and declarative infrastructure
CI/CD and deployment safety
Observability, logging and monitoring
SLOs and SLIs
Incident response and postmortems
On-call systems and operational readiness
You’ve helped application or product engineering teams take greater ownership of production systems and understand how to balance developer velocity, reliability and operational responsibility.
You’re pragmatic about tooling and comfortable deciding when to build, buy, adapt, simplify or retire infrastructure based on the underlying engineering problem.
Finally, you’re comfortable operating in a lean engineering organization where influence comes from technical credibility, judgment and execution, and can communicate effectively with the CTO and other senior engineering leaders.
Particularly Relevant Experience
Experience in any of the following would be especially valuable:
Leading SRE, Platform or Infrastructure teams in lean, high-agency organizations
Supporting globally distributed engineering teams and 24/7 production environments
Building self-service infrastructure and paved paths
Improving developer productivity through automation and toil reduction
Infrastructure security, secrets management and cloud access controls
Financial services or regulated environments
Blockchain, crypto or Web3 infrastructure
Vendor and infrastructure platform evaluation
Applying AI-assisted or agentic systems to infrastructure, operations, observability or developer workflows
About the Company
No detailed information available about this company.
More jobs at decircletalentpartner
-
Better Money - Applied AI Engineer
New York, New York, United States · · Sep 15, 2026
-
OP Labs - Senior Software Engineer, Protocol (Rust)
Remote job · remote · Aug 19, 2026
-
M0 Labs - Solutions Architect
Remote job · remote · Aug 19, 2026
-
Mural Pay - Forward Deployed Engineer
New York, New York, United States · · Aug 19, 2026
-
Mural Pay - Senior Backend Engineer
New York, New York, United States · · Aug 19, 2026