Platform Engineer - Incident Management
✨ AI Summary
Bitso, a high-scale crypto platform, is hiring a Platform Engineer 2 focused on Incident Management for a remote, full-time position. The role requires hands-on Kubernetes experience, CI/CD knowledge, and software development skills in Python or Java to automate incident response and postmortems. Candidates should have a strong automation mindset, with experience in AI agents or LLM workflows being highly desirable.
Your Purpose
At Bitso, reliability isn’t an afterthought — it’s a competitive advantage. As a Platform Engineer 2 focused on Incident Management, you’ll own the full incident lifecycle: from active response during live incidents, to driving postmortems, building automation, and eliminating the root causes that create toil in the first place. You’ll be the person who asks “how do we make sure this never happens again?” — and then actually builds it. If you thrive under pressure, love automation, and want to make a measurable dent in how a high-scale crypto platform operates, this role was designed for you.
Reports To
Incident Management Manager
Who You Are
- Proven ability to operate confidently in high-pressure incident scenarios, including communicating clearly with senior stakeholders and leadership while a production issue is live
- Hands-on experience with Kubernetes — comfortable deploying, debugging, and navigating pod-level issues
- Solid understanding of CI/CD pipelines and modern DevOps practices
- Software development background in any language; ability to read, write, and debug code is essential (Python or Java experience is a plus)
- Strong automation mindset: you identify repetitive toil and your first instinct is to eliminate it, not absorb it
- Experience building or working with AI agents or LLM-based workflows is highly desirable
- Strong interpersonal and written communication skills
- Self-directed learner who doesn’t need a fully defined path to start contributing
- Fintech or crypto industry background is a plus — familiarity with the domain vocabulary accelerates onboarding and incident triage
What You Will Do
- Own and execute on-call shifts end-to-end: acknowledge pages within SLA, declare incidents, assign roles, maintain comms cadence, and drive to resolution
- Build automation that drives the Sev1/Sev2 postmortem workflow — from scheduling and facilitation reminders to action-item assignment, ownership tracking, and due-date enforcement
- Leverage AI to identify patterns across incidents and propose systemic fixes: runbook improvements, alert tuning, platform hardening, and process changes
- Build and extend internal automation and tooling, including AI-assisted incident response workflows, to reduce manual toil and accelerate detection and resolution
- Contribute to and improve the observability ecosystem — dashboards, alert configurations, and early-warning signals across Bitso’s platform
- Participate in change and maintenance management processes, applying risk management to reduce deployment-related incidents
- Collaborate with engineering squads across the company to surface platform risks and drive preventive actions Keep incident tooling, runbooks, and severity criteria accurate, current, and useful for the broader engineering org
Research in Diversity, Equity, and Inclusion suggests that individuals may hesitate to apply for jobs if they do not meet all the listed criteria. At Bitso, we value diversity and your unique strengths could be just what we're looking for. If this role excites you but you don't match every point in the description, we still want to hear from you.
#LI-Remote
Location
Remote
Department
Platform
Employment Type
Full-Time
Privacy Policy • Terms of Service
•
© BambooHR All rights reserved.
About the Company
Latin America's leading crypto-based financial services company
More jobs at Bitso
-
Senior Engineering Manager
Remote · Full-Time · Sep 20, 2026
-
Software Engineer - Latam
Remote · Full-Time, Remote · Sep 20, 2026
-
Senior Mobile Software Engineer (React Native)
Latin America · Remote · Aug 19, 2026
-
Platform Engineer - (Site Reliability Engineering)
Latin America · Remote · Aug 19, 2026
-
Senior Software Engineer
Latin America · Remote · Aug 19, 2026
Similar Platform Engineer roles
-
Senior Platform Engineer
chronicle · Remote - Europe, UK · Aug 10, 2026
-
Senior Platform Engineer
mintlify · San Francisco · Sep 20, 2026
-
AI Platform Engineer
modus-create · Colombia · Sep 19, 2026
-
AI Platform Engineer
modus-create · Costa Rica · Sep 19, 2026
-
Senior Data Platform Engineer
endsight · Napa, CA · Sep 19, 2026