Software Engineer, Sandbox & Agent Executor
✨ AI Summary
Retool, a developer platform company, is hiring a Software Engineer to build and own the sandbox and agent executor for its AI products. The role focuses on agentic systems, sandboxing, and evaluation within a TypeScript, Node.js, and React stack, requiring Kubernetes experience. Candidates need 6+ years of professional engineering experience with production agentic systems and sandboxing technology.
Why we're looking for you
We build products where agents write and run code, and where the environment that code runs in is ours to own. As agents get more capable, users go from prompt to working app in minutes instead of hours, and increasingly they expect agents that don't just generate the app but run the work inside it.
We're looking for engineers who have shipped agentic products into production and kept them running reliably, affordably, and at scale, and who want to bring that experience to developer-facing surfaces used every day by real engineering teams.
What you'll do
As an AI Engineer, you'll build the agent platform behind Retool's AI products and own model-driven behavior in production. You will work across the product, infrastructure, and evaluation layers. Your work shapes what users experience and how confidently the team can ship. You might:
- Own the behavior of agentic features across multiple product surfaces, including quality, safety, variance, and failure modes, and shape the tool and harness surface agents operate against, including MCP servers, sub-agents, and skills
- Work across the product and infrastructure boundary, shaping agent behavior while understanding what it costs at runtime, and serve as the infrastructure team's technical counterpart on agent workloads
- Design and evolve prompting, context construction, retrieval, routing, and tool-use strategies for long-horizon workflows, and build the evaluation systems that measure them through statistical signals, distributions, and trends rather than pass/fail tests
- Detect, diagnose, and resolve non-deterministic failures such as hallucinations, partial correctness, instruction drift, or context sensitivity, working from transcripts and traces rather than logs alone
- Partner closely with product and infrastructure teams on how agent workloads are provisioned, isolated, and rolled out, including for self-hosted customers, and set the pattern for how we ship agentic products safely
You'll work across the stack (TypeScript, Node.js, React), but your leverage won't come from code volume alone. It will come from shaping runtime behavior with precision, measurement, and intent.
What this role is, and is not
It is
- Accountable for agent behavior, not just system correctness
- Designing, Building, and Deploying agentic products in both cloud and self-hosted environments
- Grounded in evaluation, iteration, and regression prevention under non-determinism
- Comfortable designing systems where outputs vary, confidence is probabilistic, and correctness is contextual
It is not
- Adding LLM calls to existing features and moving on
- Shipping AI features without owning their long-term reliability, drift, or user trust
- An SRE role, though you'll own the agent-side bugs that surface as infrastructure incidents
- Model training or research, though you'll shape model behavior, selection, and tool design
THE SKILLSET YOU’LL BRING
- 6+ years of professional engineering experience, with ownership over complex systems in production
- Production experience with agentic systems, including context engineering, tool use, and evaluation frameworks, at real user scale rather than in pilots or demos
- Hands-on experience with sandboxing technology and running agents inside sandboxed environments
- Experience in Kubernetes or equivalent in practice (EKS, ECS, or similar), owning services end to end
- Strong systems thinking, with the instinct to use AI as an augment to engineering judgment rather than a replacement for it
- Curiosity in why a model produced what it did, and the habit of checking rather than assuming
- Experience mentoring engineers on this kind of work, including when to lean on a model and when not to
BONUS POINTS
- Experience building for developer surfaces like CLIs, IDE extensions, or coding harnesses
- Experience shipping into enterprise or air-gapped environments, where you debug systems you can't directly observe
WHO YOU'LL WORK WITH
You'll join a small, senior team focused on advancing AI capabilities across the product. You'll collaborate closely with product engineers, infra engineers, designers, and PMs, often acting as the final owner of AI behavior and quality before features reach users.
Your work will set standards that others build on. If you enjoy being the person teams rely on when AI behavior matters most, even when certainty is never guaranteed, you'll thrive here.
READY TO BUILD RELIABLE AI SYSTEMS?
If you're excited to move beyond demos and take real ownership of nondeterministic behavior in production, defining quality, preventing regressions, and turning variability into a strength, we'd love to meet you.
About the Company
More jobs at retool
-
Software Engineer, Agent Platform
San Francisco, USA · hybrid · Sep 19, 2026
-
Software Engineer, Automations
San Francisco, USA · hybrid · Sep 16, 2026
-
Software Engineer, Enterprise Expansion
San Francisco, USA · hybrid · Aug 19, 2026
-
Site Reliability Engineer (SRE)
San Francisco, USA · · Aug 19, 2026
-
Solutions Architect
San Francisco, USA · hybrid · Aug 19, 2026