Search Results for "software-engineer-gpu-inference"

Found 386 jobs

AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)

elastix Seattle full_time

Location: Seattle, WA (Hybrid - 3 days/week in office)

About ElastixAI:

ElastixAI is an early-stage Software startup on a mission to reinvent AI inference infrastructure from the ground up. We're building a next-generation inference platfor...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Inference Runtime

lm-studio New York City full_time

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.

As a team, we...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer - Baseten Inference Stack

baseten San Francisco full_time

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enab...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer - GPU Kernels

baseten San Francisco full_time

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enab...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Productivity - Inference Runtime

openai San Francisco full_time

About the Team

We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hirin...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Inference Platform

cerebras Headquarters/Sunnyvale Office full_time

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in s...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Model Inference

openai San Francisco full_time

About the Team

Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never be...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Inference - Multi Modal

openai San Francisco full_time

About the Team

OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we...

Posted: Aug 18, 2026 0 views
Apply Now

Staff Software Engineer, Inference Cloud

cerebras Headquarters/Sunnyvale Office full_time

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in s...

Posted: Aug 18, 2026 0 views
Apply Now

Staff Software Engineer, Inference Platform

cerebras Headquarters/Sunnyvale Office full_time

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in s...

Posted: Aug 18, 2026 0 views
Apply Now
Previous Page 5 of 39 Next