Search Results for "engineer-inference-optimizations"

Found 1032 jobs

Staff Engineer, Inference Optimizations

digitalocean Boston, United States hybrid

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environme...

Posted: Aug 18, 2026 0 views
cuda tensorrt triton
Apply Now

Senior Software Engineer - Inference Engine (Platform Software)

furiosa-ai Seoul HQ full_time

About the Job

Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.

In this role, you will proactively research and apply...

Posted: Aug 18, 2026 0 views
Apply Now

Staff Engineer, Inference Optimizations

digitalocean Austin, United States hybrid

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environme...

Posted: Aug 18, 2026 0 views
cuda tensorrt triton
Apply Now

Staff Engineer, Inference Optimizations

digitalocean Denver, United States hybrid

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environme...

Posted: Aug 18, 2026 0 views
cuda openai tensorrt triton
Apply Now

ML Engineer, Inference & Optimization

pika Palo Alto HQ full_time

About the Role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer – AI Inference Engine

friendliai San Francisco full_time

About the Job

We are seeking a highly technical Inference Engine Engineer to optimize the performance and efficiency of our core inference engine. In this role, you will focus on designing, implementing, and optimizing GPU kernels and supporting infrastructure for next-ge...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer – AI Inference Engine

friendliai Seoul full_time

About the Job

We are seeking a highly technical Inference Engine Engineer to optimize the performance and efficiency of our core inference engine. In this role, you will focus on designing, implementing, and optimizing GPU kernels and supporting infrastructure for next-ge...

Posted: Aug 18, 2026 0 views
Apply Now

Senior Inference Optimization ML Engineer

rhoda-ai Mountain View full_time

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists c...

Posted: Aug 18, 2026 0 views
Apply Now

Software Engineer, Inference - Performance Optimization

openai San Francisco full_time

About the Team
Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, the...

Posted: Aug 18, 2026 0 views
Apply Now

Machine Learning Engineer — Inference Optimization

featherlessai Remote (world) full_time

About the Role

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that ser...

Posted: Aug 18, 2026 0 views
Apply Now
Previous Page 7 of 104 Next