Inference & Performance Engineer
Engineers in this role optimize how AI models train and run in production, focusing on the full stack from GPU kernels and inference runtimes to distributed training systems and cluster orchestration. They bridge research breakthroughs and production reality, writing CUDA/Triton kernels, tuning serving frameworks like vLLM, profiling end-to-end inference pipelines, and solving the performance and efficiency challenges that emerge when deploying transformer models at scale. They typically sit in infrastructure, platform, or systems teams at AI labs and companies, working closely with researchers and product teams to ensure models meet latency, throughput, and cost targets in real deployments.
Skills
What companies are looking for in this role.
Designing and implementing high-performance distributed inference systems serving large language models at scale
Profiling and optimizing inference workloads across the full stack from model architecture to GPU kernels
Designing and building production inference serving systems and frameworks
Architecting and managing GPU cluster infrastructure for training and inference workloads
Deploying machine learning models to production and managing their lifecycle from training to serving
Implementing model compression techniques such as quantization, distillation, and knowledge distillation for performance gains
Identifying and eliminating bottlenecks through systematic analysis and root cause investigation
Building and maintaining comprehensive benchmarking suites to measure performance and correctness
Writing and optimizing low-level GPU kernels using parallel computing frameworks
Optimizing resource utilization and cost efficiency across heterogeneous hardware environments
Debugging distributed systems issues across compute, networking, storage, and scheduling layers
Designing request routing, load balancing, and traffic management systems for distributed inference fleets
Building platform abstractions and APIs that hide operational complexity from end users
Building infrastructure-as-code systems that treat hardware provisioning and management as declarative state machines
Designing and implementing continuous performance monitoring and alerting systems
Implementing kernel-level optimizations and performance tuning within operating systems
Collaborating with AI researchers and model teams to co-design architectures optimized for production constraints
Building observability and monitoring systems for production AI infrastructure
Evaluating emerging accelerator hardware and novel AI-specific processor architectures for model workloads
Automating performance regression detection and root cause analysis across distributed training and serving jobs
Optimizing transformer model architectures for edge devices and constrained hardware with strict latency and power budgets
Managing on-call rotations and incident response for mission-critical AI infrastructure
Communicating technical findings clearly and translating performance improvements into business outcomes
Taking ownership of systems from design through production operation and maintenance
Working cross-functionally across research, product, customer engineering, and infrastructure teams
Prioritizing work in fast-moving environments with competing demands and high technical complexity
Technology
The tools and technologies that define this role.
Open Jobs
146 open Inference & Performance Engineer jobs across 40 companies.
Other Engineering roles
General-purpose software engineering roles focused on building and maintaining software systems. Covers generalist SWE positions that don't clearly fall into frontend, backend, fullstack, or other specialized tracks.
Engineers focused on server-side systems, APIs, services, and data processing pipelines. Includes roles explicitly labeled as backend or server-side development.
Engineers specializing in user-facing interfaces, web applications, and client-side development. Includes UI/UX engineering and web development roles.
Engineers working across the entire application stack, handling both frontend and backend responsibilities.
Engineers building and maintaining internal platforms, cloud infrastructure, compute systems, and developer tooling.