Companies Hiring AI Infrastructure Engineers in 2026
Who is hiring AI infrastructure engineers in 2026, by company family: frontier labs, chip makers, hyperscalers, GPU clouds, inference providers and product companies with their own ML platforms, in the US and in India, with what each family's loop weights.
7 MIN READ · UPDATED 4 SEPTEMBER 2026
PRACTICE THIS:AI Infrastructure System Design ·GPU Fleet Reliability & Observability ·LLM Inference & Serving ·Napkin Math, Cost & Capacity
Frontier labs
OpenAI posts the widest range of infrastructure roles in our catalogue: GPU Infrastructure, GPU Infrastructure HPC, Frontier Clusters Infrastructure, Fleet Infrastructure, a London GPU-cluster role serving ChatGPT, Compiler, Kernels and Runtime, and Data Infrastructure for Research. Anthropic posts Performance Engineer for inference systems and for GPUs, Software Engineer, Infrastructure at all levels and staff, a London distributed-systems infrastructure role, and Data Infra for pretraining. Google DeepMind posts Model Inference and data infrastructure for Gemini; Google posts TPU compiler and ML infrastructure roles. Meta posts Software Engineer, Systems ML and production engineering for its AI infrastructure. Microsoft AI posts Member of Technical Staff roles for pre-training infrastructure and data infrastructure on its superintelligence team. xAI posts Member of Technical Staff for ML and data infrastructure. Mistral posts inference technical lead, SRE and ML infrastructure roles in Paris and London. The loops weight the kernel, distributed and platform axes most, with design at frontier scale; the company pages have each.
Chip makers and hyperscalers
NVIDIA hires across the whole role: new-grad AI and ML Infrastructure Software Engineers for GPU clusters and DGX Cloud Performance Engineers (with posted bands), senior GPU and HPC infrastructure engineers, HPC AI cluster engineers, and in India a DGX Cloud Performance Engineer in Pune and Site Reliability Engineers in Bengaluru. AMD posts AI Cluster Validation Engineers and AI Cluster Software Engineers for its Instinct clusters. AWS Annapurna Labs posts Neuron distributed training, Neuron inference, Neuron compiler and early-career ML systems roles for Trainium and Inferentia. Google's TPU teams post compiler and infrastructure roles. Cerebras posts inference-core infrastructure software engineers and entry-level AI infrastructure operations engineers. Groq posts senior staff and principal inference-system and inference-stack engineers. These loops weight architecture depth, the vendor's software stack, benchmarking methodology and cluster validation.
GPU clouds
CoreWeave posts Senior GPU Infrastructure Software Engineers for fleet validation, Kubernetes operators and Slurm-on-Kubernetes. Lambda posts platform engineers for GPU host lifecycle and staff engineers for managed Kubernetes. Crusoe posts SREs and production engineers for its managed AI services. Nebius posts GPU performance and compute systems engineers, AI infrastructure systems engineers, inference-platform SREs and GPU cluster architects, with India among its locations. Together AI posts AI Infrastructure Systems Engineers in Bengaluru (fleet automation and AI-driven triage), staff inference and compute infrastructure engineers listing India, and backend-systems roles under an MLOps title. These loops weight the platform and cost axes: fleet operations, Kubernetes at scale, incident stories, and the economics of a rented GPU-hour.
Inference providers and platforms
Fireworks AI listed 33 open systems and ML infrastructure roles in September 2026 across inference optimization, kernel engineering and serving at scale. Baseten posts customer-facing AI Inference Engineers and internal Model Performance and inference-stack engineers. Modal posts ML performance and systems engineers for a Rust-based serverless GPU platform. Anyscale posts ML platform engineers building Ray Serve. Hugging Face has posted optimized-inference roles with CUDA kernel work. These loops weight the inference and kernel axes with pairing rounds on real code, and the design round is a serving platform with a latency budget.
Product companies with ML platforms
Uber posts staff and mid-level engineers for Michelangelo, its ML-as-a-service platform. Netflix posts L4 and L5 engineers for data and feature infrastructure and an ML platform reliability engineer. Databricks posts senior ML engineers for its GenAI platform (serving, fine-tuning, vector search). Snowflake posts software and staff engineers for Cortex AI infrastructure. These loops are backend system design with GPUs as the scheduled resource, multi-tenancy, and the platform's internal customers as stakeholders; the ML platform design prompts on this site match them.
India
Hiring in India for this role runs through three routes. Multinational engineering centres: NVIDIA (Pune and Bengaluru), Microsoft India (Azure graduate engineers placed into core compute, networking, AKS and Azure OpenAI teams, 250 to 400 per year), and the India centres of Google and AWS. US-headquartered clouds and labs hiring India-based engineers directly: Together AI and Nebius in Bengaluru. India-headquartered AI companies: Sarvam AI (GPU Infrastructure Engineer open to recent graduates, Platform Engineer for AI Infrastructure, Staff Engineer for Product Infrastructure, all Bengaluru) and Krutrim (LLM and agentic engineers, freshers, Bengaluru). GPU-cloud and datacenter operators such as Yotta, E2E Networks and Jio hire cluster and datacenter staff through job boards we could not fetch, so treat their absence here as a gap in our catalogue rather than an absence of roles. The salary guide explains why the route matters more than the company for India compensation.
Reading the directory below
The directory renders from the company registry, so it stays current as research lands. Each company page carries the reported loop, the axes it weights, the posted bands with sources, the India route, and the questions in the bank tagged to that company. We never invent a round: where a company's loop is not publicly reported, the page says so and gives the closest evidence we have.
The directory: 41 companies, and where to apply
Each entry links to that company’s own careers page, which is the only listing guaranteed to be current, and to our breakdown of its interview loop.
Filtering job boards on the exact phrase hides most of the market. The same job is posted under all of these:
Runs a real AI infrastructure program
Posts AI infrastructure, GPU, kernel or training-infra roles by name. Start here.
GPU clouds and serving platforms
Cluster, scheduling and inference-engine work under titles like infrastructure engineer, platform engineer or performance engineer.
Adjacent: strong ML and platform engineering
No classic AI infra program, but the loops overlap heavily and candidates from this site interview here regularly.
We do not host or verify individual job listings. Openings, titles and requirements change constantly, so always confirm on the company’s own careers page.
Turn the theory into offers — work the question topics this maps to:
FAQ
By open postings in our September 2026 catalogue: OpenAI (seven distinct infrastructure roles), NVIDIA (across new-grad, senior and India), Fireworks AI (33 systems and ML infra roles), Anthropic, Nebius, Together AI and the AWS Annapurna Labs teams. The frontier labs and the GPU clouds are hiring hardest for fleet and inference roles; the chip makers for cluster and performance roles.
