← All roles

GPU Performance Engineer (Inference)

Dublin, Ireland · Hybrid Full-time Mid level
Apply for this role Posted 26 Sep 2026

Overview

TensorX is a sovereign AI infrastructure platform headquartered in Dublin. We run frontier open-weight large language models on our own NVIDIA Blackwell GPUs in European datacentres, under EU jurisdiction. Customers reach them through a drop-in OpenAI-compatible API. Nothing they send is retained after the request completes. We serve regulated enterprises in finance, healthcare and government as well as developers and AI platforms. We help them adopt AI without compromising on data privacy, compliance or performance.


We are looking for multiple GPU Performance Engineers to join our growing engineering team. Reporting to the CTO, you will make each inference engine in our fleet do more work inside its latency target. You will take engine problems from report to fix, patch the open-source engines we depend on and go beneath them to the kernel when the engine is the limit.


Our margin depends on one number: how many requests each GPU answers inside a latency target. That is not the same as peak tokens per second on a benchmark. Passengers carried, not miles per hour. Our customers send very long contexts at high concurrency, so the constraint that matters is rarely the one the headline number measures.


The serving engines we depend on, SGLang first, are open source and still maturing. We hit their defects early because we run new models on day zero. We patch them, tune them and change the kernel where the engine itself is the limit. A new open-weight model worth serving arrives most weeks and more GPU capacity is going live now. We are an AI-native team. Tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems.


You will work side by side with our Inference Team, who run NVIDIA Dynamo and the serving fleet. You make each worker faster. They make the fleet of workers behave.


We hire on evidence of skill, not years of experience. We will hire across multiple levels. One of the seats may suit someone earlier in their career.


This is a hands-on individual contributor role spanning engine performance, GPU kernels, model bring-up and cache behaviour.

Responsibilities

  • Engine tickets end to end - Take engine problems from report to fix. A model runs out of memory at long context, a patch cuts throughput or a new release breaks tool calling. Reproduce it, find the mechanism, test a fix, patch it and roll it out. Then write down what you found.

  • Investigation & diagnosis - Debug across GPU memory, the engine scheduler, the container, the router and the traffic shape. Many of our hardest problems sit where two of these meet. When you raise a problem with the team, bring a read, not a question: what you think it is and why.

  • Goodput per GPU - Goodput is the number of requests each GPU answers inside the latency target. Find where it is being lost, whether in a kernel, the scheduler, the router or a configuration flag. Win it back. Measure every gain on production traffic patterns rather than a synthetic load.

  • Engine patches & upstream - Patch SGLang and vLLM defects on new models. Bake each fix into the image and keep the patch set consistent across the fleet. Where a fix is not specific to us, open the pull request upstream, not just the issue.

  • GPU kernels - When profiling shows an engine kernel is the bottleneck, write or modify it. Examples include Blackwell attention/indexer paths, FP8/FP4 paths and memory-bound decode. Profile before you form an opinion.

  • Model bring-up - Bring up new models on day zero. Run A/B bake-offs across parallelism layouts, KV cache configurations and speculative decoding settings. Test tensor, data and expert parallelism for each model. Give the Inference Team the best layout for each model so they can size the pools.

  • Cache behaviour - Tune prefix caching, KV cache behaviour and KV replication under tensor parallelism inside the engine. Work with the Inference Team on cache-aware routing and contribute to the router code.

  • Pre-production gate - Share the pre-production gate with the Inference Team so no change reaches a customer unmeasured. Run it on your own changes and on every model you bring up.

  • Documentation - Write up every finding as a report someone else can rerun. Explain trade-offs plainly enough for the Inference Team and for customers. Keep a benchmark method that works without you in the room.

Skills & Experience

  • Evidence that you can read an inference engine's source to find the mechanism behind a problem rather than only the symptom. A pull request, a write-up or a post-mortem we can check

  • Solid understanding of GPU architecture and CUDA fundamentals: memory hierarchy, occupancy and what makes a kernel compute-bound or memory-bound

  • Working knowledge of transformer inference: attention variants (MLA, DSA, GQA), KV cache behaviour, continuous batching and the accuracy impact of quantisation

  • Proficiency in Python and comfort reading C++ and CUDA

  • A measure-first habit: one variable per test arm, a pass or fail bar set before you run and a correctness check before you trust a timing

  • Honest about results, including the ones that did not work out. Losing to a baseline and saying so counts in your favour

  • A drive to learn. The stack changes every week. We value someone who reads the source over someone who already knows last year's answer

  • Comfortable using AI-assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow

  • A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non-technical audiences

Nice to Have

  • A track record of writing and benchmarking CUDA kernels on Hopper or Blackwell

  • Upstream contributions to SGLang, vLLM or TensorRT-LLM

  • Experience with Triton, CUTLASS, CuTe, ThunderKittens or similar kernel tools

  • Kernel entries in GPU Mode leaderboards, MLSys contests or FlashInfer challenges

  • Experience profiling memory behaviour and out-of-memory errors on large models

  • Hands-on time with Blackwell features (tcgen05, TMEM, TMA)

  • Technical writing in public

Why This Role

  • You work inside the engine. Many GPU clouds hire performance engineers for cold starts, storage and containers. We hire them to change the engine and the kernel. That is where our margin is.

  • The fleet is ours. Our own NVIDIA B300 GPUs in Dublin and Helsinki. You are not renting time on someone else's cluster.

  • The traffic is real. Long-context, high-concurrency production workloads that break assumptions benchmarks never test.

  • The results are real. Our inference stack answers roughly twice as many requests inside the latency target as a standard configuration. We measured this on the same hardware with the same production traffic patterns.

  • The engine is open. We patch SGLang in production and send the fixes upstream. If there is a paper in the work, we would rather it went out with your name on it.

  • The research list is long. The better the day-to-day is covered, the more time goes on KV cache beyond GPU memory, context parallelism, wide expert parallelism and attention/FFN disaggregation.

Education & Qualifications

  • BSc/MSc/PhD in Computer Science, Engineering, Machine Learning or a related technical discipline OR equivalent demonstrable ability

Remuneration

  • Highly competitive package, dependent on experience

  • 25 days paid annual leave

  • Hybrid working from our centrally located Dublin office, with remote flexibility

  • Free inference tokens!


******* NO AGENCY ASSISTANCE REQUIRED *******

Ready to apply?

Applications are handled securely through our recruitment platform.

Apply now

TensorX is an equal-opportunity employer. All inference runs on our own EU-sovereign infrastructure. Learn more about what we build.