
GPU Performance Engineer
Open Position
GPU Performance Engineer Hybrid remote in Darmstadt Full-time · PermanentDaisytuner is building the software layer for the next generation of computing. We make complex software run efficiently on any processor—from CPUs and GPUs to novel accelerators—using a self-learning compiler and cloud-scale optimization infrastructure.
Our team brings together researchers and engineers from RWTH Aachen, TU Munich, TU Darmstadt, and ETH Zurich to tackle some of the hardest problems in systems and infrastructure software. If you want to work on deeply technical challenges with real-world impact, join us and help shape the future of compute.
As a GPU Performance Engineer at Daisytuner, you will operate at the intersection of GPU architecture, performance engineering, and compiler technology. You will analyze demanding machine learning and high-performance computing workloads, identify performance bottlenecks, and develop highly optimized GPU kernels. Working closely with our compiler engineers, you will translate performance insights into compiler transformations, optimization heuristics, and autotuning strategies that enable our compiler to automatically generate high-performance GPU code.
What you will be doing- • Analyze and optimize performance-critical GPU workloads, with a strong focus on machine learning applications
- • Benchmark, profile, and hand-tune GPU kernels across modern accelerator architectures
- • Identify bottlenecks related to memory hierarchy, occupancy, instruction throughput, synchronization, and kernel launch configuration
- • Develop highly optimized CUDA kernels and GPU programming techniques for real-world workloads
- • Use profiling and performance analysis tools to understand kernel behavior and hardware utilization
- • Collaborate closely with compiler engineers to translate manual optimization techniques into automated compiler transformations and optimization strategies
- • Design reproducible benchmarking methodologies and performance evaluation pipelines
- • Evaluate optimization quality across different GPU architectures and vendors
- • Contribute to the development of performance models and heuristics for GPU optimization
- • Stay up to date with the latest GPU architectures, programming models, and optimization techniques
- • Master's degree (or equivalent experience) in Computer Science, Electrical Engineering, Mathematics, or a related technical field
- • Strong C/C++ and CUDA programming skills
- • Deep understanding of modern GPU architectures and performance optimization
- • Experience optimizing machine learning, HPC, or other performance-critical GPU workloads
- • Solid understanding of GPU memory hierarchies, shared memory, caches, register usage, occupancy, warp scheduling, and instruction-level performance
- • Experience profiling GPU applications using tools such as NVIDIA Nsight Compute, Nsight Systems, CUPTI, rocProfiler, or similar
- • Ability to analyze low-level performance bottlenecks and systematically improve kernel efficiency
- • Strong Python skills for benchmarking, automation, and experimentation
- • Comfortable working in Linux environments and with Git-based workflows
- • Strong communication skills in English and/or German
- • A structured, analytical working style and willingness to take ownership in an early-stage environment
- • Experience with GPU programming frameworks such as CUTLASS, Triton, cuBLAS, cuDNN, ROCm, or SYCL
- • Familiarity with compiler infrastructures such as LLVM, MLIR, or similar systems
- • Experience with kernel fusion, code generation, or compiler optimizations
- • Knowledge of transformer architectures, LLM inference/training, or other modern ML workloads
- • Experience across multiple GPU vendors (NVIDIA, AMD, Intel)
- • A small, highly technical team with direct impact on core technology
- • Competitive compensation and potential equity participation
- • The opportunity to work at the intersection of compilers, machine learning, and high-performance computing
- • Real ownership over GPU optimization strategies that directly influence the capabilities of our compiler and the performance of production workloads
If that is you, send us a message to hello@daisytuner.com and tell us what you're working on, what you're passionate about, and why you'd like to join our team. Don't forget to provide a CV.
Apply nowThis role is listed because Daisytuner is part of the Ignite Next portfolio. Applications are handled by Daisytuner directly — Ignite Next does not receive your application, screen candidates or act as a recruiter.
Similar roles in the portfolio
AI Network Strategist – Drug Discovery
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 27 Apr 2026
AI Tech Lead: Large Molecules
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 27 May 2026
Forward-Deployed ML Engineer – Cofolding
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 2 Apr 2026
Forward-Deployed Scientist – Computational & Medicinal Chemistry
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 26 days ago
Principal ML Scientist – Predictive Toxicology
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 23 days ago
Senior ML Research EngineerNew
Apheris AI · Remote (UTC + · Remote · Full-time · Posted 5 days ago
