Machine Learning Engineer - Kernels
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Employment Type: Full-time
Experience: 2+ years in GPU programming, parallel computing, or systems-level optimization
Core Areas: GPU Kernel Development, Machine Learning Workload Optimization, CUDA, C++, Parallel Computing, Performance Profiling, Hardware Acceleration, Distributed Computing
Compensation: $150,000 – $190,000 per year
About the Role
This opportunity is for a Machine Learning Engineer - Kernels specializing in custom GPU and accelerator kernel development for high-performance AI and machine learning workloads. The role focuses on designing efficient low-level implementations, optimizing computational performance, and translating advances in machine learning algorithms into reliable, production-ready code.
The position combines GPU programming, parallel computing, hardware-aware optimization, and machine learning infrastructure. Work involves C++, CUDA, performance profiling, benchmarking, and optimization across distributed and heterogeneous computing environments. The engineer will collaborate with researchers, evaluate hardware developments involving CUDA, ROCm, and TPUs, and contribute to engineering practices that improve computational efficiency for advanced AI workloads.
What You'll Do
- Design, develop, and implement custom GPU and accelerator kernels to maximize computational performance for machine learning workloads.
- Profile and benchmark performance-critical ML workloads, identify computational bottlenecks, and implement low-level optimizations to improve execution efficiency.
- Collaborate with machine learning researchers to translate algorithmic advances into efficient, production-ready software implementations.
- Evaluate developments in accelerator hardware and programming technologies, including CUDA, ROCm, and TPUs, to guide kernel design and optimization decisions.
- Document low-level optimization techniques and share performance engineering best practices with technical teams.
Qualifications
Required Experience
- At least 2 years of professional experience in GPU programming, parallel computing, or systems-level performance optimization.
- Experience optimizing computational workloads for distributed and heterogeneous computing environments.
Required Skills
- Strong programming skills in C++, CUDA, or comparable low-level programming languages used for performance-intensive computing.
- Familiarity with machine learning frameworks and their low-level computational backends.
- Ability to use profiling tools and performance diagnostics to investigate execution behavior and identify optimization opportunities.
- Strong understanding of performance-oriented programming and the interaction between algorithms and hardware.
- Detail-oriented approach to engineering, with a strong focus on computational efficiency and performance optimization.
- Ability to collaborate effectively on research-driven engineering challenges and contribute to technically ambitious development work.
Education
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.
Looking for more opportunities?
View All Jobs