Job Openings Remote | CUDA Engineering Expert — $60–$100/hour

About the job Remote | CUDA Engineering Expert — $60–$100/hour

We are sharing a specialised consulting opportunity for experienced CUDA Engineering Experts with strong expertise in CUDA, C++, GPU kernel optimisation, GLSL, WebGPU, profiling, and high-performance computing to contribute to an advanced GPU-engineering and AI training project.

Selected professionals will analyse and optimise GPU kernels, improve CUDA and C++ codebases, work with shader and WebGPU workflows, and document performance improvements across modern GPU architectures. The work requires strong low-level performance engineering, rigorous profiling, and the ability to translate technical findings into clear, actionable recommendations. No prior experience in AI is required.

Key Responsibilities

CUDA Kernel Analysis & Optimisation

  • Analyse and optimise GPU kernels using CUDA
  • Identify performance bottlenecks affecting computational throughput
  • Improve kernel efficiency across modern GPU hardware
  • Apply targeted optimisation techniques based on profiling results
  • Validate performance gains using quantitative benchmarks

GPU Profiling & Performance Engineering

  • Profile GPU workloads using tools such as Nsight, Visual Profiler, or comparable platforms
  • Diagnose memory, compute, occupancy, and execution bottlenecks
  • Evaluate kernel performance across different hardware generations
  • Develop data-driven optimisation strategies
  • Document measurable changes in latency, throughput, and resource utilisation

C++ & CUDA Codebase Refactoring

  • Refactor CUDA and C++ code for improved maintainability and efficiency
  • Improve architecture and organisation within performance-critical codebases
  • Reduce unnecessary complexity while preserving functionality
  • Adapt implementations for portability across GPU architectures
  • Apply high-performance computing best practices throughout development

GLSL & WebGPU Development

  • Implement shader logic using GLSL
  • Develop graphics and compute workflows using WebGPU
  • Integrate shader-based processing into existing systems
  • Evaluate performance trade-offs across GPU execution environments
  • Maintain compatibility and consistency across pipeline components

Technical Evaluation & Design

  • Contribute expertise to GPU architecture and performance-design discussions
  • Evaluate new GPU-based implementation approaches
  • Assess technical trade-offs across performance, maintainability, and scalability
  • Support definition of meaningful performance metrics
  • Recommend practical approaches based on profiling and benchmarking evidence

Technical Documentation & Collaboration

  • Document optimisation strategies, benchmark results, and technical findings
  • Produce clear reports describing performance improvements
  • Communicate complex GPU behaviour to technical stakeholders
  • Collaborate with remote and cross-disciplinary project teams
  • Share relevant developments in GPU programming and performance engineering

Ideal Profile

  • Demonstrated expertise in CUDA programming
  • Strong track record of GPU kernel performance optimisation
  • Advanced C++ development experience
  • Experience working in high-performance computing environments
  • Hands-on experience with GLSL and WebGPU
  • Strong understanding of graphics or compute shader development
  • Proficiency with GPU profiling tools such as Nsight, Visual Profiler, or equivalent
  • Ability to reason about memory behaviour, kernel execution, and hardware utilisation
  • Strong analytical skills for evaluating performance across GPU architectures
  • Experience refactoring performance-critical codebases
  • Excellent written and verbal technical communication skills
  • Comfortable collaborating in remote, cross-disciplinary environments
  • No prior AI-training experience is required

Engagement Details

  • Independent contractor engagement
  • Fully remote
  • Compensation: $60–$100/hour
  • Work will involve CUDA kernel optimisation, GPU profiling, C++ development, GLSL, WebGPU, and performance analysis
  • Strong low-level GPU performance expertise is central to this engagement
  • Assignments may involve kernel benchmarking, bottleneck analysis, codebase refactoring, shader development, and architectural evaluation
  • Project scope, workload, hardware targets, and performance requirements may evolve depending on project needs
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy