Back to the board
MercorVerified· Posted 23d ago

Trainium (NKI) Kernel Expert

Commitment
40 hrs/week
Location
Remote
Availability
3 spots left
$70–90 / hr
Apply on Mercor
Applying through our link may earn us a small commission — at no extra cost to you.

About this role

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to train and evaluate a frontier AI lab's models. You'll assess CUDA→NKI migration fidelity, Trainium-specific performance-optimization quality, and cross-platform numerical-correctness standards — and provide clear, rubric-based written feedback.

Basic Qualifications

  • 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware
  • Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration
  • Demonstrated experience assessing CUDA→NKI migration quality
  • Familiarity with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks)
  • Experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)

Preferred Qualifications

  • Direct experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries
  • Prior CUDA or Triton kernel development
  • Familiarity with Trainium hardware specifications (NeuronCore-v2 architecture, on-chip SRAM topology, supported data types: FP32/BF16/FP8/INT8)
  • Experience benchmarking ML training workloads on Trn1/Trn2 instances
About Mercor

The strongest pipeline for credentialed experts. Contracts are clear, and workers report payouts landing on schedule.

Trust score 9.2/10Read our Mercor review →

AITrainerGigs aggregates this listing from Mercor.