Back to the board
AlignerrVerified· Posted 6mo ago

Python Insfrastructure Engineer - Model Evaluation

Type
Hourly
Location
Remote
$50–75 / hr
Apply on Alignerr
Applying through our link may earn us a small commission — at no extra cost to you.

About this role

Python Infrastructure Engineer — Model Evaluation (AI Training)

About the Role

What if your Python engineering skills could directly shape how the world's most advanced AI models are built, tested, and improved? We're looking for a Senior Python Infrastructure Engineer to design and build the data pipelines, annotation tooling, and evaluation systems that sit at the heart of cutting-edge AI research.

This is a fully remote, flexible contract role working alongside leading AI labs on real production infrastructure. If you're a seasoned Python engineer who wants meaningful, high-impact work — this is it.

  • Organization: Alignerr
  • Type: Hourly Contract
  • Location: Remote
  • Commitment: 20–40 hours/week

What You'll Do

  • Design, build, and optimize high-performance Python systems supporting AI data pipelines and model evaluation workflows
  • Develop full-stack tooling and backend services for large-scale data annotation, validation, and quality control
  • Build and maintain evaluation harnesses for ML models, integrating with inference frameworks
  • Improve reliability, performance, and safety across existing Python codebases
  • Instrument systems with observability tooling and metrics collection to monitor reliability and model performance
  • Identify bottlenecks and edge cases in data and system behavior, and implement scalable solutions
  • Collaborate with data, research, and engineering teams to support model training and evaluation workflows
  • Participate in synchronous design reviews to iterate on architecture and implementation decisions

Who You Are

  • Native or fluent English speaker with clear written and verbal communication skills
  • Full-stack developer with a strong systems programming background
  • 3–5+ years of professional experience writing production-grade Python
  • Experienced building evaluation harnesses for ML models and integrating with inference frameworks
  • Solid background in observability, metrics collection, and monitoring system reliability
  • Self-directed and dependable — able to commit 20–40 hours per week and deliver consistently

Nice to Have

  • Prior experience with data annotation, data quality pipelines, or evaluation systems
  • Familiarity with AI/ML workflows, model training, or benchmarking pipelines
  • Experience with distributed systems or developer tooling
  • Background working directly with AI research teams or at AI-focused companies

Why Join Us

  • Work on production systems powering some of the most advanced AI research in the world
  • Fully remote and async-friendly — work from wherever you do your best thinking
  • Freelance autonomy with the depth and structure of meaningful, long-term engineering work
  • Collaborate directly with AI researchers and engineers at leading labs
  • Potential for ongoing work and contract extension as new projects launch
About Alignerr

A legitimate, well-funded platform (built by Labelbox) with strong rates for experts. The real catch is availability — per-approved-task pay and quiet stretches between projects.

AITrainerGigs aggregates this listing from Alignerr.