About this role
Machine Learning Evaluation Specialist (AI Training)
About the Role
AI is only as good as the challenges used to test it — and we need experts like you to design them. We're looking for machine learning specialists with deep domain knowledge to create the hard, research-level problems that state-of-the-art AI systems genuinely struggle with.
Your work will directly shape how the next generation of AI models is evaluated and improved. This is a rare opportunity to apply your hard-earned expertise at the cutting edge of AI development.
- Organization: Alignerr
- Type: Hourly Contract
- Location: Remote
- Commitment: 10–40 hours/week
What You'll Do
- Design complex, original ML problems rooted in your specialized research domain
- Craft evaluation tasks that expose the limits of today's most capable AI systems
- Draw from your own research experience to build challenges that go beyond standard ML pipelines
- Write clear problem statements, define evaluation criteria, and provide gold-standard solutions
- Review and assess AI-generated solutions for correctness, creativity, and methodological rigor
- Document problem difficulty, required domain knowledge, and expected failure modes
- Collaborate asynchronously with a global team of researchers and engineers
Who You Are
- Graduate-level expertise (MS or PhD preferred) in a scientific or technical field that intersects with machine learning
- Strong working knowledge of ML methods — model selection, feature engineering, evaluation metrics, and pipeline design
- Deep familiarity with active research problems in your domain
- Ability to identify exactly where general ML knowledge breaks down and specialized insight becomes essential
- Experience publishing or conducting original research is highly valued
- Excellent written communication — you can articulate nuanced technical problems with precision and clarity
- Self-motivated and comfortable working independently on intellectually demanding tasks
Example Domains
We're looking for experts across a wide range of specializations, including:
- Computational biology, genomics, or bioinformatics
- Climate science and environmental modeling
- Medical imaging and healthcare ML
- Materials science and computational chemistry
- Astrophysics and signal processing
- NLP for low-resource or specialized corpora
- Robotics, control theory, or reinforcement learning
- Financial modeling and quantitative analysis
Why Join Us
- Work at the frontier of AI evaluation and safety research
- Collaborate with top research labs pushing the boundaries of what AI can do
- Put your hard-earned domain expertise to meaningful, high-impact use
- Full autonomy over your schedule — work when and as much as you want
- Fully remote and asynchronous — work from anywhere in the world
- Potential for ongoing work, contract extensions, and deeper research involvement
- Build your profile as a recognized contributor to cutting-edge AI development
A legitimate, well-funded platform (built by Labelbox) with strong rates for experts. The real catch is availability — per-approved-task pay and quiet stretches between projects.
AITrainerGigs aggregates this listing from Alignerr.