About the job Remote | AI Evaluation Specialist — $30–$90/hour
We are sharing a specialised part-time consulting opportunity for experienced professionals with backgrounds in evaluation, quality assurance, editorial review, assessment, annotation, or related analytical disciplines to contribute to an advanced AI training and quality evaluation project.
Selected professionals will assess AI-generated outputs against structured quality standards, identify reasoning and tool-use failures, and provide clear written feedback that helps improve the reliability and performance of advanced AI systems. The work centres on rigorous, consistent evaluation and careful documentation across high volumes of AI-generated content.
Key Responsibilities
AI Output Evaluation & Quality Assessment
- Evaluate AI-generated outputs against detailed rubrics, guidelines, and defined quality standards
- Assess responses for accuracy, relevance, completeness, reasoning quality, and adherence to instructions
- Apply consistent and impartial judgement across large volumes of evaluation examples
- Identify outputs that fail to satisfy important quality or task requirements
- Maintain reliable assessment standards across repeated evaluation workflows
Reasoning, Logic & Tool-Use Review
- Identify reasoning gaps, logic errors, inconsistencies, and unsupported conclusions in AI assistant responses
- Detect failures involving tool use, workflow execution, or instruction following
- Analyse where AI-generated outputs diverge from expected reasoning or quality standards
- Distinguish between surface-level fluency and genuinely correct, useful, and well-reasoned responses
- Document recurring model weaknesses and opportunities for improvement
Written Feedback & Evaluation Documentation
- Produce clear, concise, and actionable written feedback on strengths and areas for improvement
- Explain the reasoning behind evaluation decisions and quality scores
- Maintain detailed documentation of assessments, findings, and recommendations
- Ensure evaluation records remain transparent, traceable, and reproducible
- Communicate complex findings clearly in professional written English
Rubric Interpretation & Process Improvement
- Participate in discussions involving rubric interpretation, ambiguous cases, and evolving quality standards
- Help refine assessment criteria as AI models and project requirements develop
- Contribute insights that support process optimisation and evaluation best practices
- Apply consistent judgement while adapting to updated guidelines and quality expectations
- Collaborate with other reviewers to improve alignment and reliability across evaluation workflows
Ideal Profile
- Experience in grading, quality assurance, editorial review, assessment, annotation, or another field requiring careful analysis and detailed feedback
- Advanced, regular use of AI assistants such as ChatGPT, Claude, or comparable tools for professional work and productivity
- Strong ability to synthesise complex information and communicate conclusions clearly in writing
- Experience with process improvement, rubric development, operational quality assessment, or structured evaluation workflows is advantageous
- Strong critical-thinking skills with particular emphasis on consistency, integrity, and fairness
- High attention to detail and comfort reviewing large volumes of similar examples
- Ability to work independently while maintaining consistent evaluation quality
- Collaborative approach to discussing ambiguous cases and refining shared assessment standards
- Excellent written English and professional documentation skills
- No prior formal experience in AI research or model training is required
- Based in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand
- Authorised to undertake contract work in the relevant country
Engagement Details
- Part-time independent contractor engagement
- Fully remote across the United States, Canada, United Kingdom, Ireland, Australia, and New Zealand
- Compensation: $30–$90/hour
- Work will involve rubric-based evaluation, AI quality assurance, written feedback, and structured assessment
- Regular use of AI assistants and strong familiarity with AI-enabled workflows are highly relevant to this engagement
- Project scope, workload, timing, and duration may evolve depending on project requirements
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy