About the job Remote | Computational Research Scientist — $55–$85/hour
We are sharing a specialised full-time consulting opportunity for researchers with strong computational, experimental-design, data-analysis, and scientific reasoning experience across STEM, computational social science, or computational humanities disciplines.
This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected researchers will translate real research methods—including study design, hypothesis testing, simulation, modelling, and rigorous evaluation—into complex multi-step tasks requiring coding, data analysis, and carefully supported conclusions.
Key Responsibilities
Research Task Design
- Transform real research workflows into challenging, multi-step benchmark tasks
- Develop assignments involving study design, hypothesis testing, simulation, modelling, or data analysis
- Create tasks that assess genuine scientific reasoning rather than surface-level pattern matching
- Ensure each task includes realistic assumptions, constraints, datasets, and evaluation objectives
Technical Implementation
- Implement complete reference solutions using Python and notebook environments
- Build reproducible analyses, simulations, models, or data-processing pipelines
- Validate calculations, code, intermediate outputs, and final conclusions
- Document methodologies clearly enough for independent review and reproduction
Evaluation Criteria & Reference Solutions
- Define what distinguishes rigorous research reasoning from plausible but unsupported analysis
- Develop reference answers, grading criteria, and structured evaluation guidelines
- Identify required methodological steps, valid alternative approaches, and material errors
- Ensure grading standards reflect the quality expected from an experienced researcher
Model Output Review
- Evaluate AI-generated attempts across computational research tasks
- Identify methodological errors, unsupported claims, statistical weaknesses, and coding issues
- Assess whether conclusions follow logically from the evidence and analysis
- Provide clear written feedback explaining errors and recommended improvements
Research Collaboration
- Work closely with researchers, task authors, and fellow subject-matter experts
- Compare evaluation decisions and support consistent benchmark standards
- Incorporate feedback into task design, reference solutions, and scoring criteria
- Surface recurring model failure patterns and opportunities for stronger evaluation coverage
Ideal Profile
Strong candidates may have:
- At least 1 year of experience in an active academic, industry, government, or national-laboratory research role
- An MSc or PhD in a STEM field, computational social science, computational humanities, or another research-intensive discipline
- Significant experience using Python for analysis, simulation, modelling, or data pipelines
- Strong grounding in experimental design, hypothesis testing, and rigorous evaluation
- Experience interpreting complex datasets and producing evidence-based conclusions
- Working familiarity with Git, IDEs, and Jupyter or Colab notebooks
- Strong written communication and technical documentation skills
- Ability to work independently through ambiguous and open-ended research problems
- Reliable availability for approximately 35 hours per week
Educational Background
- A master's degree or PhD in a relevant computational or research-focused discipline is highly relevant
- Equivalent practical experience in a research-intensive role involving coding and data analysis may also be considered
- Suitable backgrounds may include computer science, engineering, physics, chemistry, biology, mathematics, statistics, economics, psychology, political science, linguistics, digital humanities, or related fields
- Publications, conference presentations, technical reports, or impactful research contributions may strengthen an application
Nice to Have
- Experience in AI training, model evaluation, or benchmark development
- Background authoring technical tasks, reference solutions, or grading rubrics
- Familiarity with machine learning, statistical modelling, or scientific computing
- Experience designing reproducible experiments and validating analytical pipelines
- Knowledge of research-quality assurance, peer review, or methodological auditing
- Familiarity with agentic AI systems and multi-step model evaluations
- Experience reviewing code, notebooks, or technical analyses prepared by other researchers
- Strong ability to identify subtle methodological flaws and unsupported conclusions
Why This Opportunity
- Apply working-researcher expertise to advanced AI benchmark development
- Design complex tasks grounded in authentic scientific and computational workflows
- Help improve how frontier AI systems reason through research problems
- Work across experimental design, coding, data analysis, and evidence-based conclusions
- Collaborate closely with AI researchers and computational experts
- Participate in a structured full-time remote role with competitive hourly compensation
Contract Details
- Full-time W-2 contingent employment opportunity
- Fully remote within the United States
- Expected commitment of approximately 35 hours per week
- Competitive rates between $55–$85 per hour depending on expertise and project scope
- Individual tasks may require one to two days of focused research and implementation work
- Work may include task design, Python implementation, notebook development, model-output evaluation, rubric creation, and technical documentation
- Engagement scope and duration may evolve according to project requirements and performance
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.