Job Openings Remote | Python Software Engineer — $55–$85/hour

About the job Remote | Python Software Engineer — $55–$85/hour

We are sharing a specialised full-time consulting opportunity for experienced software engineers with strong Python development, debugging, version-control, technical documentation, and AI-assisted coding experience.

This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will design, implement, and review realistic multi-step software engineering tasks that test the capabilities of AI coding agents across Python development, environment setup, tooling, debugging, and technical problem-solving.

Key Responsibilities

Software Engineering Task Design

  • Create realistic, multi-step software engineering challenges based on practical development workflows
  • Design technically demanding problems that require implementation, debugging, environment configuration, and analytical reasoning
  • Define clear requirements, constraints, expected outputs, and acceptance criteria
  • Ensure tasks assess genuine software engineering capability rather than superficial code generation

Reference Solution Development

  • Build complete and verifiable reference solutions in Python
  • Create the supporting setup, dependencies, tests, and validation checks required for each task
  • Write clean, readable, and maintainable code
  • Confirm that solutions run reliably within the intended technical environment
  • Document implementation decisions and expected behaviour clearly

AI Coding Agent Evaluation

  • Use AI coding assistants and agent-based development tools within practical engineering workflows
  • Evaluate how frontier models approach complex coding and debugging tasks
  • Identify implementation errors, unsupported assumptions, inefficient approaches, and incomplete solutions
  • Analyse where AI agents succeed, struggle, or exploit unintended shortcuts
  • Document failure patterns and provide evidence supporting evaluation conclusions

Peer Review & Task Refinement

  • Review tasks and reference solutions created by other software engineering specialists
  • Assess clarity, correctness, difficulty, reproducibility, and technical fairness
  • Identify ambiguous instructions, hidden assumptions, grading gaps, and environment issues
  • Provide actionable feedback that improves task quality and benchmark reliability
  • Collaborate closely with researchers and fellow task authors

Ideal Profile

Strong candidates may have:

  • At least 1 year of experience in software engineering, research engineering, or a related coding-intensive role
  • Strong hands-on Python scripting, implementation, and debugging skills
  • Experience developing clean, readable, and maintainable software
  • Everyday fluency with Git, IDEs, repositories, and standard software development workflows
  • Comfort configuring environments, dependencies, tooling, and validation processes
  • Strong technical writing and documentation skills
  • Ability to work independently through ambiguous and open-ended engineering problems
  • Reliable availability for approximately 35 hours per week

Educational Background

  • An MSc or PhD in computer science, software engineering, another STEM discipline, or a related technical field is highly relevant
  • Equivalent practical experience in a research-intensive or engineering-intensive role may also be considered
  • Academic or professional work involving significant coding, data analysis, or technical experimentation may strengthen an application
  • Open-source contributions, technical projects, publications, or substantial software development work may also be valuable

Nice to Have

  • Experience using AI coding assistants, prompt engineering methods, or agent-based workflows
  • Previous work in AI training, model evaluation, or benchmark development
  • Background authoring technical tasks, reference solutions, or grading criteria
  • Familiarity with automated testing, CI/CD workflows, containers, or reproducible environments
  • Experience reviewing code or technical assignments created by other engineers
  • Knowledge of agentic AI systems and multi-step coding evaluations
  • Strong ability to identify edge cases, unintended shortcuts, and subtle implementation issues
  • Experience collaborating with AI research or evaluation teams

Why This Opportunity

  • Apply practical software engineering expertise to frontier AI evaluation
  • Design realistic coding tasks grounded in professional development workflows
  • Help researchers understand where advanced AI coding agents succeed and fail
  • Work across Python implementation, debugging, environment setup, and benchmark development
  • Collaborate closely with AI researchers and experienced software engineers
  • Participate in a structured full-time remote role with competitive hourly compensation

Contract Details

  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
  • Expected commitment of approximately 35 hours per week
  • Competitive rates between $55–$85 per hour depending on expertise and project scope
  • Individual tasks may require one to two days of focused engineering work
  • Work may include task design, Python development, reference-solution creation, AI agent evaluation, peer review, and technical documentation
  • Engagement scope and duration may evolve according to project requirements and performance

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.