About the job Remote | Python Software Engineer — $55–$85/hour
We are sharing a specialised full-time consulting opportunity for experienced software engineers with strong Python development, debugging, version-control, technical documentation, and AI-assisted coding experience.
This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will design, implement, and review realistic multi-step software engineering tasks that test the capabilities of AI coding agents across Python development, environment setup, tooling, debugging, and technical problem-solving.
Key Responsibilities
Software Engineering Task Design
- Create realistic, multi-step software engineering challenges based on practical development workflows
- Design technically demanding problems that require implementation, debugging, environment configuration, and analytical reasoning
- Define clear requirements, constraints, expected outputs, and acceptance criteria
- Ensure tasks assess genuine software engineering capability rather than superficial code generation
Reference Solution Development
- Build complete and verifiable reference solutions in Python
- Create the supporting setup, dependencies, tests, and validation checks required for each task
- Write clean, readable, and maintainable code
- Confirm that solutions run reliably within the intended technical environment
- Document implementation decisions and expected behaviour clearly
AI Coding Agent Evaluation
- Use AI coding assistants and agent-based development tools within practical engineering workflows
- Evaluate how frontier models approach complex coding and debugging tasks
- Identify implementation errors, unsupported assumptions, inefficient approaches, and incomplete solutions
- Analyse where AI agents succeed, struggle, or exploit unintended shortcuts
- Document failure patterns and provide evidence supporting evaluation conclusions
Peer Review & Task Refinement
- Review tasks and reference solutions created by other software engineering specialists
- Assess clarity, correctness, difficulty, reproducibility, and technical fairness
- Identify ambiguous instructions, hidden assumptions, grading gaps, and environment issues
- Provide actionable feedback that improves task quality and benchmark reliability
- Collaborate closely with researchers and fellow task authors
Ideal Profile
Strong candidates may have:
- At least 1 year of experience in software engineering, research engineering, or a related coding-intensive role
- Strong hands-on Python scripting, implementation, and debugging skills
- Experience developing clean, readable, and maintainable software
- Everyday fluency with Git, IDEs, repositories, and standard software development workflows
- Comfort configuring environments, dependencies, tooling, and validation processes
- Strong technical writing and documentation skills
- Ability to work independently through ambiguous and open-ended engineering problems
- Reliable availability for approximately 35 hours per week
Educational Background
- An MSc or PhD in computer science, software engineering, another STEM discipline, or a related technical field is highly relevant
- Equivalent practical experience in a research-intensive or engineering-intensive role may also be considered
- Academic or professional work involving significant coding, data analysis, or technical experimentation may strengthen an application
- Open-source contributions, technical projects, publications, or substantial software development work may also be valuable
Nice to Have
- Experience using AI coding assistants, prompt engineering methods, or agent-based workflows
- Previous work in AI training, model evaluation, or benchmark development
- Background authoring technical tasks, reference solutions, or grading criteria
- Familiarity with automated testing, CI/CD workflows, containers, or reproducible environments
- Experience reviewing code or technical assignments created by other engineers
- Knowledge of agentic AI systems and multi-step coding evaluations
- Strong ability to identify edge cases, unintended shortcuts, and subtle implementation issues
- Experience collaborating with AI research or evaluation teams
Why This Opportunity
- Apply practical software engineering expertise to frontier AI evaluation
- Design realistic coding tasks grounded in professional development workflows
- Help researchers understand where advanced AI coding agents succeed and fail
- Work across Python implementation, debugging, environment setup, and benchmark development
- Collaborate closely with AI researchers and experienced software engineers
- Participate in a structured full-time remote role with competitive hourly compensation
Contract Details
- Full-time W-2 contingent employment opportunity
- Fully remote within the United States
- Expected commitment of approximately 35 hours per week
- Competitive rates between $55–$85 per hour depending on expertise and project scope
- Individual tasks may require one to two days of focused engineering work
- Work may include task design, Python development, reference-solution creation, AI agent evaluation, peer review, and technical documentation
- Engagement scope and duration may evolve according to project requirements and performance
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.