Senior Software Engineer – Open Source & SWE-Bench Evaluation
Anyone AI is recruiting experienced Software Engineers for a specialized project focused on reviewing and evaluating real-world software engineering tasks derived from open-source repositories.
The work involves assessing whether coding tasks based on real GitHub issues and pull requests are technically sound, reproducible, appropriately tested, and representative of the kinds of problems professional software engineers solve every day.
What You’ll Work On
You’ll review software engineering tasks involving:
Real-world bug fixes and feature implementations
Open-source repositories and pull requests
Unit tests and test coverage
Repository setup and dependency management
Reproducibility and environment configuration
Task difficulty and complexity
Multi-file and cross-module code changes
Technical feedback and quality assessment
You’ll determine whether tasks are clearly specified, technically solvable, supported by sufficient tests, and free from issues such as flaky tests, missing dependencies, ambiguous requirements, or environment-specific behavior.
What We’re Looking For
3+ years of professional software engineering experience
Strong experience working with large, multi-file codebases
Experience reviewing pull requests, debugging issues, and maintaining production code
Strong understanding of unit testing and test coverage
Ability to evaluate whether tests correctly validate a solution without unnecessarily restricting implementation approaches
Experience with dependency management, environment setup, and reproducibility
Strong understanding of Git and GitHub-based development workflows
Ability to analyze complex technical problems and provide clear written feedback
Nice to Have
Contributions to or maintenance of open-source projects
Experience with SWE-Bench, SWE-Bench Verified, or similar coding benchmarks
Experience with major Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest
Experience with Docker, CI/CD, pip, conda, or dependency pinning
Knowledge of test fixtures, test isolation, or property-based testing
Experience designing technical assessments or reviewing coding challenges
Experience with AI/ML evaluation, data curation, RLHF, or benchmark development
What You’ll Be Responsible For
Reviewing coding tasks derived from real GitHub issues and pull requests
Assessing whether problem statements and success criteria are clear and complete
Evaluating unit tests for correctness, coverage, and robustness
Identifying flaky tests, missing dependencies, version conflicts, and environment issues
Determining whether tasks can be reliably reproduced across environments
Assessing the real-world difficulty and complexity of each task
Providing clear recommendations on whether tasks should be accepted, improved, or excluded
Engagement
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Software engineering, open-source code review, testing, and technical evaluation
This role is a strong fit for experienced engineers who enjoy debugging complex codebases, reviewing pull requests, working with open-source software, and evaluating what makes a software engineering problem well designed.
About the Company
More jobs at anyone-ai
-
Application Security Engineer – CVE & Vulnerability Research
Argentina - Fully Remote · contract · Sep 15, 2026
-
Machine Learning Engineer – ML Evaluation & Experiment Design
Argentina - Fully Remote · contract · Sep 15, 2026
-
Python Developer - Paraguay
Paraguay · contract · Sep 14, 2026
-
Backend Developer (Spain)
Spain · contract · Sep 14, 2026
-
Backend Developer (Uruguay)
Uruguay · part_time · Aug 18, 2026