Senior Software Engineer - AI Evaluation / Coding Agents
Remote — North America, LATAM or India
Special Referral gives your application an enhanced referral route, helping it stand out and potentially improving your chances of being considered.
About the role
Turing works with frontier AI labs to accelerate model development through high-quality training data, evaluations and engineering talent.
Experienced, hands-on software engineers will evaluate and improve AI coding models across real-world repositories. Review generated code and agent behavior, determine technical correctness, identify failure modes and create evaluation signals and feedback to improve model performance.
Treat the coding agent as another engineer whose work you are reviewing: assess whether it understands the task, chooses the right approach and produces correct, robust, maintainable code. Explain precisely where it succeeds or fails.
Evaluation process: AI interview (approximately 25 minutes), practical code/AI evaluation exercise (approximately 30 minutes), and hiring manager interview (approximately 20 minutes). The practical exercise focuses on reviewing and evaluating AI-generated code, not competitive programming or algorithm puzzles.
Scope of Work
- Evaluate AI-generated code and solutions across real-world software repositories.
- Review agent behavior, tool usage and code changes for correctness and quality.
- Identify technical errors, weak approaches and recurring model failure modes.
- Compare model outputs and explain why one solution is better than another.
- Create and refine rubrics and evaluation criteria for coding tasks.
- Produce high-quality evaluation and preference data used to improve coding models.
- Build and maintain pipelines and infrastructure supporting data generation, collection and evaluation workflows.
- Synthesize findings from data work into clear write-ups, updates and recommendations for the team.
- Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes.
- Share clear, actionable findings with AI researchers and engineers.
What you’ll bring
- 5+ years of hands-on software engineering experience.
- Strong proficiency in Python, TypeScript/JavaScript, Go or another major production language.
- Experience working in substantial real-world codebases.
- Strong code-review skills and technical judgment.
- Ability to clearly explain why an implementation is correct, incorrect or could be improved.
- Strong written communication.
- Experience using modern LLMs or AI coding tools.
- Located in North America, LATAM or India.
- At least 6 hours of Pacific Time overlap; availability of 40 hours per week is preferred.
Preferred Qualifications
- Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design or post-training is a plus, but not required.
Benefits
Not provided in the source listing.
Schedule and availability
40 hours per week preferred, with at least 6 hours of Pacific Time overlap.
Contract & Payment Terms
- Freelance independent contractor engagement.
- Approximately 3 months.
- Start as soon as possible.
Where you can work
Applicants must be located in North America, LATAM or India. The posting does not provide a complete country-by-country list for North America or LATAM; confirm your country eligibility with Turing.
Country eligibility has not been provided. Confirm with the employer.
Remote does not always mean work from any country. Always check the employer’s location and working-hour requirements.

