Featured Jobs
TeracSoftware Engineering Featured JobsSpecial Referral

Software Engineers: Paid Code Review for AI Agent Evaluation

Remote — United States or Canada

Special Referral gives your application an enhanced referral route, helping it stand out and potentially improving your chances of being considered.

Important note: The referral detail describes an approximately 15-minute screening/assessment, not the total project duration. 25 available spots were shown when checked on October 4, 2026 and may change.

About the role

We're running a paid study on the coding environments and programming tasks used to benchmark artificial intelligence agents. Creating robust evaluation harnesses ensures that AI models are tested against realistic software engineering scenarios. This work directly feeds into improving how autonomous agents handle complex coding objectives.

During this remote session, you will review a series of coding tasks and their corresponding evaluation harnesses. You will assess whether the programming challenges accurately reflect real-world software engineering problems. We will ask you to verify the logic, test cases, and overall structure of the environments provided. Your technical feedback will be captured through a guided conversation and screen-sharing exercises.

We are looking for practicing software engineers with strong backgrounds in building and testing complex systems. Candidates should have direct experience writing test suites, evaluation harnesses, or comprehensive code reviews. We welcome backend engineers, full-stack developers, machine learning engineers, and software architects.

Scope of Work

  • Review realistic programming tasks for accuracy and difficulty
  • Verify the logic and test cases within provided evaluation harnesses
  • Assess whether coding environments effectively measure software engineering skills
  • Walk us through your thought process while analyzing complex code structures

What you’ll bring

  • Active professional experience as a software engineer
  • Hands-on experience building or maintaining test suites and evaluation harnesses
  • Comfortable reviewing code and explaining technical concepts aloud
  • Familiarity with complex real-world software architecture
  • Must be 18 years of age or older and fluent in English.
  • Located in the United States or Canada.
  • Currently works as a Software Engineer (backend, full-stack, frontend), ML or AI Engineer, or authors coding tasks for AI evaluation.
  • Has at least 3 years of paid, full-time experience writing production software.
  • Has personally authored coding tasks with issue descriptions and test suites for AI training or evaluation.
  • Must have experience with benchmarks, harnesses, or sandboxing frameworks used to train or evaluate software-engineering agents.
  • Can commit 10 or more hours per week starting next week.

Benefits

Not provided in the source listing.

Schedule and availability

At least 10 hours per week. The source requests a start next week; confirm the actual start date with Terac.

Contract & Payment Terms

  • You will be engaged as an independent contractor.
  • This is a fully remote opportunity that can be completed on your own schedule.
  • Opportunities can be extended, shortened, or concluded early depending on needs and performance.
  • Your participation will not involve access to confidential or proprietary information from any employer, client, or institution.
  • Payments are processed weekly based on services rendered.
  • We are unable to support H1-B or STEM OPT candidates at this time.

Where you can work

This role is open to people based in: United States, Canada.

Remote does not always mean work from any country. Always check the employer’s location and working-hour requirements.

BEYOND THE JOB DESCRIPTION

Step inside the Role Room.

Explore the working day, published expectations, and an optional reflection.