Overview
The role involves evaluating the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is a contract position that requires a highly experienced software engineer, focusing on assessing AI coding interactions rather than traditional engineering tasks.
Key Responsibilities
- Evaluate AI-generated coding interactions end-to-end.
- Judge whether outputs are useful, correct (at a high level), and aligned with strong engineering thought.
- Assess the quality of explanations and reasoning, not just code.
- Distinguish between different levels of response quality.
- Provide clear, opinionated feedback on what worked, what didn’t, and what felt misleading.
- Help define what great interactions look like with AI tools.
Requirements
- Staff/Principal-level engineer (or equivalent experience).
- Strong background in TypeScript, JavaScript, or Python.
- Hands-on experience using OpenAI Codex, Claude Code, and Cursor.
- Deep familiarity with modern AI-assisted development workflows.
- Able to evaluate code without fully executing or deeply reviewing every line.
- Comfortable providing direct, opinionated feedback.
- High bar for what “good engineering” looks like.
Benefits
- Rate: $100–$200/hour.
- Hours: Approximately 10–20 hours/week.
- Contract duration: Through early May (with possible extension).
- Start date: ASAP.
Location
Remote
How to Apply
Please refer to the company's application process provided in their job listing.
Deadline
No specific deadline stated.