Overview
The role involves evaluating the quality of interactions with AI coding agents such as OpenAI Codex and Claude Code. This position is not focused on traditional engineering tasks but rather on assessing engineering judgments and responses.
Key Responsibilities
- Evaluate AI-generated coding interactions end-to-end.
- Judge outputs for usability, correctness (at a high level), and alignment with strong engineering thought.
- Assess quality of explanations and reasoning alongside code.
- Distinguish between different levels of response quality.
- Provide clear feedback on performance and areas for improvement.
- Help define best practices for AI interaction tools.
Requirements
- Staff / Principal-level engineer or equivalent experience.
- Strong background in TypeScript/JavaScript or Python.
- Hands-on experience with OpenAI Codex, Claude Code, and Cursor.
- Deep familiarity with modern AI-assisted development workflows.
- Able to provide direct, opinionated feedback.
- High standards for engineering quality.
Benefits
- Flexible hourly rate of $100–$200/hour.
- Opportunity to work around 10–20 hours per week, with the duration lasting through early May and a possible extension.
- Engagement starts ASAP.
Location
Remote
How to Apply
Details regarding application can be found via the provided Loom link.
Deadline
ASAP