Overview
The role involves evaluating the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. It is not a traditional engineering role but rather focuses on assessing AI-generated coding interactions.
Key Responsibilities
- Evaluate AI-generated coding interactions end-to-end.
- Judge whether outputs are:
- Useful
- Correct (at a high level)
- Aligned with how a strong engineer would think
- Assess the quality of explanations and reasoning, not just code.
- Distinguish between different levels of response quality.
- Provide clear, opinionated feedback on what worked, what didn’t, and what felt misleading.
- Help define what great looks like when interacting with tools like Cursor.
Requirements
- Staff/Principal-level engineer or equivalent experience.
- Strong background in TypeScript/JavaScript or Python.
- Hands-on experience using OpenAI Codex, Claude Code, and Cursor.
- Deep familiarity with modern AI-assisted development workflows.
- Ability to evaluate code without requiring full execution or deep review.
- Comfortable providing direct, opinionated feedback.
- High standards for what constitutes good engineering.
Nice to Have
- Experience with tools like Cursor or similar AI-first IDEs.
- Prior exposure to prompt design or evaluation workflows.
- Experience mentoring senior engineers or defining engineering standards.
Benefits
- Rate: $100–$200/hour
- Hours: 10–20 hours/week
- Duration: Through early May (with possible extension)
- Start: ASAP
Location
Remote
How to Apply
For more details, check the Loom video.
Deadline
No formal deadline stated.