Data Scientist β Agent Evaluations & Quality
Clera
ABOUT THE ROLE We're a small, fast-moving AI productivity startup (~25 people) building an autonomous AI executive assistant that operates across email, calendars, meetings, and business software. This role owns the measurement system that determines whether our AI agent is genuinely improving in ambiguous, real-world environments. You'll partner closely with AI Agent Capabilities engineers to produce the evidence that drives product decisions, model choices, and release quality β turning hard questions about agent behavior into rigorous, actionable answers. Visa sponsorship is not available for this role. WHAT YOU'LL DO - Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces. - Translate product capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks. - Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios. - Define and track metrics including task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability. - Design deterministic and model-based graders, calibrate LLM-as-a-jud... Click Apply to read the full job description.