Skip to main content
J

AI Safety Policy Evaluator, Violence & Threats

Jobgether
6 days ago
Full-time
Remote
Worldwide
Remote Legal
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Policy Evaluator, Violence & Threats based in United States. This role sits at the intersection of AI safety, content policy, and expert human judgment, helping improve how advanced AI models handle violent and threatening content.You’ll evaluate user requests, model responses, and conversation context to distinguish legitimate fictional, educational, historical, or defensive content from material that could enable real-world harm.The work focuses heavily on nuanced edge cases where intent, context, and a single detail can materially change the appropriate policy decision.You’ll contribute not only to evaluations, but also to adversarial testing, policy refinement, calibration, and the identification of emerging safety gaps.This is a non-engineering role for someone with deep experience in areas such as violent fiction, military or emergency response, crisis intervention, threat assessment, trust and safety, or related fields.You’ll work in a structured, feedback-rich environment where clear reasoning, consistency, and the a... Click Apply to read the full job description.