We are building an AI-native localization stack: agentic translation workflows that already publish content with minimal human touch, and a quality system that keeps that automation trustworthy at scale. As automation grows, quality assurance becomes the product. We are hiring a Senior PM to own Agent QA — the evaluation, testing, and feedback infrastructure that decides whether an AI translation is good enough to auto-publish, catches localization defects inside the live product, and continuously feeds signal back to improve our agents.
This is a bridge role for someone who is genuinely strong in both: you bring real localization/translation domain depth (you know what makes a translation wrong, what breaks in-product, how terminology and locale rules work) and the AI product instinct to turn that judgment into agents, metrics, and automated tooling. You will define what "good" means, build the tooling to measure it automatically, and make our human experts dramatically more leveraged.
AI Localization Testing Tool
Build an AI-driven product testing tool that automatically detects localization defects that translation review cannot catch — truncation, layout breakage, hardcoded strings, unlocalized images, wrong number/date/currency formats. Ship the MVP that scans mobile/web apps to accelerate localization auditing, then drive the long-term vision of integrating localization tests into the internal product testing infrastructure, so they run automatically before features go live.
Agentic Translation Quality Evaluation
Enhance the quality gate of the agentic pipeline — the quality evaluation agent that decides whether a translation is good enough to auto-publish. Own the trade-off between automation coverage and risk across content tiers. When the gate gets something wrong systematically rather than as a one-off, diagnose the root cause, and drive the fix back into the agent.
Evaluation & Annotation Infrastructure
Build the backbone that makes quality measurable and improvable: golden datasets, annotation tooling with error classification, automated evaluation, a metrics/dashboard layer, and the RLHF feedback loop that turns human evaluations into agent fine-tuning. Own annotation workflow integration (e.g. with internal AI infra) so linguists can annotate in a standardized, real-time way.