This is a 6-month contract role under Optimum Solutions. This role supports our client, a leading Southeast Asian superapp that offers a wide range of services, including mobility, food delivery, and digital payments. Work Schedule and Hours: Monday - Friday, 10am - 7pm (inclusive of a 1-hour lunch break) Key Responsibilities: • Take ownership of the ACS embedded-evaluation workstream. • Establish reliable regression evaluation for at least one additional agent team. • Make its distributed traces consistently usable for debugging and automated evaluation. • Ship at least one reusable improvement to EvalsHub or its SDK/tooling. • Reduce manual evaluation effort through automated datasets, evaluators, and CI integration. • Document ownership, operational procedures, and onboarding guidance. • Embed with AI product teams and own measurable agent-quality outcomes. • Build and maintain golden datasets, regression suites, evaluation pipelines, and quality dashboards. • Convert product rules and human review procedures into deterministic, numeric, LLM-as-judge, trajectory, tool-selection, SOP-adherence, and multi-turn evaluators. • Diagnose evaluation and production-trace failures, then convert findings into dataset improvements, evaluator changes, or fixes to prompts, tools, skills, and agent workflows. • Instrument distributed agent systems using OpenTelemetry, including Temporal workflows and Go/Python services. • Ensure traces contain consistent inputs, outputs, tool calls, metadata, feedback, and root-agent results. • Develop reusable capabilities across EvalsHub backend, SDK, CLI/plugins, and supporting frontend interfaces. • Integrate evaluations into CI/CD and model-release workflows. • Run model, retrieval, embedding, and agent-architecture experiments; balance accuracy, latency, and cost. • Drive migrations from legacy evaluation and observability systems. • Write technical designs and documentation, run demonstrations, and enable engineering teams to use evaluation tooling independently. • Participate in the LLMOps operational rotation and support production-quality integrations. Qualifications: • Minimum 2 years of relevant experience in one or more of the following areas: > EvalsHub, LangSmith, LangGraph/LangChain, FastAPI, Temporal, Grafana, Redis, or similar technologies. > React/TypeScript and internal developer-tooling interfaces. > LLM-as-judge, multi-turn evaluation, tool/MCP evaluation, and agent-trajectory analysis. > Text-to-SQL/Text-to-DSL, RAG, semantic retrieval, embedding evaluation, or model benchmarking. > Building platform capabilities that can be reused across several AI products. • Strong analytical and problem-solving skills related to AI systems, evaluation frameworks, and data-driven quality improvements. • Excellent project management and cross-department communication abilities. • A detail-oriented and metrics-driven approach to continuous improvement