This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Automation QA (JavaScript) Engineer based in India.
This role offers the opportunity to shape quality engineering for a next-generation enterprise Agent Development Platform.
You will own automation across frontend, backend, workflows, security boundaries, reliability, and performance.
The position focuses on testing AI-driven, inherently non-deterministic systems where statistical quality matters as much as traditional assertions.
You will build robust automation that validates reliability, tenant isolation, sandboxing, resilience, and customer-facing functionality.
Working closely with AI, platform, and evaluation engineers, you will embed quality and testability directly into product design.
Your test suites will become release gates and generate evidence used for critical delivery and client milestones.
This is a high-impact environment where reusable frameworks, adversarial testing, and engineering judgment directly influence production readiness.
Accountabilities
Design, develop, maintain, and scale automated test suites covering frontend and backend functionality, including contract testing for evolving APIs and SDKs.
Build automated validation for long-running, asynchronous workflows, including failure injection involving worker crashes, provider outages, retry storms, and recovery scenarios.
Verify durable-execution guarantees such as zero lost workflow runs, idempotency, eventual consistency, and correct compensation or saga behavior.
Automate multi-tenant isolation testing, including cross-tenant data access, configuration leakage, and accurate cost attribution.
Test agent sandboxing and egress controls through adversarial validation of unauthorized tool calls, network destinations, and data-exfiltration scenarios.
Develop testing strategies for LLM-driven and non-deterministic behavior using statistical assertions, repeated-run consistency, semantic similarity, confidence thresholds, and flakiness quarantine.
Build and expand adversarial test datasets covering prompt injection, tool-call hijacking, document-based attacks, and data-exfiltration attempts aligned with recognized AI security practices.
Validate Human-in-the-Loop workflows and ensure confidence thresholds and authorization gates operate correctly before irreversible or restricted actions are executed.
Partner with evaluation teams to develop golden datasets, synthetic-data pipelines, test infrastructure, and CI integrations that make AI quality measurable and repeatable.
Automate performance and reliability verification for invocation latency, concurrency, queue backpressure, provider failover, and other platform non-functional requirements.
Use observability and tracing data, including Langfuse and OpenTelemetry, to validate trace completeness, token and cost accounting, and anomaly detection.
Integrate automation into CI/CD pipelines as hard release gates with automated regression detection and per-metric reporting.
Produce reproducible and auditable test evidence that supports release decisions and client milestone sign-offs.
Maintain test environments, mocked LLM and provider layers, and synthetic data generators to keep testing efficient, deterministic where possible, and cost-effective.
Collaborate with AI and platform engineers from the design stage to establish acceptance criteria, observability, and testability requirements.
Mentor QA engineers, develop reusable AI testing frameworks and patterns, and contribute to broader quality engineering practices.
8+ years of experience in test automation, API testing, and building reusable test frameworks adopted by other engineers.
Strong hands-on experience testing LLM-powered or other non-deterministic systems, including statistical assertions, semantic scoring, model-variance management, and AI-specific regression testing.
Mandatory proficiency in Playwright and TypeScript for developing sophisticated automated test suites.
Extensive experience building automation frameworks and integrating them into secure, high-quality CI/CD environments.
Strong understanding of LLM behavior, prompt sensitivity, RAG failure modes, and common reliability challenges in AI systems.
Experience testing asynchronous and event-driven architectures, including failure injection, idempotency, eventual consistency, and workflow recovery; experience with Temporal or a similar workflow engine is highly desirable.
Strong CI/CD expertise, particularly with GitHub Actions or equivalent platforms, and experience implementing automated release gates.
Solid SQL and data-validation capabilities, including experience creating synthetic datasets and evaluating dataset quality.
Strong security-testing and adversarial mindset, with familiarity with prompt injection, application security testing, or a demonstrated willingness to specialize in AI security.
Strong statistical and analytical skills, with the ability to interpret evaluation metrics and provide actionable recommendations to AI engineering teams.
Demonstrated judgment in determining appropriate quality thresholds under delivery and milestone pressure, with the ability to communicate and defend test evidence to technical and non-technical stakeholders.
Experience with API automation and validation across complex enterprise platforms.
Excellent collaboration and communication skills, with the ability to work effectively with AI engineers, platform engineers, product teams, and evaluation specialists.
Experience with performance and load-testing tools such as k6 or Locust is an advantage.
Experience testing multi-tenant SaaS isolation is preferred.
Familiarity with AI evaluation and observability tools such as Langfuse, LangSmith, Promptfoo, DeepEval, or Arize Phoenix is beneficial.
Experience with compliance evidence such as SOC 2 or GDPR, or accessibility testing using Section 508/WCAG, is a plus.
Familiarity with commercial real estate workflows, lease accounting, CAM reconciliation, OCR, or document-extraction systems is advantageous.
Experience using LLMs to generate test cases, adversarial payloads, or synthetic documents is desirable.
Experience with high-concurrency background processing within agentic systems and familiarity with qTest are additional advantages.
Fully remote opportunity based in India.
Work on a high-impact enterprise AI and Agent Development Platform.
Exposure to cutting-edge technologies across AI, LLMs, agentic workflows, automation, observability, and cloud-native engineering.
Collaboration with experienced AI, platform, evaluation, and quality engineering professionals.
Opportunities to work on technically challenging reliability, security, performance, and non-deterministic testing problems.
Strong professional development opportunities through internal meetups, conferences, workshops, and knowledge-sharing initiatives.
Access to Udemy and language-learning programs.
Company-supported professional certifications.
Opportunities for internal mobility and exposure to diverse technology domains.
Company-paid medical insurance.
Mental health support.
Financial and legal consultation services.
Collaborative, open-door working environment focused on continuous learning and career growth.

