Nexxa
QA Engineer (AI Systems)
SF Bay area · Posted Aug 27, 2026
About the role
Nexxa is building the best AI systems for heavy industries — enabling machines, systems and operations to think, decide and act autonomously across manufacturing, large-scale infrastructure, logistics and legacy environments. Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry. ROLE OVERVIEW We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems — products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions. You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously. KEY RESPONSIBILITIES - Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence. - Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts. - Define and track quality metrics beyond simple accuracy — groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations. - Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD. - Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams. - Test agent behavior across the full action loop — planning, tool selection, tool execution, error recovery, and final output — not just the final response. - Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic. - Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports. - Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments. - Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems. - Advocate for testability and observability in agent architecture from day one. QUALIFICATIONS - 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems. - Hands-on experience testing LLM-based products, chatbots, or AI agents — you understand why traditional deterministic test assertions break down for generative systems. - Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own. - Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling. - Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks. - Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time. - Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism. - Comfortable operating in ambiguity — defining what "correct" means for a task when there's no single right answer. - Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders. PREFERRED - Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rub
Apply in minutes, not hours
Godspeed drafts a cover letter, tailors your resume to this role, and finds the hiring manager's email — on a curated feed of jobs worth your time.
Start for free