面向 AI 智能体开发团队的自动化质量测试平台,检测幻觉、评分正确性并验证策略合规,避免问题回复上线。
热门评论
PH 用户
嗨 Product Hunt!👋
今天非常激动地跟大家分享 QAgent。
现代 AI 工程有个肮脏的小秘密: 我们会给一个 20 行的后端函数写 50 个单元测试。但当我们上线一个处理真实客户、非确定性的 AI 智能体时,我们的“测试流水线”就是往 OpenAI playground 里敲 3 个 prompt,看它礼貌地回一句,然后点部署。
问题在哪?Prompt 回归是无声的。你为了修复 edge case A,在 system prompt 里改了一句话,结果它悄悄搞坏了另外 3 个本来正常的客户流程,连一个 runtime error 都不报。⟪NL
PH 用户
"Stop shipping on vibes" is uncomfortably accurate for where most of us are with agent QA right now - regression testing against ground truth instead of eyeballing transcripts is exactly the gap. Scoring policy adherence specifically is the part I'd actually pay for, since that's the failure mode that's hardest to catch by just reading a few sample outputs.
今天非常激动地跟大家分享 QAgent。
现代 AI 工程有个肮脏的小秘密:
我们会给一个 20 行的后端函数写 50 个单元测试。但当我们上线一个处理真实客户、非确定性的 AI 智能体时,我们的“测试流水线”就是往 OpenAI playground 里敲 3 个 prompt,看它礼貌地回一句,然后点部署。
问题在哪?Prompt 回归是无声的。你为了修复 edge case A,在 system prompt 里改了一句话,结果它悄悄搞坏了另外 3 个本来正常的客户流程,连一个 runtime error 都不报。⟪NL