ConciergedemoSource

How it was tested

54 test conversations, every one shown.

30 of them are attacks: prompt injection, someone else's bookings, fake admin claims, and pressure over several messages. Each case says what counts as correct. Measured on 2026-10-06 against a throwaway database, with the same cases run on two models. Source and raw results are in the repository.

Claude Opus 5

shipped model
52 of 54
cases passed
29 of 30
safety cases passed
0
refunds paid without a person
$0.0082
cost per conversation
2.1 s
median turn time
4.9 s
slowest-case turn time (p95)

Claude Haiku 4.5

comparison
48 of 54
cases passed
25 of 30
safety cases passed
0
refunds paid without a person
$0.0012
cost per conversation
1.8 s
median turn time
4.1 s
slowest-case turn time (p95)

Read these numbers honestly

Finding the right help-center answer

Before the model answers, the system looks up the help-center sections that might contain the answer. This part is free to test, so it runs on every change. “Right section” counts a hit only when a retrieved piece is the section that actually answers the question, not just any section of the right article.

Search modeRight article in top 4Right section in top 4Average rank score
hybrid1.0000.9760.899
vectorshipped1.0000.9760.927
keyword0.8570.8100.766

Every test case

Showing 54 of 54