Evaluation

V1 → V2 → V3 on the 120-case development set, across quality, cost, and latency. Every value below is read from the real, persisted comparisons — nothing is hand-typed.

Answer correctness

Higher is better

Faithfulness

Higher is better

Hallucination rate

Lower is better

Abstention accuracy

Higher is better

Cost per query

Lower is better — generator (model inference) cost only

p95 latency

Lower is better

Full metric table (dev, 120 paired cases)

V1/V2 from the real v1_vs_v2_dev comparison; V3 from the real v1_vs_v3_dev comparison.

MetricV1V2V3
Answer correctness73.15%73.37%76.71%
Faithfulness82.99%72.22%86.07%
Hallucination rate5.00%1.67%1.69%
Instruction following53.33%53.33%60.00%
Abstention accuracy6.67%93.33%93.33%
Prompt injection block rate100.00%100.00%100.00%
Sensitive-information protection100.00%100.00%100.00%
Benign false positives (guardrail suite)55.56%55.56%11.11%
Benign false positives (all non-attack cases)12.12%12.12%6.06%
p50 latency524 ms467 ms523 ms
p95 latency1194 ms1315 ms1036 ms
Cost per query$0.0823 / 1k queries$0.0756 / 1k queries$0.1026 / 1k queries
Failure rate0.00%0.00%0.00%