Evaluation record
How reliable are the suggestions?
These recorded tests check how EvidenceDesk finds documentation and generates answers. Correct quotations do not prove correct answers; human review of answer quality is still pending.
Human review pendingGenerated responses passed these checks. This is not an accuracy score.
Failed checks or answers that disagreed with the expected result.
No human answer-quality score has been assigned.
Live model evaluation
Real Cloudflare calls · Recorded run40 of 40 selected cases processed (all split). 30 generation attempts; 30 completed provider responses. 28 passed structured-output and citation-integrity checks.
Models and test details
Finding the right documentation
All methods used the same real BGE query embeddings and eligible corpus. Duplicate chunks from one document count once. Unlabeled questions do not contribute to retrieval scores.
| Method | Cases | Recall@5 | MRR |
|---|---|---|---|
| lexical | 16 | 1.000 | 0.703 |
| vector | 16 | 1.000 | 0.922 |
| hybrid | 16 | 1.000 | 0.875 |
Recall@5 measures whether the expected document appears in the first five results. MRR measures how early it appears. Neither measures answer correctness.
Generated-answer failures
credential-rotate-1— Rejected: invalid_model_output. No validated answer.unknown-region-1— Answer status disagreed with the expected abstention.obsolete-vs-current-1— Answer status disagreed with the expected abstention.obsolete-vs-current-2— Rejected: invalid_model_output. No validated answer.foreign-routing-2— Answer status disagreed with the expected abstention.foreign-export-1— Answer status disagreed with the expected abstention.foreign-export-2— Answer status disagreed with the expected abstention.
Valid source IDs and exact quotes do not establish semantic correctness. No human quality score has been assigned.
Timing and scope
Latency n=30: p50 1.15 s; p95 2.44 s. Includes embedding, three retrieval methods, optional generation, and failed cases. These samples describe this run only.
Exact cost is not reported. The full dataset has 40 cases: 24 development and 16 holdout. Ten action/provider/malformed cases refer to separate engineering tests.
Fixture / offline engineering evaluation
No live model calledDeterministic test vectors measure the fixture retrieval path. These numbers are separate from live embedding and generated-answer quality.
| Method | Cases | Recall@5 | MRR |
|---|---|---|---|
| lexical | 16 | 1.000 | 0.703 |
| vector | 16 | 0.000 | 0.000 |
| hybrid | 16 | 1.000 | 0.703 |
Fixture provenance and report
40 cases processed; control cases retain engineering-test references. No live generation or latency is measured.
Read the fixture report ↗What remains to be reviewed
- A human must assess correctness, completeness, supported recommendations, and appropriate abstention. Review status remains pending.
- The cosine cutoff is a starting heuristic, not a calibrated factuality guarantee. Retrieval can surface passages that do not answer the question.
- Approval, isolation, stale proposals, quotas, and provider failures are covered by separate PostgreSQL/API tests. Fixture tests are never counted as live model passes.
- Free-tier limits can make analysis temporarily unavailable. The application never substitutes a successful simulated answer for failed live inference.