Evaluation record

How reliable are the suggestions?

These recorded tests check how EvidenceDesk finds documentation and generates answers. Correct quotations do not prove correct answers; human review of answer quality is still pending.

Human review pending
Format and source checks28 / 30

Generated responses passed these checks. This is not an accuracy score.

Cases needing attention7

Failed checks or answers that disagreed with the expected result.

Human quality reviewPending

No human answer-quality score has been assigned.

Live model evaluation

Real Cloudflare calls · Recorded run

40 of 40 selected cases processed (all split). 30 generation attempts; 30 completed provider responses. 28 passed structured-output and citation-integrity checks.

Models and test details
Generation model
@cf/meta/llama-3.1-8b-instruct-fast
Embedding model
@cf/baai/bge-small-en-v1.5 · 384 dimensions
Tested revision
b1943513cef1
Run date / prompt
2026-09-10 · support-v3
Corpus version / hash
1 · 16e987278345dbf6
Dataset hash
cd577f8a745391e1

Finding the right documentation

All methods used the same real BGE query embeddings and eligible corpus. Duplicate chunks from one document count once. Unlabeled questions do not contribute to retrieval scores.

Document-level retrieval · 16 labeled answerable cases
MethodCasesRecall@5MRR
lexical161.0000.703
vector161.0000.922
hybrid161.0000.875

Recall@5 measures whether the expected document appears in the first five results. MRR measures how early it appears. Neither measures answer correctness.

Generated-answer failures

  • credential-rotate-1 Rejected: invalid_model_output. No validated answer.
  • unknown-region-1 Answer status disagreed with the expected abstention.
  • obsolete-vs-current-1 Answer status disagreed with the expected abstention.
  • obsolete-vs-current-2 Rejected: invalid_model_output. No validated answer.
  • foreign-routing-2 Answer status disagreed with the expected abstention.
  • foreign-export-1 Answer status disagreed with the expected abstention.
  • foreign-export-2 Answer status disagreed with the expected abstention.

Valid source IDs and exact quotes do not establish semantic correctness. No human quality score has been assigned.

Timing and scope

Latency n=30: p50 1.15 s; p95 2.44 s. Includes embedding, three retrieval methods, optional generation, and failed cases. These samples describe this run only.

Exact cost is not reported. The full dataset has 40 cases: 24 development and 16 holdout. Ten action/provider/malformed cases refer to separate engineering tests.

Full run after a bounded smoke check. Historical failures and the ten-case smoke remain available in the report history. The full run used an explicit maintainer budget within the unchanged 40-attempt environment cap. Public sandbox limits were unchanged.
Read the full live report ↗

Fixture / offline engineering evaluation

No live model called

Deterministic test vectors measure the fixture retrieval path. These numbers are separate from live embedding and generated-answer quality.

Document-level retrieval · 16 labeled answerable cases
MethodCasesRecall@5MRR
lexical161.0000.703
vector160.0000.000
hybrid161.0000.703
Fixture provenance and report
Generation model
fixture-v1
Embedding model
fixture-hash-v1 · 384 dimensions
Tested revision
31696bab8c6f
Run date / prompt
2026-09-09 · support-v1
Corpus version / hash
1 · 16e987278345dbf6
Dataset hash
cd577f8a745391e1

40 cases processed; control cases retain engineering-test references. No live generation or latency is measured.

Read the fixture report ↗

What remains to be reviewed

  • A human must assess correctness, completeness, supported recommendations, and appropriate abstention. Review status remains pending.
  • The cosine cutoff is a starting heuristic, not a calibrated factuality guarantee. Retrieval can surface passages that do not answer the question.
  • Approval, isolation, stale proposals, quotas, and provider failures are covered by separate PostgreSQL/API tests. Fixture tests are never counted as live model passes.
  • Free-tier limits can make analysis temporarily unavailable. The application never substitutes a successful simulated answer for failed live inference.