Eval Suite & Diagnostics

Benchmark evaluation suite (`tacet-eval`), statistical sign tests, and microsecond diagnostics.

Statistical Benchmark Suite (`tacet-eval`)

`tacet-eval` measures tool selection accuracy across English and Turkish prompt suites under reproducible benchmark conditions.

  • Sign Test Metric: Calculates exact p-values (e.g. p = 1.0) over paired test suites to verify statistically significant model improvements.
  • Noise-Free Evaluation: Disables sampling temperature and random noise during benchmark runs for 100% reproducible results.
  • Detailed Fingerprinting: Benchmark reports record model path, tensor quantization, hardware device, wall time, and active tool catalog fingerprints.

Diagnostic Command: `tacet why`

The `tacet why` command analyzes trigger matching, term boundary scoring, and tool budget allocation in milliseconds without initializing model weights:

tacet why "Dolar kuru şu an ne durumda?"