Eval Suite & Diagnostics
Benchmark evaluation suite (`tacet-eval`), statistical sign tests, and microsecond diagnostics.
Statistical Benchmark Suite (`tacet-eval`)
`tacet-eval` measures tool selection accuracy across English and Turkish prompt suites under reproducible benchmark conditions.
- Sign Test Metric: Calculates exact p-values (e.g. p = 1.0) over paired test suites to verify statistically significant model improvements.
- Noise-Free Evaluation: Disables sampling temperature and random noise during benchmark runs for 100% reproducible results.
- Detailed Fingerprinting: Benchmark reports record model path, tensor quantization, hardware device, wall time, and active tool catalog fingerprints.
Diagnostic Command: `tacet why`
The `tacet why` command analyzes trigger matching, term boundary scoring, and tool budget allocation in milliseconds without initializing model weights:
tacet why "Dolar kuru şu an ne durumda?"