Evidence policy v1.0Updated 02.10.2026

How we assess AI recorders and mini PCs

Our evidence policy distinguishes personal experience, manufacturer figures and outside testing. This version describes how we report evidence and the conditions required for a repeatable comparison; it does not claim that every product has completed a shared bench protocol.

01

How we buy and set up products

An article should identify the product variant, configuration and information source. For a hardware test, record the operating system, firmware, drivers and subscription tier before collecting results. A setting change can explain a performance difference as easily as a different product can.

02

How we test AI voice recorders

A repeatable recorder comparison needs the same speech, room, placement and transcription language. Accuracy should be scored against a checked reference transcript, with speaker-attribution errors reported separately. Battery claims require a timed recording rundown. Convenience and summary usefulness are qualitative judgements and should be described as such.

03

How we test AI mini PCs

A repeatable inference comparison needs the same model file, quantisation, context, prompt and backend. Prompt processing and output generation should be recorded separately. Power and sound readings require the measurement conditions. Results from different operating systems or memory configurations should not be presented as a clean head-to-head.

04

How scores are weighted

The Testbench score is David Wilson’s editorial rating out of 10. It summarises the article’s recommendation, weighing the intended workload, compatibility, ownership cost and relevant limitations. It is not a measured lab result, and no weighting formula is applied to produce it; readers can see the reasoning in the article.

05

Test rigs and software versions

Configuration details belong beside any numerical result. A reproducible record includes RAM capacity and arrangement, power mode, BIOS, graphics driver, application build and model settings. Our current launch copy does not publish an invented test-rig inventory or treat source specifications as measurements.

06

Protocol changelog

Evidence policy v1.0 — 2 October 2026. Launch coverage separates editorial preference from measured results, labels advertised maximums and links to source documentation. Future benchmark datasets will include the conditions needed to interpret and reproduce them.

Frequently asked questions

Do you test every product the same way?

A fair measured comparison requires matching conditions. The launch coverage combines hands-on opinion and researched assessments, and does not claim that every product completed the same test.

Can I reproduce your results?

A numerical result is reproducible only when its hardware, software and workload settings are supplied. The launch copy avoids publishing figures without those conditions.

Why did a score change?

The Testbench score is an editorial rating, not a lab measurement. A score or recommendation can change when a price, configuration, software feature or supporting evidence changes.