How We Test

A test should leave an evidence trail. This is the framework ChalkBench intends to apply to hands-on reviews; it is not a claim that any product has already been tested.

1. Define the classroom job

Each test begins with a specific educator workflow and a realistic success criterion. Record the subject, grade range, learning goal, output format and time constraints. Use fictional student examples or approved material. Never upload identifiable student records simply to evaluate a tool.

2. Record the conditions

Identify the product, plan, test date, device, browser or operating system where relevant, and material usage limits. Retain prompts and representative outputs when permitted. If an integration is only described in vendor documentation, label it as documented rather than tested.

3. Review the output and the work left

Check facts, answer keys, relevance, accessibility and editing needs. Inspect exports in the format an educator would use. Repeat representative tasks to avoid presenting one unusually strong result as typical. Report substantial errors and areas where the test cannot support a conclusion.

4. Measure the complete workflow

Timing should include setup, generation, review, correction and export. A claim of time saved needs a stated comparison method and the same end task. If no reliable baseline is available, report elapsed time without inventing a saving.

The proposed ChalkBench Score

For sufficiently tested products, each dimension is rated from 0 to 10 using documented observations. Multiply each rating by its percentage weight and add the contributions to produce an overall score out of 10. Publish the component ratings and explain the result. Do not issue an overall score if material evidence is missing.

Review dimensions and weights
DimensionWeightWhat it examines
Output quality25%Accuracy, relevance and completeness for the defined task.
Classroom readiness20%Editing needs, accessibility and practical usability.
Time saved15%End-to-end time against a documented comparable workflow.
Ease of use15%Setup, navigation, editing and export.
Privacy10%Clarity of data practices, controls and school-relevant limitations.
Free plan5%Whether meaningful tasks can be completed within its limits.
Integrations5%Documented and tested connections to educator workflows.
Value5%Usefulness relative to the checked price and alternatives.

Rating anchors and important limits

A 0 means the dimension fails the defined requirement; 5 means it is usable with significant compromises; 10 means it meets the documented test criteria exceptionally well. Intermediate ratings require an explanation. Missing evidence is “not assessed,” not a zero and not an invented midpoint. The privacy dimension is an editorial assessment of available information, not a security audit or legal certification.

Hands-on is not the same as classroom-tested

A reviewer using a tool with synthetic examples is hands-on testing. Classroom testing requires actual use in a relevant educational setting with appropriate permissions. “Teacher Tested” requires a real teacher to have performed the stated test. These labels are not interchangeable.

Hardware and troubleshooting

The AI-tool weighting does not automatically suit a printer, projector or laptop. Hardware coverage needs task-specific criteria, device details and a disclosed method. Troubleshooting should use current vendor documentation, start with reversible checks and identify when school IT must take over.

Corrections, retests and commercial access

Record pricing checks separately from test dates. Update or qualify findings when product changes affect them. Disclose vendor-provided access and affiliate relationships; neither can buy a score. Read the editorial policy and affiliate disclosure.