Research
Test: retrieval, source quality, citation correctness, context handling, synthesis and whether the artifact can be reused.
Do not infer from: one fluent answer or a fast search demo.
TOOL LAB / TEST THE JOB
Tool Lab starts with a task specimen: same input, same criteria, same conditions. We only say a tool fits a job when the evidence supports that scope; without a benchmark, the page says what still needs to be measured.
JOB BENCH
Research does not fail like coding; coding does not fail like automation. A universal leaderboard often hides the question that matters: where will this tool fail in your work?
Test: retrieval, source quality, citation correctness, context handling, synthesis and whether the artifact can be reused.
Do not infer from: one fluent answer or a fast search demo.
Test: patch quality, tests, repo context, diff quality, latency, cost and recovery when the agent goes wrong.
Do not infer from: a clean-repo new-file demo.
Test: prompt adherence, consistency, editability, export, speed and cost on your actual asset class.
Do not infer from: the best image selected for a gallery.
Test: triggers, auth, retry, observability, data sensitivity and human approval points.
Do not infer from: integration count or the word “agentic.”