Showwork / powered by Jev
Check the proof.
Your agent said “done.” Does its evidence support that claim?
Jev compares the claim, check, and supplied output. You see the assessment and uncertainty. No account required.
Start with an example, or write your own
Illustrative reconstructions of evidence failures and a stronger check. These are not live test receipts or measured Jev results.
Jev / evidence assessment
A passing check can miss the point.
Paste the claim next to what was actually checked. The assessment will appear here.
Claim → check → evidence
Does the scope match?
Advisory only. Jev does not execute tests, inspect your repository, authenticate receipts, or approve a release. Showwork's deterministic checks stay separate.
Make the claim fail when it is false.
Showwork's quickstart walks through a refused close, a fix, and a checked receipt.
The pilot results, including the miss
On September 17, 2026, we ran 20 development cases, then 40 held-out cases with the same frozen question and jev-1.13.0. These are authored examples, not an independent sample of real projects. A matching label does not prove that any submitted evidence is authentic.
| Measure | Jev | Text baseline |
|---|---|---|
| False support / 20 inadequate | 0 | 0 |
| Inadequate flagged / 20 | 20 | 10 |
| Adequate recognized / 10 | 10 | 0 |
| Missing context recognized / 10 | 10 | 10 |
Provider latency: median 178 ms, p95 347 ms. Usage: 28,656 input and 2,300 output tokens across 40 calls. 5 answers required low-confidence review. There were no service errors or holdout label mismatches.
The development miss: a mocked email queue acceptance was labeled a scope gap against the claim that the customer received the email. Jev called it a contradiction with confidence 0.54. It flagged the weak evidence, but disagreed with our label. Because the supplied evidence explicitly said no email left the process, that label boundary is debatable. We kept the original label and made no rubric change.
This small pilot passed its launch thresholds. It does not establish a general error rate or replace independent review. The baseline is deliberately conservative and recognized none of the adequate cases.
Development results and missHoldout matrix, usage and hashesBaseline results and failures