Skip to content
[bmdpat]

Showwork / powered by Jev

Check the proof.

Your agent said “done.” Does its evidence support that claim?

Jev compares the claim, check, and supplied output. You see the assessment and uncertainty. No account required.

Start with an example, or write your own

Illustrative reconstructions of evidence failures and a stronger check. These are not live test receipts or measured Jev results.

What was the agent asked to accomplish?

What did it say was done?

Paste the actual assertion or relevant test code. Leave blank if missing.

Paste what the check observed. A test name alone is not execution evidence.

Submitting sends these four fields to TypeSafe for assessment. Omit secrets and private data. This app does not save your submitted text. TypeSafe processes it under its privacy policy.

Jev / evidence assessment

A passing check can miss the point.

Paste the claim next to what was actually checked. The assessment will appear here.

Claim → check → evidence

Does the scope match?

Advisory only. Jev does not execute tests, inspect your repository, authenticate receipts, or approve a release. Showwork's deterministic checks stay separate.

Make the claim fail when it is false.

Showwork's quickstart walks through a refused close, a fix, and a checked receipt.

Try showwork locally

The pilot results, including the miss

On September 17, 2026, we ran 20 development cases, then 40 held-out cases with the same frozen question and jev-1.13.0. These are authored examples, not an independent sample of real projects. A matching label does not prove that any submitted evidence is authentic.

40-case holdout. Lower is better for false support.
MeasureJevText baseline
False support / 20 inadequate00
Inadequate flagged / 202010
Adequate recognized / 10100
Missing context recognized / 101010

Provider latency: median 178 ms, p95 347 ms. Usage: 28,656 input and 2,300 output tokens across 40 calls. 5 answers required low-confidence review. There were no service errors or holdout label mismatches.

The development miss: a mocked email queue acceptance was labeled a scope gap against the claim that the customer received the email. Jev called it a contradiction with confidence 0.54. It flagged the weak evidence, but disagreed with our label. Because the supplied evidence explicitly said no email left the process, that label boundary is debatable. We kept the original label and made no rubric change.

This small pilot passed its launch thresholds. It does not establish a general error rate or replace independent review. The baseline is deliberately conservative and recognized none of the adequate cases.

Development results and missHoldout matrix, usage and hashesBaseline results and failures