Provenance you re-run.
Obsign is a Python SDK, a CLI, a native core + OFX plugin, an evidence workstation, and an independent browser verifier. Create a signed case package; let any recipient verify it locally.
Quickstart
The SDK is in private preview -- request access. It needs only numpy
and cryptography (torch, GPU, C2PA are optional extras). Once you have it:
# health-gate obsign selfcheck # conformance + signing + determinism, exit 0 if HEALTHY # accountable enhance (deterministic, reversible, signed) + a self-verifying certificate obsign accountable-enhance photo.png --out out.png --receipt r.json --sign key.json --certificate obsign verify-accountable r.json --input photo.png # signed, self-contained evidence package obsign case-create photo.png --out CASE-2026-042 --case-id 2026-042 --examiner "A. Rivera" --key examiner-key.json obsign case-verify CASE-2026-042
Everything is reproducible: the same input + the same quantized model gives a bit-identical output and a receipt that re-runs on any machine.
The CLI
One accountable-transform primitive runs across modalities -- image, video, audio, and document -- anything with a deterministic canonical form. obsign assess verifies any of them.
Receipt format
A receipt is canonical JSON. The receipt_sha256 is SHA-256 over the claim
(everything except the signature and the hash itself, keys sorted, no whitespace). The signature is
Ed25519 over that hash string. Array hashes are SHA-256 over the int64, C-contiguous bytes.
{
"spec": "obsign/accountable/v1",
"producer": { "name": "obsign", "version": "1.0.0", "spec": "obsign/spec/1" },
"tier": "A", "guarantee": "bit-exact deterministic (cross-device)",
"model_sha256": "...", "scales": { "activation": 1024, "weight": 1024 },
"original_sha256": "...", "enhanced_sha256": "...", "diff_sha256": "...",
"accountability": { "mean_change": .., "max_invention": .., "reversible": true },
"receipt_sha256": "<sha256 of the canonical claim>",
"signature": { "alg": "ed25519", "signer": "..", "public_key": "..", "sig": ".." }
}
Verification workflow
Drop a complete .obcase into the package verifier. It checks every layer locally and returns exact findings.
obsign case-verify evidence.obcase verifies an archive without buying the creator or trusting a server.
For a standalone JSON receipt, use the receipt verifier for integrity, signature, issuer status, and optional source binding.
The desktop creator keeps intake, signing, verification, export, and case data on the examiner's Windows workstation.
A signature proves who signed and that nothing changed. Whether to trust that signer is the issuer-trust step -- see the security model.
Provability tiers
Any model fits; the guarantee degrades gracefully, never to zero.
- Tier A -- bit-exact. Quantized integer inference, identical on CPU, GPU and native. Re-run anywhere.
- Tier B -- reproducible. Any model, fixed seeds + deterministic ops, wrapped accountable.
- Tier C -- accountable. Even a black-box generator: preserved original, signed diff, coherence + invention maps.
AI-detection: every number, and how to re-derive it
A detector does not have “an accuracy”. It has a curve, and a threshold picks a point on it. Publishing one number without saying which point it came from is how a figure goes stale on a website while the lab notebook moves on, so all three points ship, selectable, each reporting its own cost with the verdict.
Engine of record: aigen_clip_v9 (sha256 97335b18…f50638,
recorded in obsign/models/MANIFEST.json and verified before every load).
Measured on Synthbuster+ held out: 1800 fakes across 9 generator families unseen in
training (DALL·E 2, DALL·E 3, SDXL, SD2, SD3‑medium,
FLUX.1‑dev, FLUX.1‑schnell, Firefly, Imagen3; 200 each) and
200 pristine RAISE DSLR photographs.
| operating point | threshold | AI caught (unseen) | real photos flagged as AI | balanced |
|---|---|---|---|---|
high_recall |
0.496 | 95.1% | 46% | 0.7453 |
calibrated (default) |
0.570 | 93.6% | 38.5% | 0.7756 |
balanced |
0.814 | 80.2% | 11.5% | 0.8433 |
Per‑generator at the high‑recall point: DALL·E 3 100%, SDXL 99.5%, Firefly 98.5%, SD3 97.5%, FLUX.1‑schnell 97%, SD2 96.5%, FLUX.1‑dev 94.5%, Imagen3 87.5%, DALL·E 2 84.5%. The real set is 200 images, so the false‑positive figure carries roughly ±3.5pp of sampling error at one sigma. We would rather say that than round it off.
It is white‑box evadable. A PGD attack by an attacker holding the weights drives detection from 87.5% to 0.0% at an imperceptible L‑∞ budget of 1/255. We built the standard input‑transform defense, measured that it does not work, and did not ship it. So a detection flag is calibrated evidence, never proof, and the absence of a flag is not authenticity. The path an adversary cannot forge is a signed provenance receipt.
Nothing above has to be taken on trust. The per‑image scores are published, so every number on this page is re‑derivable in about a second, without a GPU:
python tools/aigen_operating_point.py --sweep # reads docs/evidence/aigen_v9_scores.npz
The registry behind these figures is docs/evidence/measurements.json, and
tools/check_claims.py fails our build if any number on this site stops
resolving to a measured entry — or if a recall figure ever appears without its
false‑positive rate beside it.