chopmark — independent TEE scorecard

Six inference providers, graded on what they can actually prove about the machines serving your requests. A grade measures coverage of a proof surface: a vendor-rooted quote answering a live challenge, the endpoint bound into that quote, the boot measured, the release pinned, the GPU attested. Only independently verified checks earn points. Provider claims set expectations and never score.

Evidence is checked against vendor roots of trust — TDX quotes against Intel PCS, SEV-SNP against AMD KDS, GPU evidence against NVIDIA NRAS, supply-chain bundles against Sigstore. Every response is signed, and every grade is reproducible: chopmark verify <provider> runs the same checks this page does.

Questions, corrections, or a provider you want covered: disputes@chopmark.dev

providers

providergradecoveragelevelcheckslast verifiedidentity anchors
chutesA · 85si✓ bd✓ lv✓ mb✓ gpu✓ sc✓ fl✓L544✓ 0✗2026-10-08 12:12 UTCinstance_count=4
nearaiB · 82si⚠ bd✓ lv✓ mb✓ gpu✓ sc⚠ fl✓L5 / clean L027✓ 0✗2026-10-08 12:13 UTCmeasured_os-image-hash=da9a3d5cc196a1a7… compose_hash=8bc802c9c8043a9c…
phalaB · 80si✓ bd✓ lv✓ mb✓ gpu· sc⚠ fl·L4 / clean L310✓ 0✗2026-10-08 12:13 UTCkeyset_digest=ef8a03c0c5e34931… measured_os-image-hash=bd369a8c2f9edb2b…
redpillA · 85si✓ bd✓ lv✓ mb✓ gpu✓ sc✓ fl✓L568✓ 0✗2026-10-08 12:13 UTCinstance_count=4
tinfoilA · 85si✓ bd✓ lv✓ mb✓ gpu· sc✓ fl·L46✓ 0✗2026-10-08 12:13 UTCrelease_tag=v0.0.155
veniceF · 73si✓ bd✓ lv✓ mb✗ gpu✓ sc⚠ fl✓L214✓ 2✗2026-10-08 12:13 UTCmeasured_os-image-hash=a6eafc5f007f642d… compose_hash=55db164f4f8c6a83…

coverage: si silicon_root · bd endpoint_binding · lv liveness · mb measured_boot · gpu · sc supply_chain · fl fleet — ✓ proven ⚠ warned ✗ expected but absent · not applicable. Points are earned only by independently verified checks; claims never score.

level: L1 challenge · L2 bound · L3 measured · L4 reproducible · L5 full. The first number is how far the proof reaches; a second number is how far it reaches with no warnings at all. A caveated root of trust leaves no clean rung, so "L5 / clean L0" means full scope and nothing unqualified.

incidents

whenproviderseveritykinddetail
10-08 12:13redpillminorcoverage_changepoints 84 → 85
10-08 12:13redpillinfofact_changeinstance_count changed: 3 → 4
10-08 12:13redpillinfofact_changeinstance_set changed: sha256:ec9d3746e23f9f1d36777d8ded602d6e → sha256:ca2eb8f7bb594befb28a482f7ca74e4b
10-08 12:13redpillinfograde_transitiongrade B → A ()
10-08 12:12chutesinfofact_changeinstance_count changed: 3 → 4
10-08 12:12chutesminorcoverage_changepoints 84 → 85
10-08 12:12chutesinfograde_transitiongrade B → A ()
10-08 12:12chutesinfofact_changeinstance_set changed: sha256:36765361c923b986c8cc5b4528c17433 → sha256:90645e68b3f286c164f696f0b743d74d
10-08 11:13nearaiinfofact_changemeasurement_config changed: 8xh200 [10.2.1, numa-236c-1128g-nvsw-node1] v1.4.1 → 8xh200 [10.2.1, numa-124c-1128g-nvsw-node1] v1.4.1
10-08 11:13nearaiinfofact_changeinstance_id changed: b5413fa3-ee8c-4e55-9c9c-f480455f7c65 → 3202c6b3-f6de-4d32-a6e7-3da1e18dd9ec
10-08 10:13redpillminorgrade_transitiongrade A → B ()
10-08 10:13redpillinfofact_changeinstance_count changed: 5 → 3
10-08 10:13redpillinfofact_changeinstance_set changed: sha256:dd358201706f10d12b4b8c7ad9aafe0b → sha256:ec9d3746e23f9f1d36777d8ded602d6e
10-08 10:13redpillminorcoverage_changepoints 86 → 84
10-08 10:13nearaiinfofact_changemeasurement_config changed: 8xh200 [10.2.1, numa-124c-1128g-nvsw-node1] v1.4.0 → 8xh200 [10.2.1, numa-236c-1128g-nvsw-node1] v1.4.1
10-08 10:13nearaiinfofact_changeinstance_id changed: 2b971a36-a2b2-4034-bf7a-64df579eaaa1 → b5413fa3-ee8c-4e55-9c9c-f480455f7c65
10-08 10:12chutesinfofact_changeinstance_count changed: 5 → 3
10-08 10:12chutesinfofact_changeinstance_set changed: sha256:c001456494cdc6c8b753e8ccdff301af → sha256:36765361c923b986c8cc5b4528c17433
10-08 10:12chutesminorgrade_transitiongrade A → B ()
10-08 10:12chutesminorcoverage_changepoints 86 → 84

default history

providergradeworstdowngradesincidentsunresolvedreported loss
chutesAF3233233—
nearaiBF311041104—
phalaBB01616—
redpillAF6629629—
tinfoilAB11919—
veniceFF310931093—

A rating is only as good as its record of failures. An unreachable sweep never counts as a downgrade or a worst grade. GET /defaults serves this signed.

quality labels

task setmodelresultjudgeendpoint integrity
judged@1.0.0e2ee-deepseek-v4-flash6 pass · 0 failmodel:deepseek-ai/DeepSeek-V3.2-TEE@1 (A 87 L5/clean L5)B 73 L5/clean L0
instruction@1.0.0e2ee-deepseek-v4-flash7 pass · 0 failexact-match@1B 73 L5/clean L0
formatting@1.0.0e2ee-deepseek-v4-flash6 pass · 0 failjson-has@1B 73 L5/clean L0
classification@1.0.0e2ee-deepseek-v4-flash6 pass · 1 failone-of@1B 73 L5/clean L0
extraction@1.0.0e2ee-deepseek-v4-flash7 pass · 0 failexact-match@1B 73 L5/clean L0
reasoning@1.0.0e2ee-deepseek-v4-flash9 pass · 0 failexact-match@1B 73 L5/clean L0

Each set is signed and carries a verifiable reference to the scorecard's verdict on the endpoint that produced the outputs, so one artifact answers what was served, whether it was any good, and what machine served it. A judge id beginning model: is a model verdict carrying the judge's own attested standing. GET /labels streams the archive.

verify these claims

GET /scores · GET /score?provider=X · GET /provider?provider=X · GET /grade?provider=X&at=T · GET /allow?provider=X&min_grade=B · GET /allows?providers=X,Y · GET /spec · GET /defaults · GET /incidents · GET /evidence · GET /labels · GET /pubkey — all responses signed ed25519/JCS.

signer pubkey: d0488b99f10d265b78ad40cc22fed544dce45c4668a4468b947692dee0f31314

generated 2026-10-08T13:11:15Z