Benchmark · TRACE (working name)

TRACE scorecard — run-1

Run window 2026-10-06 · 5 products · scorecard.json · results · what we are fixing · plan · re-run and challenge

Phase 0: operator self-assessment. Questions, answer keys and scoring are run by Asymmetric Intelligence, which operates the systems scored. Results are not yet independently validated or maintained. They measure the fleet's progress against the TRACE method, not independent quality. Challenges are open now through our contact page; the checks become re-runnable by anyone when the repository opens.

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Not yet in place:

No composite score and no rank. Every figure is a separate measure, shown as found.

What TRACE measures, and why

Floor

Your own deep-research report

The strongest version a reader could run themselves: a frontier AI model with deep research switched on, challenged by a second model, and asked once to check its own answer for errors and omissions. This is the floor. It is set deliberately high, never a straw man.

Measured

Our products

What a reader of Advennt, Crypto, Data Protection, Financial Integrity or World Payments actually sees: the published jurisdiction records, the reports and the Co-Pilot answers. This run checks the published records.

Ceiling · measured towards, never claimed

A major international law firm's research

An answer key written by senior practitioners for each question. This is the ceiling. We measure how far we have closed the distance towards it, dimension by dimension. We never claim to have reached it.

TRACE is the public, re-runnable part of the method we use internally to steer our own work. That method is called ARRF. It places each product between the two reference points above and asks, for each part of an answer (accuracy, currency, authority, application to the facts and so on), how much of the distance from the floor to the ceiling we have closed. There is no single overall score: each dimension is reported on its own.

Before we may say anything about a product, it has to earn a step on a ladder. The first step (R1) is "every claim traceable to a source we read, with its status computed as at a stated date". Later steps are about being more rigorous and repeatable than AI research (R2), applying the law to a reader's facts (R3) and structuring answers to the standard of specialist advice (R4). A product never states a step it has not earned. We never claim to be a law firm, to give legal advice, or to make anyone fully compliant.

How confident each answer should be is the job of our risk engine, RRIE. Calibration (whether a stated confidence matches how often the answer is right) cannot be measured yet, because our published records do not yet carry a stated confidence for each claim. We say so rather than estimate it.

Half of the internal question bank is held back and never published, so that our own monthly re-scoring stays honest. Only the other half will ever appear here.

Where each product stands

Track B checks each product's published records: 170 Advennt, 170 Crypto, 170 Data Protection, 170 Financial Integrity, 170 World Payments Monitor. Products are listed alphabetically. This is not a ranking. B1 and B2 count every failing row the fleet provenance gate finds, split into fatal rows (the gate fails them) and warnings. FATAL on B1 or B2 means new fatal rows exist against that product's debt baseline. Failing rows are listed in the repository's per-product failing files.

Advennt advennt.io

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.

Track B, dataset audit: 170 records, 2272 citations.

CheckFailing / checkedFatalWarningResult
B1 Source retrieved
Was each cited source actually fetched and read?
510 / 2272 145 365 FATAL: new fatal rows against the debt baseline
B2 Provenance match
Is the page we read the same publisher as the page we cite?
21 / 2272 21 0 No new fatal rows against the debt baseline
B3 Tier validity
Is every source we grade as official (T1) on the official-domain list for that jurisdiction?
73 / 740 — — FATAL
B4 Date coherence
Do the checked, published and changed dates on each record agree?
pending first run
B5 Gap honesty
Where we have no data, do we say so explicitly instead of leaving a blank?
pending first run
B6 Integrity
Is any published record empty or hollow?
0 / 170 — — Pass
B7 Sampled truth (operator-keyed)not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b.

Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.

Crypto cryptoassets.gi

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.

Track B, dataset audit: 170 records, 2498 citations.

CheckFailing / checkedFatalWarningResult
B1 Source retrieved
Was each cited source actually fetched and read?
867 / 2498 27 840 FATAL: new fatal rows against the debt baseline
B2 Provenance match
Is the page we read the same publisher as the page we cite?
6 / 1610 6 0 FATAL: new fatal rows against the debt baseline
B3 Tier validity
Is every source we grade as official (T1) on the official-domain list for that jurisdiction?
21 / 249 — — FATAL
B4 Date coherence
Do the checked, published and changed dates on each record agree?
pending first run
B5 Gap honesty
Where we have no data, do we say so explicitly instead of leaving a blank?
pending first run
B6 Integrity
Is any published record empty or hollow?
0 / 170 — — Pass
B7 Sampled truth (operator-keyed)not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b.

Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.

Data Protection dataprotection.gi

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.

Track B, dataset audit: 170 records, 4004 citations.

CheckFailing / checkedFatalWarningResult
B1 Source retrieved
Was each cited source actually fetched and read?
784 / 4004 146 638 FATAL: new fatal rows against the debt baseline
B2 Provenance match
Is the page we read the same publisher as the page we cite?
5 / 2980 5 0 FATAL: new fatal rows against the debt baseline
B3 Tier validity
Is every source we grade as official (T1) on the official-domain list for that jurisdiction?
17 / 300 — — FATAL
B4 Date coherence
Do the checked, published and changed dates on each record agree?
pending first run
B5 Gap honesty
Where we have no data, do we say so explicitly instead of leaving a blank?
pending first run
B6 Integrity
Is any published record empty or hollow?
0 / 170 — — Pass
B7 Sampled truth (operator-keyed)not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b.

Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.

Financial Integrity sentinel.gi

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.

Track B, dataset audit: 170 records, 21913 citations.

CheckFailing / checkedFatalWarningResult
B1 Source retrieved
Was each cited source actually fetched and read?
5636 / 21913 874 4762 FATAL: new fatal rows against the debt baseline
B2 Provenance match
Is the page we read the same publisher as the page we cite?
42 / 14428 42 0 FATAL: new fatal rows against the debt baseline
B3 Tier validity
Is every source we grade as official (T1) on the official-domain list for that jurisdiction?
100 / 3641 — — FATAL
B4 Date coherence
Do the checked, published and changed dates on each record agree?
pending first run
B5 Gap honesty
Where we have no data, do we say so explicitly instead of leaving a blank?
pending first run
B6 Integrity
Is any published record empty or hollow?
0 / 170 — — Pass
B7 Sampled truth (operator-keyed)not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b.

Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.

World Payments Monitor payments.gi

phase-0operator-runoperator-keyedunblindedoperator-affiliated

Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.

Track B, dataset audit: 170 records, 30487 citations.

CheckFailing / checkedFatalWarningResult
B1 Source retrieved
Was each cited source actually fetched and read?
6105 / 30487 1266 4839 FATAL: new fatal rows against the debt baseline
B2 Provenance match
Is the page we read the same publisher as the page we cite?
178 / 21645 178 0 FATAL: new fatal rows against the debt baseline
B3 Tier validity
Is every source we grade as official (T1) on the official-domain list for that jurisdiction?
279 / 7561 — — FATAL
B4 Date coherence
Do the checked, published and changed dates on each record agree?
pending first run
B5 Gap honesty
Where we have no data, do we say so explicitly instead of leaving a blank?
pending first run
B6 Integrity
Is any published record empty or hollow?
0 / 170 — — Pass
B7 Sampled truth (operator-keyed)not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b.

Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.

What we are fixing because of this run

Each finding below is an internal fix with its own reference, re-measured in the next run.

Action plan

Every month:

  1. Before each run, we record in advance what will be measured, which models judge it and which questions are used.
  2. We run the checks on every product's published records, and the answer tests once the question set exists.
  3. We publish the results here as found. Each run is kept; past runs are never overwritten. Bad results get the same space as good ones.
  4. Every failure is tied to the internal fix that owns it.
  5. The next run measures it again. A fix is not done until its measure moves.

What comes next, in order:

StepWhat it unlocks
Choose the benchmark's final nameTRACE is a working name until this is decided.
Move the repository to a neutral organisation and open it to the publicAnyone can then re-run the checks themselves and file challenges directly.
Fix what run-1 exposed (the items above)Run-2 shows whether each fix moved its measure.
Make date coherence and gap honesty measurableB4 and B5 move from pending to measured.
Publish the development question set, with primary-text quotes for every answerAnswer testing (Track A) can start. The held-back half stays unpublished.
Run the answer tests: our products against the deep-research floor, repeated runs, changed-fact tests and planted errorsTrack A results, and the first distance-closed figures per dimension.
Carry a stated confidence on each published claim, checked against an expert-reviewed sampleCalibration moves from not measurable to measured.
Run monthly from run-2, each run registered in advanceA history for every measure.
Open alpha: outside domain reviewers, keyholders from three organisations, a sealed question set written by senior practitionersResults stop being a self-assessment. Until then this page makes no comparative claims.
Add the AI Monitor (AIC) when it joins the shared research pipelineA sixth product row.

Before results can stop being a self-assessment (the open-alpha gates):

Open-alpha gateNowNeeded
External domain reviewers03
Keyholders from outside organisations03
Outside systems scored on a sealed set03
Judge–human agreement κnot yet measured≥ 0.7

Re-run and challenge

The repository is not yet public: it stays private until the work is more advanced. The commit and re-run command below are shown anyway, so that this run is pinned to exact code and inputs. They become usable by anyone when the repository opens.

Challenges are open now through our contact page: email [email protected] naming the product, the figure and why you think it is wrong, or write to Asymmetric Intelligence Limited at the registered office (Unit G02, Eurocity, Europort Avenue, Gibraltar GX11 1AA). Challenges and our rulings will be published with the next run. Open challenges: 0.

Runrun-1 (2026-10-06 to 2026-10-06) · previous: run-0
Repo commit83fcd7c (+ branch tr2/bind; consumer snapshot commits in runs/run-1/NOTES.md)
Item-set sha256n/a: no Track A items in run-1
Preregistrationnone: run-1 was not pre-registered
Judgesnone in this run
Harnesstrackb-adapter (tr2/bind, --gate check_source_provenance.py blob fb17d37a)
Re-runsee runs/run-1/NOTES.md section 'Reproduce' (adapter.py per consumer, then runs/run-1/build_scorecard.py)
Page rendered fromtrace-benchmark@a3f751cb347b, template a-i-gi.html.j2

Changelog

Data: CC BY 4.0. Information, not legal advice. Machine-readable: runs/run-1/scorecard.json.