Detector validation baseline
How we validated the Accordia detector against 3.7 million real AP records
Before onboarding our first paying customer, we ran the detector's mechanics against two large public procurement books and reported per-axis calibration alongside every primary result. This page carries the full baseline inline, with the numbers, the tables, the method, and the caveat that keeps them honest.
Own a small QuickBooks shop and would share anonymized books? Skip to the design-partner ask →
This baseline does NOT prove SMB-specific prevalence.
Municipal and national procurement books are not proxies for small-business bookkeeper behavior. What the baseline proves is that Accordia's detector measures the phenomenon correctly at scale, on real invoice-level AP, and that the multi-axis calibration approach prevents the misleading headlines a single-axis detector would produce. SMB-shaped data is what we are pursuing through aggregate partners and direct customer onboarding.
What we did
Two public books, one common schema, one detector-mechanics replica.
We downloaded two public procurement and municipal AP datasets: the City of Houston Checkbook (2016 to 2025, its full 9-year history, USD) and the French Public Procurement Awards dataset (FOPPA, 2010 to 2020, EUR). Both were normalized into a common schema of payment id, vendor, department-or-project attribution, amount, and date.
We then ran a detector-mechanics replica of Accordia's production QuickBooks Online detector against 3.7 million records, and reported per-axis calibration alongside every primary result so no single axis could quietly carry a misleading headline.
The full baseline report is embedded inline below (executive summary, cross-dataset comparison, Houston and FOPPA detail, multi-axis calibration tables). The full attribution-gap paper extends this baseline with methodology, honest limits, and the multi-axis discipline behind it: read the whitepaper.
How the numbers are measured
What "tagged" means, why one axis lies, and why we report the honest range.
The detector rule. A payment record is tagged when it carries a non-empty value in the field that says which piece of work paid for it. In QuickBooks Online, that field is the customer or project reference on the Bill line. In municipal AP, the same idea lives in one of several fields (a WBS code on a capital project, a purchase-order number on an operational buy, a contract number on a long-lived vendor agreement). A record is untagged when none of those fields carries a value. Untagged is exactly what our production detector flags in QuickBooks: a cost that landed with no route back to the work it paid for.
Per-axis calibration, in plain terms. A single field is called an axis. If you measure untagged prevalence against one axis and only one axis, you can be structurally, systematically wrong, in either direction, without the headline number ever looking suspicious.
Worked example: Houston, WBS alone versus the composite
Measured against the WBS-ID field alone, Houston looks like 92% untagged. That number is technically correct and structurally misleading. WBS captures capital-project work only. Most municipal AP is opex, attributed through purchase orders and contracts, not WBS codes. If we published 92% untagged, we would be reporting the shape of the WBS schema, not the shape of the checkbook.
Measured against the composite ANY-OF(WBS | PO | Contract), where a
record is tagged if it carries any of the three references, which mirrors the way
a QuickBooks Bill can carry customer, project, or class attribution, Houston reads
20.18% untagged. That is the honest headline. The 92%-versus-20%
gap is not a detector failure, it is what the calibration table is there to prevent.
Why we report every axis, always. The finding above is not a one-off. FOPPA has the same shape in the other direction: the buyer-id axis alone reports 0% untagged, because FOPPA's schema does not record buyerless lots at all. The awardPrice-present axis reports 30.87%, which is the honest number and matches what the FOPPA maintainers independently documented. Both cases are structural artifacts of the source data, and both would silently mask a real detector regression if the report showed only one axis. Every table on this page carries every axis so the composite primary can be read in context, never on its own.
Why it matters for the product. This is not just report hygiene. The Accordia product carries the same discipline into the surfaces a shop owner actually sees. A job margin, a cohort accuracy, a coverage number, the product surfaces these with the axes on which the measurement could have varied, one click away or in the primary view. The formal name for this rule is the Confidence-Bounded Presentation Principle, and it is the reason a customer never sees a single-number "X% untagged" without the breakdown that would have told them a different story on a different axis.
What we found
Three findings. Only one was the one we expected.
Houston: the rare untagged bills are the largest
2.3 million records, $65.1 billion in AP over 9 years. 20.18% of Bills are untagged on the composite ANY-OF(WBS, PO, Contract) axis. Those untagged Bills carry 55.66% of the dollar volume. In the largest-amount bucket (Bills over $1 million), 60.22% are untagged, dominated by financial-flow items like debt service and interbank settlements that are not attributed to a specific project. The rare untagged Bills tend to be the largest, so they distort disproportionately. That pattern is the direct SMB parallel our engine is designed to catch.
FOPPA: external validation of the detector
1.4 million records. 30.87% of lots untagged on the awardPrice-present axis. That exactly matches the FOPPA maintainers' independently documented ~31% missing-counterpart figure in their published technical report. External validation, from a separate research team using different methodology, that Accordia's detector measures the phenomenon accurately.
A methodology catch we did not expect
A naive single-axis detector produced misleading headlines in both datasets. Houston-WBS-alone would report 92% untagged (WBS captures capital-project work only, and most municipal AP is opex attributed via PO or contract). FOPPA-buyer-alone would report 0% untagged (FOPPA's schema does not record buyerless lots at all). Both are structural artifacts of the data, not detector failures. The multi-axis calibration approach is what surfaces the honest number. This is now a product-UX rule for Accordia: never show a customer a single-number "untagged %" without axis breakdown.
The numbers, in one look
Three charts. Records versus dollars, bill size, and how one axis lies.
20.18% of the Bills carry 55.66% of the dollars. The rare untagged Bill is a large one.
The Over-$1M bucket is 60.22% untagged. Debt service, interbank settlements, and other financial-flow items dominate it. That asymmetry is the direct SMB parallel: the vendor bill you never see attributed to a job is usually the largest bill of the month.
The primary axis reads 30.87%, matching the FOPPA maintainers' independently documented figure. The buyer-id sidecar reads 0.00% because the schema does not record buyerless lots. Either sidecar, published alone, would have deceived the reader. The multi-axis view is what makes the honest number readable.
The baseline report, inline
Every table, every axis, on the page.
Cross-dataset comparison
| Dataset | Records | Tagged | Untagged | Untagged $-volume |
|---|---|---|---|---|
| Houston Checkbook (USD) | 2,293,699 | 79.82% | 20.18% | 55.66% |
| FOPPA (EUR) | 1,380,965 | 69.13% | 30.87% | 0.00% |
Dollar figures are reported per dataset in their native currency (Houston USD, FOPPA EUR). They are never summed across the two: no single 2010 to 2020 EUR-to-USD rate is honest (the range across the FOPPA window was roughly 1.05 to 1.40).
Houston: multi-axis calibration
Each axis measured independently. The composite ANY-OF is the honest primary; the sidecar rows are calibration transparency.
| Axis | Untagged % | Untagged $-volume % | Note |
|---|---|---|---|
| WBS ID (single-axis) | 92.11% | 74.55% | Structural: WBS captures capital-project work only. |
| Purchase Order (single-axis) | 20.90% | 56.78% | |
| Contract Number (single-axis) | 28.62% | 59.24% | |
| Department (single-axis) | 0.00% | 0.00% | Structural: every record carries a department. |
| ANY-OF(WBS | PO | Contract): primary composite | 20.18% | 55.66% | The honest headline. Anchored to whichever route the work was attributed through. |
Houston: distribution by amount bucket
| Bucket | Records | Untagged % |
|---|---|---|
| Under $1k | 1,560,847 | 22.10% |
| $1k to $10k | 459,779 | 10.73% |
| $10k to $100k | 200,705 | 21.92% |
| $100k to $1M | 62,390 | 29.89% |
| Over $1M | 9,978 | 60.22% |
FOPPA: multi-axis calibration
The awardPrice-present axis is the honest primary and matches the independently documented figure. The buyer and supplier rows are calibration sidecars.
| Axis | Untagged % | Note |
|---|---|---|
| awardPrice-present (primary) | 30.87% | Matches the FOPPA maintainers' documented ~31%. |
| supplier_id (sidecar) | 10.08% | Different measurement; kept for calibration transparency. |
| buyer_id (sidecar) | 0.00% | Structural artifact: FOPPA does not record buyerless lots. |
Data sources
Houston Checkbook. Landing: data.houstontx.gov/dataset/checkbook. Years loaded: 2018 to 2026, all pulled live at run-time.
FOPPA v1.1.3 (French Public Procurement Awards). Landing: zenodo.org/records/10879932. Full CSV bundle, 2010 to 2020.
Reproducible on request. The replica method Accordia used to produce every Houston and FOPPA number on this page is available from Accordia on request. It additionally covers top-20 untagged-vendor tables and data-quality notes (FOPPA has known sentinel-value rows capped at EUR 1B, which never changes tagged / untagged attribution, only tames dollar-volume math).
What this proves, and what it does not
Two lists. The honest one on the right is why publishing the first list is credible.
Proves
Detector mechanics work correctly on real invoice-level AP at scale, across 3.7 million records and two independent source systems.
False-positive rate on adversarial mostly-tagged books is quantified and low. Procurement-discipline data lands where it should; the detector does not fabricate exceptions where none exist.
Multi-axis calibration is necessary to prevent structural-artifact misreads that would otherwise mislead an owner with a single-number headline.
SMB-specific untagged prevalence.
Municipal and national procurement books are not proxies for small-business bookkeeper behavior. That gap is real, and it is why we are pursuing SMB-shaped data through aggregate partners such as Codat and Rutter, and through direct customer onboarding.
The SMB parallel
The largest bills, quietly untagged. That is the pattern our engine catches in QuickBooks.
The most notable pattern in the municipal data, that the largest bills are the ones most likely to be untagged, is the direct SMB parallel our engine is built for. In an owner-operated QuickBooks Online shop's books, the vendor payments that do not get tagged to a customer or job tend to be either the outliers, a big equipment purchase or a subcontractor draw, or the recurring back-office ones, rent, utilities, insurance.
Both categories quietly inflate the reported margin on the jobs that did get tagged. The owner sees profitability they have not actually earned. Our detector reads QuickBooks AP the same way this baseline reads Houston's checkbook, and it surfaces the missing attribution before the next quote goes out.
See the same pattern applied to a single job at See the hidden cost, and the mechanism it works through at How it works.
Accordia is looking for owner-operated QuickBooks Online shops where work is priced by job, project, repair, install, order, or contract, and where the owner is willing to share anonymized books under an appropriate NDA.
In return, the partner gets free margin visibility on their own book while the partnership runs, plus a first look at the product decisions their data helps shape.
Every partner materially shortens the honest gap between this municipal baseline and the SMB findings Accordia can publish once enough customer-shaped books have been measured.
Before you write in, you may want the longer methodological version: Read the full attribution-gap whitepaper →.
For aggregate partners and researchers
If you have SMB-shaped data at scale, we would welcome a conversation.
If you have or know of anonymized SMB accounting data, particularly QuickBooks Bill-level records with customer or project attribution, we would welcome a conversation about a research or NDA framework. Accordia is in production on the Intuit App Store and preparing for its first paying customers.
For research collaboration, NDA, or media inquiries:
Contact us.
Independent replication
Independent replication and peer review are welcome.
Researchers are invited to replicate the baseline using publicly available datasets (Houston Checkbook, FOPPA). Contact for the replication method.
If you are a researcher or practitioner who has read the baseline and would like to share a review, please contact us.
Reminder, in plain terms. This baseline proves the detector's mechanics on public procurement data. It does not prove SMB-specific prevalence. The proves and does-not-prove section above is the honest bound of every number on this page.