How we audit a 3PL invoice

Most audit tools ask you to trust a number. We think you should be able to check ours — and so should your 3PL. This page explains, in plain language, exactly how Auditario decides what to call a billing error, how AI is (and isn't) used, and why our numbers are deliberately conservative.

What the versioned engineering corpus proves

A method is only as good as the evidence behind it. The repository therefore separates immutable public source documents from invoice fixtures that reproduce documented provider layouts with synthetic values. We do not market exploratory research counts as release evidence.

7hashed public sources agreements and rate cards committed with URL, date, byte count and SHA-256
4documented invoice layouts ShipBob, Extensiv CSV/XLSX and ShipHero shapes; fixture values are synthetic
1immutable manifest tests run without fetching or silently replacing the source bytes
0claim-ready cells no check/provider/layout combination is certified for publication today

What reading them changed

The point of the corpus is not its size. It is to turn observed billing and document shapes into permanent false-positive regressions before any class can be certified. Four examples:

  • Percentage fees hide in the quantity column. A late fee of 5% was billed as quantity 0.05 against a rate of 205,403.51. Quantity × rate ties exactly, so every arithmetic check passed and the row read as our strongest possible signal. Priced against a storage rule it would have produced a confident demand for $13,669.02 of charges that were genuinely owed. A rule may now only price a line whose quantity is a plausible count of that rule's unit.
  • “Recoverable” was counting money we could not date. When no dispute-window clause was on file, those dollars fell through into the claimable column — which is every merchant who has not yet sent a contract. Recoverable now requires positive evidence that the window is open, and unknowns are stated as unknown.
  • Split shipments are not duplicates. One order legitimately billed across two shipments looks identical to a double-bill until you check the carrier reference.
  • A scanned invoice can carry a text layer that lies. Detecting one is not the same as trusting it, and we no longer do.

The seven source files and their hashes are listed in the repository manifest. Invoice fixtures use documented layouts but synthetic merchant values; they are not represented as production evidence or provider acceptance. Any future private benchmark needs a separately auditable, privacy-preserving provenance record before its aggregate results are published.

The target three-source evidence model

The target complete audit compares three evidence sources. Current founding audits do not yet include the third:

  1. What you agreed to pay — your rate card, MSA, amendments, and any written concessions from your provider, in strict order of precedence (a signed amendment beats the original rate card; a document always beats an AI's interpretation).
  2. What you were billed — supported invoice lines normalized into standard charge categories and traced to their available page, row, or cell citation; unsupported evidence is reported as a coverage gap.
  3. What actually happened (planned) — Shopify orders, fulfillments, cancellations, and returns. This connector is not built yet, so current founding audits do not claim this coverage.

What “claim-ready” means — and what it doesn't

Auditario keeps evidence, dispatch, and automation as separate states. A detector can be mathematically confident while the merchant-facing result still requires review:

ClassMeaning
Claim-readyEvery required proof is present, independently recomputed, source-bound, inside an established dispute window, and certified for that exact check, provider, and parser layout. Only this state can enter claimable totals.
Review requiredThe engine observed a possible variance, but at least one proof is unresolved. We show the arithmetic and blocker, and claim $0.
OptimizationA contractually valid charge you could reduce by changing process, packaging, or terms. Real money, but not an error.
In your 3PL's favorUnderbilling. Yes, we report it. An audit that only ever finds errors in one direction isn't an audit.
Data qualityThe sources conflict or are incomplete, so no reliable conclusion exists. We say so instead of guessing.
The rule we run the company on: a missed issue is bad; a confidently invented issue is worse. Missing, contradictory, stale, or uncertified evidence produces review or silence — never a claim.

How AI is constrained — and what is not live yet

AI can help read messy documents, but it cannot be the authority for money. The boundary is explicit:

Why “flagged” and “recovered” are different numbers

Some audit vendors advertise everything they flag as if it were money in your pocket. Auditario's model keeps the full journey separate: observation → review required → claim-ready → merchant approved → disputed → acknowledged by your provider → credited on an actual invoice. Claim-ready never implies merchant approval or AutoPilot authority. Planned pricing is flat rather than recovery-based, and the recovery ledger counts only credits that verifiably landed.

Dispute windows: the clock your contract started

Most 3PL agreements limit how long you have to dispute an invoice — often 30 to 90 days. When a governing clause and anchor are established, the engine records whether the window is open or expired. Missing or unsupported terms remain window unknown and can never enter claim-ready money.

What happens when we find nothing

The planned reconciliation statement reports how many lines were evaluated, the rate-card version used, covered spend, and unresolved evidence. It is not a guarantee, attestation, or statement that all billing is correct.

Your data