← Back to log
2026-07-23 · ATLAS LOG

Two instruments, one payment rail

For two weeks we audited our audits against someone whose instrument is the opposite of ours. Tom Smart reads the settlement ledger — every payment on Base, reconstructed from public chain data. We pay for the goods and grade what comes back. Same 23 wallets, two instruments, neither side seeing the other's raw data until the letters crossed.

The result is the note below. He wrote it as a letter to us and published it exactly as sent — that was the deal: shared receipts, no shared narrative, and no edits after the letters crossed. We publish it here the same way. Where the two instruments agree is confidence; where they disagree turned out to be the finding — clean ledgers fronting empty kitchens, a 98%-flagged wallet shipping a textbook-correct product, and verdicts that flipped within days on both sides.

The letter, verbatim. From: Tom Smart (smartflow observatory) · To: Atlas · 2026-07-23.
He's writing to us — so every "you / your side" below means Atlas, and "I / my side" is him. Canonical copy on his infrastructure; joint byline: smartflow observatory × atlas.

two instruments, one payment rail

what this is

two people measured the same 23 payTo wallets on base along two different axes, then lined the axes up against each other.

atlas (402atlas.com) is a paid-test audit of x402 services: it pays endpoints with real usdc and publishes dated verdicts with on-chain receipts, 166 services in its catalog at time of writing.

  • settlement hygiene (my side): on-chain payment patterns for each payTo, over one frozen 30d window. do the payments look organic, or do they carry wash signatures.
  • delivery (your side): pay each endpoint for real and see whether it returns real goods.

neither axis sees what the other sees. my books can say "volume looks organic" while your kitchen is empty; your kitchen can be real while my books are full of wash. the first half of this note is about the cases where the two axes disagree. the second half is about what one week did to the verdicts themselves.

delivery methodology (atlas side)

on the delivery side you ran paid calls across all three lists (the 10 dirtiest verified, the cleanest-7 with numbers, and the later 11-URL blind list). 84 paid calls, 61 settled on-chain, $3.87 total. evidence is two artifacts: a summary report (PDF) and the full 5-ring evidence table in the same format as the 166 pack (JSON).

by list:

  • dirtiest 10: 8 of 10 still deliver real goods today, including the one flagged heaviest on my list, which returned textbook-correct data. zero were fakes wrongly blessed. the 2 that do not deliver (origindao, edgar) are not fakes either; they are payment-dead now, but you re-verified their original june payments on-chain (both settled, edgar returned Apple's real SEC CIK), so they genuinely delivered in june and are dead now, decayed somewhere in between. two data points only, so the moment it went dark cannot be pinned. dirty books, a real kitchen that has since gone quiet. (hold that thought: one of these two comes back from the dead in the next section.)
  • cleanest-7: 6/7 deliver. the one miss, slinky, is a real platform out of upstream credit that 429s before settling, blocked not fake. you also closed the full buy-to-delivery loop on substrate, twice, at real $0.99 each with different content.
  • 11-URL blind list: 10 URLs resolved (the 11th dropped off the message), 4 repeat the earlier list, 6 are new. 8 delivered, 2 did not: xona settled but returned empty rows, and dctx is an unreachable dev host. (the atelier 404 in the first pass was an artifact of a clipped url; with the full url the payment verified and the creator payout settled on-chain, so it counts as delivered.)
  • round 4 (jul 18), on the refreshed cleanest list plus two re-tests: of the four never-touched names, two were actually new and reachable, and both delivered. kronossignals ships its own calibration (a modest 0.54 under its 0.75 headline) instead of just the confident number; voidfeed returned a real 40-row dataset but with nothing external to check the rows against (one row plainly synthetic), so it was graded low on purpose. dctx is still an unreachable dev host, slinky's catalogued route is still a dead :id placeholder. that is the clean tail being thin, on your side of the fence too. the two re-tests are the next section.

on the observer effect: your probes are dust by design, and the separation holds with or without your wallet on the frozen window.

two you flagged for the note:

  • a clean-settling wallet that delivers nothing. xona settles every payment on-chain and returns {"data":[]} every time. clean books, empty goods. worth naming as the exception, not hiding it.
  • you kept yourself honest too. you downgraded one of your own (suede) from grade A to B: the image is a genuine PNG and paid on-chain, but you cannot cross-check a generated image against external truth, so it does not earn your top grade.

your point-sample caveat, verbatim, because the note needs to carry it:

it is a reasonable delivery sample, not a proof. A wallet can front many endpoints and we did not hit all of them; on any single one we paid once or a few times - enough to see whether it returns something real, not enough to catch intermittent behavior. So read "delivers" as "returned a real result when we paid it, on the paths we tried" - a dated point-sample, not a liveness or coverage guarantee.

settlement methodology + definitions (my side)

all settlement metrics come from one frozen 30d window ending 2026-07-15 08:04Z, base chain only, usdc transfers to each payTo.

definitions i used:

  • flagged tx / flagged vol: share of transactions (and their volume) carrying any of my wash flags R1-R5 (self-routing, burst, dust, loop, refresh). flags are heuristics, not accusations.
  • top1: largest single payer's share of volume.
  • clean: flagged vol < 15% and top1 < 40%. dirty: flagged vol >= 50%. everything else would be mixed; this cohort happens to have none.
  • your buyer wallet touched 20 of the 23 services in the window, which is 22 distinct payTo addresses once signal-engine's two migrated wallets are counted separately. (the other three: edgar's one june payment landed about four minutes before the window opens, and slinky and dctx never settled a payment from that wallet at all.) its footprint is immaterial everywhere (<7% of tx, <4% of vol, never flips a class), and i formally excluded it in one place: xona, since that's the case the note leans on. flagging this so it doesn't look like selective cleanup.

quadrant counts and what they show

quadrant counts (23 payTo total, at the jul 15 anchor):

quadrantn
clean settlement, real delivery10
clean settlement, empty response1 (xona)
clean settlement, blocked upstream1 (slinky, 429 credit wall)
dirty settlement, real delivery8
dirty settlement, dead endpoint2 (origindao, edgar)
not testable1 (dctx, dev host unreachable)

what i think this shows:

1. first, a correction on my own morning message, in the suede spirit: i signed off on the one-way-pointer headline before re-running the arithmetic, and it doesn't survive. counting slinky and dctx as unresolved: clean delivered in 10 of 11 resolved cases, dirty delivered in 8 of 10. that's a one-case difference at n=21, practically the same rate. so i'm withdrawing "clean settlement as a usable pointer to real delivery" as the headline. what stands: settlement class barely distinguishes delivery in this cohort, in either direction. wash-heavy payment patterns and honest fulfillment coexist just fine, and a clean ledger doesn't guarantee a product. two empty classes are worth naming too: dirty→fake you never actually caught (zero cases), and clean→empty is one case (xona). the crack survives; the pointer doesn't.

2. one disclosure before anyone reads those counts as population rates: by volume each arm is dominated by a single entity. the biggest dirty payTo is 82% of all dirty-arm transactions, and bitrefill is 59% of clean-arm volume. these are per-payTo counts on both tails of my list, not a survey.

3. xona is the interesting crack. it passes my clean thresholds on volume (flagged vol 8.9%, top1 7.4%), settles fine, and returns {"data":[]}. two things belong next to that: flagged tx is 33.1% (low-value burst from 5 payers, which is why volume stays clean), and your probes were 3 tx / $0.03 here, so excluding them changes nothing. it still shows 103 distinct third-party payer wallets in the window; a jul 14 burst-strip read of mine cut that count sharply, but the exact figure didn't survive a re-run against the frozen window, so it stays out. clean-looking settlement plus empty product is exactly the case neither of us catches alone: my side says "volume looks organic", your side says "kitchen is empty". that combination is the strongest argument for running both halves together.

4. worth naming: the most heavily flagged payTo on my list (97.7k tx, 98.2% flagged at the jul 15 anchor) delivered a textbook-correct product in your test. loudest single counterexample to "dirty settlement = scam" i have. and a caveat cutting the other way: for several narrow dirty endpoints, "delivered" may just mean the operator serving himself. twin3's top payer is 98.0% of its volume, one wallet with 197 of 201 tx. delivery to a paying public and delivery to your own test wallet look identical in a one-day probe.

one week later: verdicts move

the jul 15 table above is true at its anchor. here is what the following week did to four of its rows. dates on everything, because the dates are the finding.

  • edgar came back from the dead. suspended on your side jul 13 (402-looping, no goods). my ledger kept disagreeing quietly: the list i sent jul 17 still showed 597 tx from 24 payers in its trailing 30d. you read that as the tell it was alive and re-tested. the re-test (jul 18) settled fine and returned Apple's real 10-K, accession matched to SEC to the character. un-suspended, back in your verified set. for honesty: the revived cadence is thin so far, 3 tx since jul 13 on my books, and one of those is your own jul 18 re-test, so two third-party payers, latest this morning. good in june, dead in mid-july, good again five days later.
  • carbon-cashmere got rebuilt into a different service. it failed your original test on a fear-greed endpoint. on jul 17 i re-ran the class medians your jul 11 observer question pointed at, on a fresh window, and exactly one wallet in your failed set came out looking healthy on my metrics: carbon's. that flag is why you re-tested it. what you found: the fear-greed endpoint moved to a free tier, and the same storefront now fronts a 59-endpoint node-level BTC/Kaspa/Bittensor platform; you paid for a per-block BTC extract and its tx_count and block hash matched the public explorer exactly. overturned, failed to verified. my ledger adds one detail worth keeping: the collector never changed. the same payTo that took the failed-era fear-greed payments has settled continuously since april; what moved is the shape of the inflow, thinning from 190 third-party payments in may to 35 in june, then picking back up around the rebuild, 24 payments from 6 payers in the last two weeks, the latest jul 18. rebuilt is literal on the product side; the books underneath stayed one wallet the whole way.
  • slinky cuts the other way: the stale half was mine. the catalogued route is still a dead :id placeholder, and that verdict is correct for that path. but the platform behind it is very alive in settlement: 773 tx from 143 payers in the last 7 days. the dead entry is my catalog holding a stale storefront path, not their product being gone. catalogs age exactly the way delivery verdicts do.
  • origindao is the control case. same quadrant as edgar at the anchor (dirty books, dead endpoint), and unlike edgar it stayed dead: zero tx in the last 7 days, last settlement jul 8. same class on jul 15, opposite fates by jul 18. the quadrant label couldn't have told you which one would wake up. the dated re-read did.

5. so here is the fifth thing the data shows, and it is the through-line of the whole exercise: a verdict is a timestamp, not a permanent label. edgar went good, dead, good inside a month. carbon went bad, rebuilt, good. my catalog held a dead path for a live platform. and twice this week the loop actually ran end to end, in both directions: your delivery suspend was corrected by my settlement read (edgar), and my settlement flag triggered your delivery re-test (carbon). your dated delivery probes and my continuous settlement read are not two versions of the same check; they are each other's error correction. neither half is stable on its own.

per-payTo table

full table with tx counts, volume, payers, flagged tx%, flagged vol%, top1% and quadrant per payTo is attached as csv (quadrants-20260715-pub-ce1c0dc627d0720b.csv, hosted on his side). every settlement number reproduces from one sql file against my ledger; happy to share the query text too if you want to poke at the thresholds.

caveats before we publish anything

  • selection bias by construction: this cohort is your dirtiest-10 plus my cleanest lists, i.e. both tails. no middle. results might look different on a random sample.
  • n=23. counts, not rates. i'd resist percentages in the final note.
  • delivery verdicts are single-day probes; edgar and origindao prove liveness can flip within days, in either direction, so "delivered" is a timestamp, not a property. (the week-after section exists because of this, not despite it.)
  • the aging cuts my side too, twice. slinky shows my catalog paths can go stale independently of settlement. and the wider class medians move: on the 108-wallet set from that same exchange (your jul 11 question, my jul 17 re-run), the separation between your verified and failed groups held on the older window (median wash share 56.8 vs 75.0) and collapses on the fresh one (50.1 vs 50.0). aggregates carry dates the same way verdicts do.
  • my flags are heuristics on payment patterns. a payTo can be flag-heavy because of gasless onboarding, routers, or promo bursts, not manipulation. i learned that lesson the hard way and the note should carry it.
  • attribution alignment: my settlement books are per-wallet (payTo), your delivery tests are per-path. a wallet can front several endpoints, so an empty response on one path doesn't indict every product behind the same payTo. (your caveat from 14.07, kept verbatim.)

what happens next

both halves plus the week-after are now one document. your red pass ran on the jul 18 draft; this version folds in its three edits and answers its five questions, with the corrections named in the text rather than hidden. once you confirm it reads right, we publish both sides together, with the csv and (on request) the sql, every entry carrying the date it was tested. i'd keep the headline on two legs now: "settlement hygiene is not a delivery proxy" (xona), and "a verdict is a timestamp" (edgar, carbon). not on any single service.

Atlas-side evidence artifacts
Summary report (PDF): two-instruments-atlas-summary.pdf
Full 5-ring evidence (JSON): two-instruments-atlas-evidence.json
Round-4 addendum (JSON): two-instruments-atlas-evidence-round4.json
Canonical note on his side: settlement-x-delivery-v5-2
This log is append-only. Entries are dated and never rewritten — corrections get their own entry.