Guides
Budgeting Categorization: Synthetic Example, Not a Test
A correction to the former eight-app benchmark and a reproducible explanation of the 120-row synthetic categorization example.
Research correction — The former claim that Vault Daily tested 500 real transactions across eight named apps is withdrawn. The available public file contains 120 synthetic example rows and two anonymous prediction columns. It cannot establish any named app's accuracy or ranking.
What can actually be calculated
In the labelled synthetic CSV, app_a_correct totals 111 and app_b_correct totals 104 across 120 rows. The illustrative proportions are therefore 111 ÷ 120 = 92.5% and 104 ÷ 120 ≈ 86.67%. “App A” and “App B” have no verified mapping to real products.
| Anonymous example | Correct flags | Rows | Illustrative proportion |
|---|---|---|---|
| A | 111 | 120 | 92.5% |
| B | 104 | 120 | 86.67% |
What the file does not prove
It does not prove that a product processed these transactions, that the merchant descriptors came from real customers, or that any app is better than another. No score for Noruvo, Quenzio, Monarch, YNAB or another product can be derived from an anonymous synthetic example.
Reproduce the example responsibly
Count the data rows, sum each correct flag and divide by the row count. If you use the file to demonstrate a calculation, retain its synthetic label and do not present the result as a field experiment. A real comparison would require the evidence described in our evaluation standard. Browse the hub for current source-based guides.