Guides
How We Evaluate Budgeting Apps: Evidence and Limits
Our current source-based evaluation standard, the records required for a real test, and a correction to earlier claims about hands-on testing.
Evidence standard — Current Vault Daily comparisons are based on cited public documentation unless a page provides verified records of an actual test. A plausible scenario, generated dataset or numerical score is not such a record. We are not currently presenting a verified four-week, real-money testing programme.
What a source-based comparison checks
- The exact product, plan, region and billing period being discussed.
- The primary page supporting each material claim, with an access date.
- Whether permissions, costs and limitations are described accurately.
- Whether examples are hypothetical and calculations can be reproduced.
- What remains unknown or could not be independently verified.
What would be required to call something a test
A real test needs a recorded protocol, actual dates, product versions, appropriate consent, traceable observations and a way to reproduce the reported calculation. Identifying details must be protected. We should not claim purchases, account connections, interviews or device trials without records supporting those events.
A test score would also need a documented rubric, transparent weights and underlying observations. We have withdrawn unsupported legacy scores rather than using precise-looking numbers as a substitute for evidence.
The legacy CSV is synthetic
The categorization example explains what the published sample actually contains. It is illustrative synthetic data, not evidence of a measured competition between named apps. It does not validate old rankings.
How AI is used
The publication system uses AI to research, draft, review and publish material through automated checks. A second AI review is not a human fact-check. Source records and deterministic checks help detect mistakes but cannot guarantee truth. See the editorial policy, corrections log and contact page.