Solutions · Spend data foundation
How to build a procurement spend cube without implementing a new S2P platform
A procurement spend cube can be built directly from a raw accounts-payable or ERP export. The work is supplier normalization and line-level classification to a taxonomy, and neither needs a source-to-pay platform. ValueChaser Procurement Categorization takes the raw export and returns a categorized dataset at L2 to L4, with the method and confidence recorded for each row.
Analysis is generated in minutes, with completed client diagnostics available within 24 hours. A Procurement Categorization run completes within 10 minutes.
What a spend cube is
A spend cube organizes spend on three dimensions: what was bought (category), who it was bought from (supplier) and who bought it (cost center, location or business unit), usually over time. It is the dataset that savings analysis, supplier consolidation and category strategy are built on.
What you need to start
- A raw spend or AP extract in Excel, CSV or TSV, with one row per transaction line.
- Three signals on each line: a supplier name, an amount, and something that indicates what was bought, such as a category field, a GL code or a description.
- Dates are optional. 12 to 36 months is the recommended window.
No template, required column names or data cleaning are needed. Multi-currency files are normalized to one reporting currency. Aggregates such as pivot tables, dashboards and PDFs of tables cannot be categorized line by line.
The method
- Field detection and sufficiency check. Column headers are detected automatically and each mapping is recorded with a confidence score. A sufficiency check confirms the three signals are populated before anything is processed or charged.
- Supplier normalization. Raw vendor name variants are consolidated so the same supplier is treated the same way across the whole file. The report shows how many raw vendor entries resolved to how many unique suppliers.
- Taxonomy selection. Three modes are available: standard UNSPSC, the ValueChaser industry taxonomy, or your own taxonomy, validated at intake. Classification runs to four levels.
- Classification in a fixed order. Each transaction runs through four methods: a direct match on the file's own category field, a keyword and rule match, a known supplier match, and AI classification for the residual that deterministic methods cannot resolve. AI output is constrained to valid taxonomy nodes.
- Audit trail. Each row records the method that classified it and the confidence. The categorization rate is reported by rows and by spend, and uncategorized spend is listed with the suppliers, amounts and the fields that would resolve it.
What the output looks like
| Output | What it contains |
|---|---|
| Categorized dataset (CSV) | Each transaction with its L1 to L4 category, spend type, normalized supplier, classification method and confidence, next to the original source fields |
| Excel workbook | Category, supplier, data quality and field mapping tabs |
| Categorization report | Field mapping with confidence, method distribution, a category-level audit trail and the uncategorized spend list |
The categorized dataset is the spend cube. It can be pivoted in Excel, imported into an ERP or procurement system, or passed to Procurement Insights for procurement savings analysis.
Quantified examples
European power utility (anonymized). 200,000 spend transactions were classified to UNSPSC family level across 16 categories, with 179 suppliers normalized and a 100% categorization rate.
Food and beverage producer (anonymized). 5,000 transactions covering €488M of spend were categorized to the company's own uploaded taxonomy across 34 L2 categories, with supplier entries normalized from 98 to 93.
“What you produced in about 10 to 15 minutes would normally take me roughly two weeks of work.”
Experienced procurement advisor · live client spend dataset. A recent example. His review of the output took about three hours.
When this is the wrong tool
This produces a spend cube from an export at a point in time. It is not a live data pipeline connected to your ERP, and it does not refresh on a schedule. If you need continuous spend reporting, a spend analytics platform is the better fit. See AI procurement diagnostics vs. traditional spend analytics platforms.
Frequently asked questions
Can AI build a reliable spend cube directly from an AP export?
Yes, when deterministic methods do most of the work and a model handles only the residual. Reliability comes from a fixed taxonomy, a recorded method and confidence on each row, and an audit trail that lets you check any classification.
Does it classify to UNSPSC or to my own taxonomy?
Both. You can choose standard UNSPSC, the ValueChaser industry taxonomy or your own taxonomy, which is validated at intake and then applied.
How many lines can it handle?
The examples on this site include files of 200,000 transactions. Deterministic rules apply the same way to each row, and supplier normalization runs across the whole file.
What happens to inconsistent supplier names?
Raw vendor variants are consolidated into normalized supplier names, and the report shows the before and after counts.
Do I need to clean the data first?
No. Inconsistent formatting, missing fields and multi-currency files are handled. If the data cannot support the analysis, the sufficiency check stops the run and lists the gaps before any charge.
Related
- Spend categorization to UNSPSC or your own taxonomy with Procurement Categorization
- How to identify procurement savings opportunities from AP spend data
- How raw spend gets categorized, a detailed walkthrough
- ValueChaser for consulting teams and PE portfolio diagnostics
