← All posts

What is spend analysis? Making spend data usable

10 September 2026 / BoreaTech Team

Scattered faint data points on a dark background converging into a small set of solid stacked bars, representing fragmented purchasing records aggregated into categories

Spend analysis is the process of collecting purchasing data from every system that holds it, cleaning it, classifying it against a category taxonomy, and aggregating the result so an organisation can see what it actually buys, from whom, and at what price.

Described that way it sounds like a reporting exercise. It is not. Three of those four steps are data work, and the fourth is only ever as good as the three beneath it. Programmes that disappoint almost never fail at the analysis. They fail because the data arriving at it could not support the question being asked.

Why it sits underneath everything else

Sourcing, contract negotiation, supplier consolidation and category strategy all begin from the same two numbers: how much is spent on a category, and with whom. An organisation that cannot produce them cannot run a credible sourcing event. There is no baseline to negotiate against, no volume to commit, and no way to tell afterwards whether the negotiated saving materialised.

The same data decides less visible things: which suppliers are worth the effort of a full onboarding cycle, which categories justify catalog content and which can stay on free text, where a contract is being ignored. Spend analysis is the measurement layer the other disciplines are steered by, which is why its weaknesses propagate quietly into decisions that look unrelated to it.

Collection: the data is in more places than anyone expects

The first step is gathering transactions, and the difficulty is that they are not in one place and have never been designed to be combined.

A typical estate holds purchase orders and invoices in one or more ERPs, often several after acquisitions, each with its own vendor master and chart of accounts. Card transactions arrive as a bank feed with a merchant name, a total, and nothing else. Expense claims sit in a separate system with a different approval trail. Travel and logistics are often booked in tools that never touch the ERP until the invoice lands. Accounts payable holds everything eventually, but as posted documents rather than purchases.

Then there are the decisions that quietly change the answer, each of which has to be made once and applied identically on every refresh, or two runs of the same period disagree and nobody trusts either:

  • Which value. Gross or net of recoverable tax, before or after credit notes, including or excluding freight and surcharges.
  • Which date. Purchase order date, invoice date, or general ledger posting date. These can differ by months, and a category that looks seasonal may just be a posting artefact.
  • Which exchange rate. Transaction-date rate, monthly average, or a single closing rate for the period. On a multi-currency base the choice moves totals enough to change a conclusion.
  • What counts as spend. Payroll, taxes, intercompany transfers and statutory charges are money leaving the organisation but are not addressable by procurement, and leaving them in makes every percentage meaningless.

Cleansing: one supplier, several vendor records

The second step is making the supplier dimension true. In the source systems it usually is not.

The same supplier appears as several vendor records: a name typed slightly differently on each creation, a legal entity renamed after a restructuring, a local subsidiary set up as a separate vendor because it invoices from a different country, and an acquired business still trading under its old name in one system and its new name in another. A large vendor master routinely represents a materially smaller number of real trading relationships.

The goal is not deduplication but a supplier hierarchy: rolling transactions up to the entity an organisation would actually sit across a table from. Name matching alone is fragile, because the strings that look closest are often different companies and the ones that look distant are often the same. VAT and company registration numbers are stronger keys where populated, bank details help, and addresses are useful until a supplier moves. Each of those fields has holes, so cleansing ends up as a scored combination of signals with a person ruling on the borderline cases.

The hierarchy then has to be maintained. One that was correct at the time of the project decays with every acquisition, new vendor record and entity rename.

Classification: the step where projects stall

The third step is assigning every transaction to a node in a category taxonomy, whether that is UNSPSC, eCl@ss, or a bespoke internal scheme. This is the hardest part of spend analysis, and it is where most programmes lose momentum.

The obvious shortcut, using general ledger account codes as categories, does not work. GL accounts exist for financial reporting, not category management. An account called consumables spans a dozen separately sourceable categories, professional services mixes legal, audit, recruitment and consulting, and in most charts of accounts the largest single line is some variant of “other”.

So classification has to be done against the transaction data itself, and the level at which it is done determines what the analysis can ever show.

Supplier level versus line-item level

There are two ways to classify, and the choice is the most consequential design decision in the exercise.

Supplier-level classification assigns a category to a vendor, and every transaction with that vendor inherits it. It is fast, needs only header data, and is accurate often enough to be tempting. It is wrong the moment a supplier sells across categories, which the largest suppliers invariably do. A national distributor sells stationery, cleaning chemicals, personal protective equipment and small tools on one account, and all of it lands in whichever bucket that vendor was tagged with. The error is not random, either: it concentrates in the biggest suppliers, which is exactly where the analysis needs to be right.

Line-item classification assigns a category to each line of each document. It is accurate, it survives suppliers that sell widely, and it is the only way to answer item-level questions. Price variance for the same product across sites cannot be seen at supplier level, because at supplier level the product does not exist as an entity.

Line-level analysis requires line-level data

This is the part that gets skipped, and it is not a tooling problem.

Classifying at line level requires line-level detail in the source: an item identifier, a description worth reading, a quantity, a unit of measure and a unit price. That detail exists only when the transaction was structured at the moment it was created, and in practice that means it came from a catalog. A requisition raised against a catalog line carries the supplier part number, the unit of measure and the contracted price, and every document downstream of it inherits those fields. See what is a procurement catalog for what that structure consists of, and what is PunchOut for the case where the structured cart is returned from the supplier’s own site rather than held locally.

A free-text requisition carries a sentence somebody typed and a total. No analytics tool recovers what was never captured. An organisation that buys predominantly through free text has a permanent ceiling of supplier-level classification, however good its spend analytics platform is. Buying a better tool does not move that ceiling. Increasing catalog coverage does, which is why catalog management and spend analysis are the same programme viewed from two ends.

The questions worth asking

Once the data holds up, four analyses repay the effort of getting there.

Supplier consolidation. How many suppliers serve a category, how the spend is distributed across them, and how much sits in a long tail of low-value relationships each carrying a fixed administrative cost. The finding is rarely that consolidation is impossible, it is that the tail is longer than anyone believed.

Price variance for the same item. The same part number, bought in the same month by two sites at different unit prices. This is the most directly actionable output spend analysis produces, and it requires line-level data, which is why many organisations never see it.

Contract coverage. What share of category spend runs through a negotiated agreement and what share went elsewhere. Off-contract buying is not usually defiance, it is a catalog that could not be searched successfully or a supplier that was easier to phone.

Payment terms. The distribution of terms across the supplier base, and whether the terms held in the ERP match the signed contract. They frequently do not, and the gap is both a working capital question and a compliance one.

How spend analysis fails in practice

The one-off project. A team or an external firm assembles a clean dataset, presents findings everyone agrees with, and leaves. Nobody built the refresh. Six months on the numbers are stale, refreshing them is a second project at full cost, and the organisation concludes that spend analysis is expensive. The deliverable should be a repeatable pipeline, with the report as a by-product.

Classification done once. Mappings are completed during the project and never maintained. New suppliers and new items arrive unclassified, land in an “unclassified” bucket, and within a year that bucket is the largest category in the organisation. At that point the analysis is not wrong, it is just silent about the fastest-growing part of the spend.

A taxonomy that ages. The category tree was designed around the business as it was. Spend patterns move, new types of purchase appear, and categories that were immaterial become significant while still buried inside some general node, while a third of the tree holds no spend at all. A taxonomy is a model of the business and needs a scheduled review like any other model.

Numbers that do not tie to the ledger. If the spend dataset does not reconcile to accounts payable totals, the first question in every meeting is why the two disagree, and the meeting never reaches the findings. Agree the reconciliation, and the reasons for any remaining difference, before presenting anything.

Findings without an owner. An opportunity with no named category owner, no mandate and no follow-up date is an observation, not a finding. That is a governance failure rather than an analytical one, and it is the most common reason good analysis produces nothing.

Spend analysis is funded as an analytics problem and constrained as a data-structure one. The binding limit is the shape of the transaction at the moment it was created, and that is decided upstream, in whether the buyer selected a catalog line or typed a description. SupplierForge works on that upstream end: ingesting supplier content through whatever channel a supplier can manage, then validating, classifying and normalising it centrally, so the lines reaching the P2P systems already carry identifiers, units and categories. Analysis at line level is only ever possible on transactions that were structured before anyone thought about reporting.

procurement-strategyprocurement-fundamentalsdigital-procurement