Carbon & GHG Accounting

The Biggest Problem in Carbon Accounting Isn't Carbon. It's Data.

The difficult part of carbon accounting is not the equation. It is knowing whether you can trust what goes into it.

2026-08-26|Carbon & GHG Accounting

Corporate data wasn't designed for carbon accounting

The basic mathematics of carbon accounting is surprisingly simple: activity data × emission factor = emissions. In practice, the difficult part is rarely the calculation itself. It is everything that comes before it.

From our experience working with large corporate carbon datasets, far more time is spent understanding what a transaction actually represents than calculating the emissions associated with it. Where did the data come from? What was actually purchased? Who supplied it? How should it be classified? Is it within the reporting boundary? Has it already been captured somewhere else? Is the information sufficiently granular to apply an appropriate emission factor? And, critically, could someone else trace the calculation back to its source and understand how the final number was produced?

These are fundamentally data questions. As carbon reporting becomes more sophisticated, organisations need to start treating carbon accounting less as an isolated sustainability calculation and more as an enterprise data problem.

Most of the information required to calculate a corporate carbon footprint already exists somewhere within an organisation. The problem is that very little of it was collected for that purpose. Procurement systems are designed to purchase goods and services, finance systems to record expenditure, expense systems to reimburse employees, and supplier databases to manage vendors. None of them were primarily designed to answer the question: what greenhouse gas emissions resulted from this activity?

That becomes very apparent when you start working with the data at scale. A procurement transaction might contain a supplier name, value, cost centre, category and short description, but that does not necessarily tell you what was actually purchased. Descriptions may be vague, supplier names may differ between systems or countries, categories may be missing or inconsistently applied, and the same supplier may provide completely different products and services.

This creates a strange situation where the calculation can be technically correct while the answer is still wrong. Multiplying a value by the wrong emission factor perfectly simply produces a very accurate calculation based on a poor assumption.

The real work therefore starts with understanding the data well enough to decide what that value actually represents.

A supplier is not necessarily an activity

Supplier-level classification is a good example. Suppose an organisation spends £10 million with a particular supplier. A straightforward spend-based approach would classify that supplier into an industry, select the corresponding emission factor and calculate the associated emissions. At a high level, that may be a perfectly reasonable approach.

But what happens if that supplier provides engineering consultancy, construction materials and specialist equipment? Applying one industry classification to the entire £10 million creates a level of consistency that does not really exist. The supplier may be one legal entity, but the underlying economic activities are very different.

The obvious solution is to classify everything at a more granular level. In reality, this quickly runs into another problem: scale. Large organisations can generate hundreds of thousands or millions of procurement records, and attempting to investigate every individual transaction is rarely practical or proportionate.

One of the most important lessons from working with large Scope 3 datasets is that the question should not simply be 'How granular can we make the data?' It should be 'Where would additional granularity materially improve the result?'

Materiality before perfection

In one multi-year analysis, our experience covered more than a million procurement records across different suppliers, expenditure categories, business units and reporting periods. The important data-quality problems were not evenly distributed across those records. A relatively small number of suppliers, categories and classification decisions could have a disproportionate effect on the calculated footprint.

That changes how you approach data quality. If a small supplier represents a negligible proportion of total emissions, spending hours determining the perfect classification may make almost no difference to the final inventory. If another supplier represents a significant proportion of the footprint, a broad or incorrect classification can materially distort the result across multiple reporting years.

This is where conventional data analytics becomes extremely useful in carbon accounting. Rather than starting with individual records, start by profiling the dataset. Understand completeness and coverage, identify duplication and anomalies, measure classification quality, and rank suppliers and categories according to their calculated contribution to the footprint. A Pareto-style analysis can quickly show where a relatively small number of decisions are driving a large proportion of the result.

The next stage is then targeted investigation. Where would a better supplier classification materially change emissions? Which categories rely heavily on generic assumptions? Where is poor-quality data being applied to large amounts of spend? Which suppliers warrant transaction-level analysis rather than a broad supplier-level classification?

The objective should not be to create a perfect dataset. In most large organisations, that is probably unrealistic. The objective is to create a dataset that is sufficiently reliable in the places that matter.

Sometimes the information you need isn't in the dataset

Even after cleaning and analysing the available data, there will be situations where the information simply isn't there. A finance system might tell you that £250,000 was paid to a contractor, for example, without telling you exactly what the £250,000 purchased. The information needed to improve the carbon calculation may instead sit inside a purchase order, contract or invoice.

Historically, extracting that information at scale was difficult. Someone might have needed to manually open hundreds or thousands of documents, interpret descriptions, extract line items and classify them. The potential improvement in the carbon calculation often wouldn't justify the time and cost required to obtain the information.

AI and automation are starting to change that equation. OCR, large language models and other document-processing tools can increasingly turn information contained within invoices and other unstructured documents into usable datasets. Our work has tested this approach on complex invoice data, extracting individual line items, descriptions and values before structuring and classifying them for emissions calculations. Instead of treating a large invoice as one homogeneous activity, it becomes possible to understand what was actually purchased and apply more appropriate assumptions at line-item level.

AI can add significant value to carbon accounting, but it is important not to confuse better data extraction with better carbon accounting. AI may help an organisation access information that was previously difficult to use, but it does not determine whether something belongs within the reporting boundary, whether the classification is conceptually correct, whether the emission factor is appropriate or whether the final methodology is defensible.

Automating a weak methodology simply allows you to produce weak results faster.

The data also has to be consistent over time

Another complication is that carbon data does not remain static. Supplier classifications improve, new data sources become available, emission factors are updated, reporting boundaries evolve and errors are discovered. This is a positive part of developing a more mature carbon inventory, but it creates an important issue when organisations compare performance over time.

Our multi-year rebaselining work has shown that improvements to supplier classifications, category mappings, emission-factor selection and historical methodology can materially change the calculated footprint even though the underlying activity has not changed. The goods were still purchased, the services were still delivered and the emissions still occurred. What changed was the understanding of them.

This is why materiality thresholds and a clear base-year recalculation policy matter. Not every correction should trigger a rebuild of years of historical data. Carbon datasets will continually improve, and constantly rewriting the baseline for immaterial changes would be both impractical and potentially confusing. But where an improvement is significant enough to cross an organisation's established materiality threshold, historical recalculation may be necessary to maintain a meaningful comparison.

Otherwise, a reported movement in emissions can contain two different things: real-world changes in emissions and changes in how those emissions were measured. If those aren't separated, an organisation can appear to have decarbonised when the reduction is actually methodological, or appear to have gone backwards simply because its data improved.

This is a subject in its own right, but it demonstrates why carbon accounting cannot be separated from data governance. The methodology, data and historical record all need to work together.

Traceability matters more than false precision

Carbon inventories often produce remarkably precise numbers. An organisation might report Scope 3 emissions of 482,316 tonnes of CO₂e, giving the impression that the footprint is known with considerable accuracy. Underneath that number, however, may sit supplier classifications, spend-based emission factors, inflation adjustments, currency conversions, estimates, proxies and other assumptions.

There is nothing inherently wrong with estimation. For many Scope 3 categories, it is unavoidable. The problem is confusing numerical precision with data certainty.

A more useful question is whether the organisation understands where uncertainty exists and whether the reported number can be traced back through the methodology to the underlying evidence. For every material component of the footprint, someone should be able to explain where the source data came from, how it was transformed, why a particular classification was selected, which emission factor was applied and what assumptions were made along the way.

That data lineage becomes particularly important as carbon reporting moves towards greater scrutiny and external assurance. If a reported number cannot be reconstructed or explained, adding another decimal place does very little to make it more credible.

Carbon accounting is becoming an information system

Our experience with corporate emissions data shows that carbon accounting is an information-management problem rather than simply an annual sustainability calculation.

Reliable carbon reporting still requires a strong understanding of greenhouse gas accounting, but it increasingly also requires skills and systems from data engineering, procurement, finance, governance and analytics. Automation and AI will become part of that architecture as well, particularly where organisations need to process volumes of information that would previously have required extensive manual work.

The organisations that do this well will not necessarily be those with the most sophisticated carbon calculator. They will be the ones that understand where their information comes from, where its weaknesses are, which assumptions materially affect the result and where investing in better data will actually improve decision-making.

The carbon equation was never really the difficult part. The difficult part is knowing whether you can trust what you put into it.

If you would like to discuss this topic in the context of your organisation, get in touch.

Need specialist sustainability support?

Tell us what you're trying to solve. We will assess the problem, identify where specialist support would add value and propose a practical way forward.