Skip to content

[AI] Track on-premise burning of fossil fuels for location based carbon intensity #152

Description

@mrchrisadams

Outline Action Item Details

I'm opening this issue because:

  1. we're now seeing a trend in datacentres using increasing amounts of onsite generation, acting as the primary source of power, or regularly used, source of supplementary power rather than as backup. I think representing this in the cloud metadata dataset would be useful
  2. it looks like there will be precedents being set in a new law in Europe that will require disclosure of this info as part of a labelling scheme, meaning the data is more likely to be collated and disclosed.

I'll address each separately as I think they're both relevant.

Increasing amounts of onsite fossil generation from hyperscalers

There's a fair few examples you can link to, but the chart below from Latitude Intelligence (the research arm of Latitude Media) gives a quick graphical summary of the announced gas generation specifically tied to new datacentre projects for 2026:

Image

These figures look to be concentrated on specific campuses rather than spread across the entire fleet of datacentres for a given company. I expect this means they'd likely influence the location-based carbon intensity figures for specific cloud regions (in both directions - making them higher on clean grids, and possibly making them lower on coal-heavy grids).

Affecting current cloud regions, not just future, announced ones

While the chart above points to announced projects, we also have examples of operational projects, that would likely impact some regions in use now, which ideally would be represented here too. One concrete example might be the northeurope Azure cloud region with in Dublin in Ireland, where recently the operators of the facility applied for, and were granted a permit to install more than 160MW of reciprocating internal combustion engines running on methane gas.

I've pasted a relevant passage from the permit submitted to the Irish EPA, detailing how it's expected to be used:

The Dub 15 facility is supported by an onsite gas generation plant as the development is located in what is noted as a constrained area in terms of electrical grid capacity. The standby gas generation plant (comprising 22no. generators with 22 flue stacks (c.25m high) is planned to meet the requirements of the utilities flexible demand policy. The capacity of the plant will be 167.2MWth. It is anticipated that it will operated up to 8 hours a day, during peak demand periods, 365 days a year.

Elsewhere in the document, there are some numbers for estimating how much energy will be generated onsite from these generators - it seems to imply less than 8 hours of use per day, but it's still a meaningful figure.

Anyway, given that this is a trend we're now seeing, I think it's worth starting a conversation on how we'd represent it in the dataset - particularly if you are a consumer of these services.

Precedents for disclosure

While these are early days, I think there's actually a precedent that might mean this could be disclosed.

The driver - a new public labelling scheme for datacentres in the European Union

There's a follow on piece of legislation in the European Union after the E.E.D. led to some (somewhat patchy disclosure), that I think shows this will need to be reported.

It came up in the policy working group discussions, and there is an issue there to respond to the current consultation. I've added the early draft version of the label, plus a supporting diagram defining the different kinds of power consumption to be disclosed.

I need to stress this is early draft language and subject to change, but I think it's promising, and worth tracking

Here's the early version of the proposed label:

Image

There's a supplementary diagram related to the detail about what kind of power is disclosed.

How is different to the underwhelming experience of the earlier EED data?

One of the key things is that the the new labelling scheme seems to be much more granular and useful to this project. The design of the new label explicitly has a QR code that is intended to link to much more detailed info, and also includes new kinds of data that I think would be relevant.

Here's the specific quote from the language of the current draft about providing per-datacentre level disclosure:

Article 2 - Definition

‘quick-response (QR) code’ means a matrix barcode included on the label of the data centre that links to the location in the publicly accessible space of the European database where this label is stored.

I take this to mean that each datacentre will have a label with a URL/URI now. This goes further than cloud region level info, and would allow for meaningfully rolling up data from one or more datacentres to supplement cloud-region-level info, for specific regions offered by cloud providers.

This bit also suggests per datacentre level labels being created:

Article 3 - Generation of labels

  1. By 15 August 2027 and every year thereafter, an electronic label in the format set out in Annex II shall be automatically generated by the European database and supplied by electronic means to data centres that have communicated information and key performance indicators to the European database in accordance with Article 3 of Delegated Regulation (EU) 2024/1364.

  2. Labels for data centres shall be publicly available in electronic form in the European database in all official languages of the Union

...and this bit here suggests labels will be publicly accessible for projects like the Real Time Cloud or others:

Article 4 - Obligations of data centre operators

  1. Data centre operators shall ensure that the label issued in accordance with Article 3 is made available in electronic form to any physical or legal person requesting it. For this purpose, data centre operators can refer to a free-access website they manage or to the European database.

Where does local fossil generation come in?

We already have some EED derived columns in the dataset, but because information wasn't really disclosed at the per-DC level, nor disclosed at an aggregated, cloud-region level either, we haven't been able to use it much.

However there's a change to what needs to be disclosed now here under the part of the EED laws that set what out companies have to disclose, even if they don't publicly disclose this themselves. I'll post the bit from the relevant draft annex which goes into the detail (the bold emphasis is mine):

‘(d) Total energy consumption (‘EDC’, in kWh) of the reporting data centre shall be measured as defined by, and by using the methodology in the CEN/CENELEC EN 50600-4-2 standard or equivalent methodology. The total energy consumption includes the use of electricity, fuels and other energy sources used for cooling.

The amount of EDC coming from on-site, non-renewable sources such as generators or backup generators (EDC-BG, in kWh) shall be also reported separately. EDC-BG shall include energy consumption related to both periods of normal operation of the data centre and of maintenance of the generators or back-up generators.

Total energy consumption shall be measured at the input of the data centre system before the supply transfer switchgear. The measurement points shall be set at the primary and secondary supply of energy and at every additional supply, for example, back-up generation.

What is this EDC-BG thing? I interpret that to be the regulators trying to re-use an existing datapoint, and use it to also capture this new, secondary onsite generation we're seeing in the Azure northeurope region, or in some of the new campuses referenced in that latitude research chart.

Here's the diagram from the Annex to the draft law, that shows onsite non-renewable generation with the EDC-BG definition.

Image

The diagram is on page 6 of the annex document I've added to the issue in the GSF working group policy issue to respond to this consultation. I've added the link below

Green-Software-Foundation/policy-wg#172

What would this look like in practice?

We currently track some of the EED labels already. My suggestion, if this actually goes ahead would be to add a new column to the dataset following the current convention.

It might look like this:

RTC Cloud metadata region label unit example Notes
renewable-energy-consumption kWh 90 000 000 Total renewable energy (EED requirement)
renewable-energy-consumption-goe kWh 10 000 000 Total renewable energy from Guarantees of Origin/Renewable Energy Certificates (EED requirement)
renewable-energy-consumption-ppa kWh 75 000 000 Total renewable energy from power purchase agreements (EED requirement)
renewable-energy-consumption-onsite kWh 5 000 000 Total renewable energy from on-site generation (EED requirement)
fossil-energy-consumption-onsite kWh 10 000 000 Total non-renewable energy from on-site generation (recast EED requirement)

I know there are other regions than Europe, but this is one part of the world where there are some standards outlining what is reasonable to disclose, and I think this helps address a need end users of cloud services have - especially ones with existing climate commitments, and where hyperscalers form part of their supply chain.

Issue dependency with other WGs Groups

Policy WG

Metadata

Metadata

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions