Sheet 3 · Edition 2026-08-30 02:52 UTC
Open data file formats, counted across 31 government catalogues
Every entry in a government catalogue advertises a format for each downloadable resource. Read across these catalogues, that is 3,828,453 resources naming their format in 3,547 different strings. One string per resource, and the counts sum to the resource total exactly. This sheet is the whole of that column, sorted. If you are about to filter one of these catalogues by format, it is what you would be filtering against. One further catalogue is registered and has not been opened, and nothing from it is in these figures.
- 3,828,453
- Resources counted
- 3,547
- Distinct format strings
- 300
- Strings published in full
- 31
- Catalogues read
The 20 most common values, verbatim
Exactly as the portals wrote them, lower-cased and trimmed and nothing else. No two of these rows have been merged, which is why the same format appears more than once in the list.
- html562,391 · 14.7%
- csv517,427 · 13.5%
- pdf266,761 · 7.0%
- xls232,438 · 6.1%
- xlsx212,942 · 5.6%
- zip179,866 · 4.7%
- json131,056 · 3.4%
- http://publications.europa.eu/resource/authority/file-type/csv129,342 · 3.4%
- unknown122,438 · 3.2%
- xml116,718 · 3.0%
- http://publications.europa.eu/resource/authority/file-type/netcdf108,072 · 2.8%
- application/x-netcdf107,261 · 2.8%
- wms80,991 · 2.1%
- http://publications.europa.eu/resource/authority/file-type/html70,900 · 1.9%
- png59,144 · 1.5%
- wfs49,792 · 1.3%
- tiff31,854 · 0.8%
- http://publications.europa.eu/resource/authority/file-type/xml31,382 · 0.8%
- http://publications.europa.eu/resource/authority/file-type/pdf30,965 · 0.8%
- geojson30,435 · 0.8%
Below the published 300, a further 3,247 distinct strings cover 15,712 resources, 0.4% of the corpus. They are counted but not listed: portals declare formats as free text. The export records only how many strings and how many resources, so the most that can be said about their spread is the average, 4.84 resources per string.
HTML leads on single values, CSV once spellings are added up
The single most declared value in the corpus is html, on 562,391 resources. A further 7 strings say the same thing in different words, and together the 8 of them account for 635,794 resources, 16.6% of everything catalogued. The two rankings disagree, and the disagreement is the finding. Added up across their spellings, CSV totals 650,304 resources and HTML totals 635,794, so which format is the most common depends entirely on whether the spellings of one format are counted as one format. Neither figure is adjusted, and both are in the table below. A resource declared as HTML is a link to a page. It may be a page with a table on it, or a landing page with a download button, or a page that no longer exists. What it is not is a file, and no catalogue field records which of those it is.
- html562,391
- http://publications.europa.eu/resource/authority/file-type/html70,900
- text/html; charset=utf-8931
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/html617
- html_simpl518
- file:///srv/udata/ftype/html179
- shtml159
- htm99
One format, several spellings
Grouping these is a judgement, so it is shown as one. Each row prints the exact strings it is the sum of, and the total beside it is the sum of those strings and of nothing else. No row claims to be every occurrence of its format: only the published 300 strings were searched, so the 3,247-string tail is not in any of these figures.
15 formats, 3,139,344 resources, 82.0% of the corpus, spread across 89 different spellings. A compressed or profiled variant is never folded in: csv.gz is an archive and kmz is a zipped KML, so neither counts as its parent format here.
CSV
650,304 · 17.0%Plain text, one record per line, values separated by a delimiter. 7 spellings.
- csv517,427
- http://publications.europa.eu/resource/authority/file-type/csv129,342
- сsv2,381
- https://www.iana.org/assignments/media-types/text/csv637
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/csv220
- file:///srv/udata/ftype/csv152
- cvs145
HTML
635,794 · 16.6%A web page. Not a data file: whatever is on the page still has to be read off it. 8 spellings.
- html562,391
- http://publications.europa.eu/resource/authority/file-type/html70,900
- text/html; charset=utf-8931
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/html617
- html_simpl518
- file:///srv/udata/ftype/html179
- shtml159
- htm99
A fixed page layout. Readable by a person, machine-readable only with work. 3 spellings.
- pdf266,761
- http://publications.europa.eu/resource/authority/file-type/pdf30,965
- рdf282
XLS
260,058 · 6.8%The pre-2007 Microsoft Excel binary workbook. 5 spellings.
- xls232,438
- http://publications.europa.eu/resource/authority/file-type/xls25,683
- хls1,621
- vnd.ms-excel190
- r.xls126
XLSX
232,777 · 6.1%The Office Open XML workbook: a zip container of XML parts. 9 spellings.
- xlsx212,942
- http://publications.europa.eu/resource/authority/file-type/xlsx9,201
- хlsx4,779
- excel (.xlsx)3,764
- vnd.openxmlformats-officedocument.spreadsheetml.sheet1,097
- https://www.iana.org/assignments/media-types/application/vnd.openxmlformats-officedocument.spreadsheetml.sheet560
- xlxs212
- xslx114
- ехсеl (.xlsx)108
NetCDF
218,370 · 5.7%Array-oriented scientific data, the usual container for climate and earth observation. 4 spellings.
- http://publications.europa.eu/resource/authority/file-type/netcdf108,072
- application/x-netcdf107,261
- nc2,193
- netcdf844
ZIP
192,791 · 5.0%A container. What is inside it is not declared anywhere in the catalogue. 5 spellings.
- zip179,866
- http://publications.europa.eu/resource/authority/file-type/zip12,331
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/zip329
- file:///srv/udata/ftype/zip162
- https://www.iana.org/assignments/media-types/application/zip103
JSON
158,344 · 4.1%Nested key and value text. 4 spellings.
- json131,056
- http://publications.europa.eu/resource/authority/file-type/json26,265
- https://www.iana.org/assignments/media-types/application/json881
- file:///srv/udata/ftype/json142
XML
148,444 · 3.9%Tagged text. The schema it follows is a separate question the catalogue does not answer. 4 spellings.
- xml116,718
- http://publications.europa.eu/resource/authority/file-type/xml31,382
- хмl226
- https://www.iana.org/assignments/media-types/text/xml118
WMS
108,368 · 2.8%OGC Web Map Service: a live endpoint returning rendered map images, not a file. 5 spellings.
- wms80,991
- http://publications.europa.eu/resource/authority/file-type/wms_srvc24,434
- ogc:wms1,842
- ogc wms888
- wms_srvc213
WFS
65,194 · 1.7%OGC Web Feature Service: a live endpoint returning map features, not a file. 5 spellings.
- wfs49,792
- http://publications.europa.eu/resource/authority/file-type/wfs_srvc13,717
- ogc:wfs1,171
- ogc wfs339
- wfs_srvc175
Esri Shapefile
63,468 · 1.7%A vector dataset held in several files that must travel together, so it is usually zipped. 14 spellings.
- esri shapefile (shp)24,939
- shp22,452
- http://publications.europa.eu/resource/authority/file-type/shp7,488
- shape-zip4,301
- shape1,104
- esri shapefile706
- shp.zip695
- shapefile380
- esri shape362
- shapefiles337
- zip (shp)217
- shapefile (zip)179
- shp / zip179
- esri shapefile - zipped129
TIFF
41,115 · 1.1%A raster image. GeoTIFF is the same file carrying coordinate tags. 6 spellings.
- tiff31,854
- tif4,878
- geotiff1,941
- http://publications.europa.eu/resource/authority/file-type/tiff1,219
- geotif884
- http://publications.europa.eu/resource/authority/file-type/geotiff339
GeoJSON
39,567 · 1.0%Geographic features expressed as JSON. 5 spellings.
- geojson30,435
- http://publications.europa.eu/resource/authority/file-type/geojson8,636
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/geojson210
- file:///srv/udata/ftype/geojson152
- application/vnd.geo+json134
KML
26,742 · 0.7%Geographic features expressed as XML, the Google Earth lineage. 5 spellings.
- kml22,566
- http://publications.europa.eu/resource/authority/file-type/kml3,736
- https://data-lra-cha.opendata.arcgis.com/api/feed/dcat-ap/ftype/kml209
- file:///srv/udata/ftype/kml128
- https://www.iana.org/assignments/media-types/application/vnd.google-earth.kml+xml103
Four things that break a format filter
These overlap the groups above and each other on purpose: one string can be both a spelling of CSV and a URI. Nothing here is added to anything there.
1. The value is a URI, not a format name
67 of the published 300 strings are URIs, covering 578,305 resources, 15.1% of the corpus. Most of them, 570,059 of the 578,305, point at the EU Publications Office file-type authority, which is the controlled vocabulary DCAT-AP points implementers at, so those are a portal following the standard rather than a portal getting it wrong. The rest are not that: one is the IANA media-type registry, one is a persistent identifier for a vocabulary term, two are individual portals’ own hostnames, and one is a path on a server’s own disk that reached publication.
- http://publications.europa.eu40 strings570,059 · 14.9%
- https://www.iana.org14 strings4,704 · 0.1%
- https://data-lra-cha.opendata.arcgis.com5 strings1,585 · <0.1%
- file:///srv/udata/ftype6 strings915 · <0.1%
- https://w3id.org1 string766 · <0.1%
- https://data.overheid.nl1 string276 · <0.1%
2. The format name is spelled with letters that are not the letters
8 published strings, covering 10,159 resources, contain a Cyrillic letter where a Latin one is expected. They are not Cyrillic words: they are Latin format names in which a letter was typed on a Cyrillic layout. On screen they are identical to the Latin spelling. To a comparison they are a different string, so a filter written as format == 'xlsx' matches none of them. A prefix test is not a way around it, and is not even consistent: 7 of the 8 carry the substituted letter in first position, where a prefix test sees it immediately, and the other one carries it further along, where a prefix test passes and only the equality test fails.
- х U+0445 reads as x
- с U+0441 reads as c
- е U+0435 reads as e
- р U+0440 reads as p
- м U+043C reads as m
- хlsx4,779
- сsv2,381
- хls1,621
- ехel507
- рdf282
- xlsх .zip255
- хмl226
- ехсеl (.xlsx)108
3. The value names no format at all
24 published strings, covering 191,566 resources, 5.0% of the corpus, say nothing about the file. At least three different things are in that figure. unknown is what this harvester writes when a resource declared no format whatsoever, so it is our label for an absence rather than a portal’s word. inconnu, autre and ostalo are portals saying “unknown” and “other” in their own language. Others describe how the thing is delivered (service, api and download), where it lives (url and web page) or how its bytes are encoded (utf8), instead of what it is. Those are examples and not the whole list: this paragraph names 10 of the 24 strings, and the other 14 are in the run below.
- unknown122,438
- online29,485
- other13,804
- service7,099
- api3,949
- url3,515
- web page2,635
- data1,700
- download1,628
- inconnu1,121
- view975
- spatial data format674
- page web593
- webpage564
- net320
- ostalo160
- http142
- map viewer126
- https115
- spatial viewer106
- utf8106
- website link106
- si/103
- autre102
4. One field, several formats
8 published strings, covering 4,220 resources, put a list where a single value belongs. Every one of them fails an equality test against each of the formats it contains.
- application/zip, application/octet-stream, application/x-zip-compressed, multipart/x-zip1,978
- shp, tab, fgdb, kmz, gpkg726
- xls(x)521
- doc(x), xls(x), zip316
- shp tab jpeg2000 - gda2020 - mga2020212
- csv/zip210
- csv zip130
- xlsx, csv127
How much of this has been checked against the file
A declared format is a claim the portal makes in its catalogue. Everything above counts those claims; it does not verify them. This harvester has fetched 6,540 of 3,828,453 resources, 0.2% of the corpus, and the download phase is what the rest is waiting on. Of the ones fetched, 74 returned a web page in place of the file they advertised and 4 came back empty. Those are counts of what happened, not a rate: the fetch order is not a random sample of the corpus, so nothing here can be scaled up to it. The per-format byte figures for what was mirrored are on the harvest status sheet.
Where the per-catalogue figures are
This sheet is the corpus total. Every catalogue also has its own declared-format breakdown, on its register entry, and the same figures are in the JSON behind this page: formats.json for the corpus and portals/<portal_id>.json for one catalogue. Both are linked, with their terms, from the export. How the harvester reads a catalogue in the first place is on the method sheet.