What EXIF Data Do Photos Still Carry?
This study read 2984 JPEG files drawn at random from Wikimedia Commons on 10 Sep 2026.
- Still in the file: a camera make in 57.2%, a serial number in 13.4%, and a location coordinate in 10.8%.
- Who writes what: phone-made files carry a location in 52.7%; every other device carries a serial in 28.5%.
- Half are hidden: opening the maker’s own notes finds a serial in 26.7%, against 13.3% for the standard fields alone.
- Why it matters: a serial number follows one camera across every photograph it takes.
What this study measured
The question was narrow. Of the JPEG photographs published on Wikimedia Commons, what share carries a camera serial number, what share carries a location coordinate, and how do those two shares differ between camera makes and between camera models? The sample is Commons itself, and it was drawn by the platform’s own random generator rather than chosen by us: 2984 files, the ones a visitor would land on by opening the library at random.
EXIF is the block of notes a camera writes into a photograph next to the picture. It is where a camera leaves its name, its settings and sometimes its serial number and the place the picture was taken, and it is present in 65.5% of the files measured here.
The XMP packet holds a similar note in a different format, and the maker’s own note is a private area each manufacturer fills in as it likes. This study reads all of them and keeps them apart.
- The sample: 2984 files, drawn at random from the Commons file library on 10 Sep 2026, kept only when the file is a JPEG.
- The reading: the first 131072 bytes of every file, plus every byte of every tenth file, so the cheap reading could be checked against the complete one.
- The exclusions: 16 files were left out because the fetch or the reading failed. The rule that removes them was fixed before the first file was fetched, and the count travels beside every share on this page.
A file counts as carrying a location only when its coordinate values are not all zero. That rule was fixed in advance, because a camera that never got a satellite fix often writes a coordinate block full of zeros, and a cleaning tool that removes the value while leaving the block produces the same shape. Of the files with a coordinate block at all, 1.5% hold nothing but zeros, and those are not counted as carrying a place.
The prediction, and what the measurement did
The pre-registration was written down before the first file was fetched, and it made claims that could fail. It predicted that a location coordinate would appear in somewhere between 8% to 25% of the files. The measurement came back at 10.8%, with an interval of 9.7% to 11.9%, inside that range. That prediction held.
The second prediction was the interesting one, and it was written as something a script could check: the serial number is a property of the maker, not of the file. In the words of the pre-registration, more than half of the serial-bearing files would come from a handful of makes, and at least one make with 20 or more files would sit at or below 1%. Both held. 90.9% of the files carrying a serial come from the makes that carry the most of them, and 4 of the makes with 20 or more files sit at or below 1%.
28.2 points
is the gap in serial numbers between files that come from a phone maker and files that come from every other kind of device. A serial number is not something a photograph either has or lacks; it is something a particular kind of camera writes.
The pre-registration also named what would prove it wrong, in these words: no make differs from the overall serial rate by more than five percentage points. The largest gap measured is 33.4 points, so the falsifier did not fire. Any statement about camera serial numbers in photographs therefore has to name the camera, because the average describes almost nobody.
Who writes the location, and who writes the serial
Split the sample by what made the file and the two fields swap places. Files from a phone maker carry a location coordinate in 52.7% of cases, and files from every other kind of device carry one in 10.4%, a gap of 42.3 points with an interval of 36.6 to 48.1 points. The serial number runs the other way: it is in 28.5% of the files from cameras and scanners, and in 1 of the phone-made files, which is a share of 0.3%.
| What made the file | Files | Location | Camera serial |
|---|---|---|---|
| A phone maker | 315 | 52.7% | 0.3% |
| Every other kind of device | 1389 | 10.4% | 28.5% |
The pattern is not that phones write less. A camera make is present in 100% of the phone-made files and in 100% of the rest, and an EXIF block in 100% and 100% of them, so the two groups carry the same container and differ only in what is inside it.
Most of the location on the library is not in the file at all. A coordinate added to the file page by a person or a program is present in 23.4% of the drawn files, and one written by the camera in 10.8%, which is 12.7 points more, with an interval of 11.2 to 14.1 points. A page coordinate is metadata about the library entry rather than about the photograph, so it places the upload rather than the camera, and it stays behind when the file is downloaded. The two are worth keeping apart: the file's own coordinate is what leaves with the picture.
Interpretation
The reason behind that split is a guess added after the measurement, not a measurement. This study never asked a manufacturer why it writes what it writes, and no file can answer that question. What the files show is the split itself, and the split is the part a reader can rely on.
Where a serial number hides
A serial number is not in one place. On the whole files, this study’s own reading of the standard fields found one in 13.3% of files, and a tool that also opens the maker’s own note and the XMP packet found one in 26.7%, a difference of 13.3 points with an interval of 9.7 to 17.3 points. Both numbers are correct. They are answers to different questions, and the second is the question a reader means when asking what a photograph gives away.
The two readings are nested, and that is the sharper fact. Of the 80 files the wider reading finds a serial in, the narrower one misses 40. Going the other way it misses none: not one file holds a serial this study finds and the wider tool does not. So the narrower reading is not a different answer to the same question, it is a strict subset of the wider one. 13 of the whole files carry a serial in the manufacturer’s own internal field and in none of the standard ones at all.
The wider reading was adopted after the narrower one had already been taken, when a second method written independently reported finding more serials than this study had: the narrow pass had 22.3% of these files, against 26.7% now. That order matters and is stated rather than left out, because a reading widened after a result is known is how a study talks itself into a bigger number.
- Same files, not a second sample. The wider request was made on the same whole files, so it is a closer look at one sample rather than a new one, and both readings can be compared file by file.
- It moved toward the independent method, not away from it. The second method was written without seeing this one, and the wider reading agrees with it rather than with what this study had said before.
The narrow rows are still in the ledger, marked as superseded, so a reader who prefers them can bind these claims to those instead.
| Where the serial number sits | Files |
|---|---|
| In the maker’s own note, and nowhere else | 32 |
| In the XMP packet, and nowhere else | 10 |
| In the standard EXIF field, and nowhere else | 2 |
| In at least one of those three, found by the second tool | 80 |
| In at least one of the two this study reads itself | 40 |
| Files read whole and compared | 300 |
The same split appears on the whole frame, where the platform’s own metadata tool reports a serial in 11.5% of the files while reading them directly finds one in 13.4%, a difference of 2 points with an interval of 1.1 to 2.9 points. The platform’s field is not wrong; it is narrower than the file. Its documentation describes the field as the metadata for the file and says nothing about where a value comes from, so a reader of that field cannot tell whether a missing serial means the camera wrote none or the reader did not look.
The make does not settle it
Grouping by manufacturer shows how far the difference goes, and it also shows that a brand is not a promise. Canon files carry a serial in 45.4% and files under the name NIKON CORPORATION in 46.8%, while Sony files carry one in 1.2%. The clearest case in the table is Apple, whose 111 files carry a location in 72.1% of cases and a serial in none of them. The zero has a boundary too, and it is narrow: across every phone make taken together the serial rate stays at 0% to 0.9%.
| Camera make | Files | Location | Camera serial |
|---|---|---|---|
| Canon | 430 | 9.8% | 45.4% |
| NIKON CORPORATION | 331 | 7% | 46.8% |
| SONY | 168 | 10.7% | 1.2% |
| Panasonic | 137 | 15.3% | 8% |
| Apple, a phone maker | 111 | 72.1% | 0% |
| samsung, a phone maker | 84 | 41.7% | 0% |
| NIKON, the same manufacturer spelled differently | 67 | 3% | 0% |
| FUJIFILM | 66 | 9.1% | 16.7% |
| OLYMPUS IMAGING CORP. | 39 | 5.1% | 12.8% |
| Xiaomi, a phone maker | 33 | 57.6% | 0% |
The odd row is the one where a single manufacturer appears twice. NIKON CORPORATION is the name Nikon’s interchangeable-lens cameras write, and 46.8% of those files carry a serial. NIKON, with no second word, is what its compact cameras write, and not one of those 67 files carries a serial number. The zero is a boundary rather than a proof: a clean run of that size still permits a rate as high as 5.4 per cent. Every file in that second group came from a Coolpix compact. The serial is a property of the model line, not even of the brand.
| Camera model, with the most files in the sample | Files | Location | Camera serial |
|---|---|---|---|
| NIKON D5 | 33 | 0% | 97% |
| COOLPIX AW100 | 27 | 0% | 0% |
| NIKON D4 | 27 | 0% | 92.6% |
| Canon EOS 6D | 22 | 27.3% | 100% |
| NIKON D3S | 21 | 0% | 52.4% |
| NIKON D2Xs | 20 | 0% | 70% |
Interpretation
These are the models with twenty files or more, the threshold below which two values cannot be compared at all; every other model is in the dataset. Read the table as a count rather than as a rate: an interval narrow enough to set one model against another would need about a hundred files in every row, which a random draw of this size does not produce. The spread inside one brand is the part worth noticing, and the reason for it is not something this study measured, because the sample records the cameras enthusiasts chose to publish from rather than the cameras that were sold.
How the measuring program was checked
A number is only worth what its instrument is worth, so the instrument was tested before it was used. It is scored against 15 test files built for the purpose, whose contents were written down in advance, and it answers 165 checks on them without a failure. Each new check was then falsified on purpose: the reading was broken in the way the check exists to catch, and the suite had to fail before the check was trusted.
58%
is what a second execution of this method found for a camera make, against 57.2% in the first. It ran from the beginning on a fresh draw of its own, written by a session that could not see the first one’s data, and it reproduced every share of the drawn files: a location coordinate in 10.3% against 10.8%, and a serial readable in the file in 14% against 13.4%. That rules out one unlucky draw. It does not rule out a fault in the reading program, which is a different question and has its own checks below.
Two further checks ran against real files. A second reader was pointed at the same random frame, and over 150 files taken from the library the two readers agreed on every field compared, disagreeing on 0 of them, and a run of that size still permits a disagreement rate as high as 2.4 per cent. That is what makes a disagreement about the sample a fact about the sample rather than about the software.
The cheaper reading was checked against the complete one on 300 files read whole. Reading only the first 131072 bytes missed none of the serial numbers that the same reader finds in the whole file. That is a statement about where the reading stops, and its boundary is wide: a clean run of this many files still permits a miss rate as high as 1.2 per cent. It is not a statement about what the reading looks at. Both readings walk the standard fields, and neither opens the maker’s own note, which is where the 40 missed serials above actually sit.
Where this measurement is weak
- Commons is not the web. These are files that volunteers chose to publish in a library with an explicit licence, so the sample describes enthusiasts and institutions rather than everybody who owns a camera. It is not a sample of photographs in general and this page does not claim to be one.
- The maker’s note is read only on the subsample. The share of 26.7% rests on the files read whole, not on the whole frame, because the note sits in different places for different manufacturers and only a complete reading can be trusted to find it.
- The group split is ours, not the data’s. A file is counted as phone-made because its maker is on a list of phone manufacturers written into the method. The per-make table lets any reader rebuild the grouping differently.
- A make is a string in a file. One manufacturer can write more than one of them, which is why the table above shows two NIKON rows and two for samsung. Nothing here verifies that a make string is true.
- Presence, never content. The study never reads a serial value or a coordinate, so it cannot tell a real serial from a placeholder, and it cannot tell an original from a copy re-saved by software that carried the fields forward.
What would change the answer
- A different population would. Ordinary phone libraries and social platforms are the interesting comparison, and this method cannot reach them without an account on each. Commons is what a study can sample anonymously today.
- A larger sample would not. More files would narrow the interval around 13.4% and around the family difference of 28.2 points. It would not move the split itself, which is a property of what manufacturers write.
- Manufacturers changing their firmware would. The serial is a decision made in software, and a camera maker could stop writing one in the next generation exactly as Nikon’s compact line appears never to have written one.
What Viallo does with the file you upload
Viallo stores the file as it arrived and returns it through the download button, so the stored file is the same kind of file this study found a serial in 28.5% of. That is deliberate: keeping a photograph unaltered is what the product is for.
- The smaller copies are the cleaner ones: the display copy and the thumbnail are re-encoded when they are made, and the encoder writes a fresh file containing the picture and nothing else. The fields this study measures do not survive that step, so a link that is only looked at gives away far less than the file behind it. That is mechanical rather than a policy, and it holds for every photo on the service.
- What that means for you: a photograph you would rather not hand over with a serial number attached has to be cleaned on your device first, because a service that keeps your file faithfully keeps that too.
Sources
- ExifTool, EXIF tag names. The tag table this study used to separate the two serial fields: EXIF 0xA431 is the body serial number and 0xA435 is the lens serial number. exiftool.org/TagNames/EXIF.html
- MediaWiki, API help for imageinfo. Documents the metadata property this study compared against, and says nothing about which part of a file its value comes from: metadata for the version of the file. If the file has been revision deleted, a filehidden property will be returned. commons.wikimedia.org/w/api.php
Both pages were read and archived on 10 Sep 2026, the day the files were fetched. The platform documents what its metadata field is for and not where its value came from, which is why the difference between the two readings had to be measured rather than looked up.
The data behind this article
Every number on this page comes from the same table, and the files are listed below so anyone can check the work or reuse it.
- The measurements. One row per fetched file and one column per field read. It carries no image content, no coordinate values and no serial numbers, only whether a field was there. DATASET.csv
- The run and the instrument. Where the sample came from, how the reading program was checked, and the counts behind every group. DATASET.json
- The prediction, written first. The question, the predicted shares and the counting rules, all fixed before the first file was fetched, with every later change recorded beside it. PREREGISTRATION.md
- Licence. Released into the public domain under CC0, so it can be copied, changed and reused for any purpose without asking. LICENSE.txt
- The program that read the files. It asked for the bytes, read the metadata out of them, and wrote one row per request. analyze.py
- The dataset page. The published data, its licence and the citation in one place, for anyone who wants to cite it rather than read it. /studies/what-exif-data-photos-still-carry
Citing this study
Viallo, What EXIF Data Do Photos Still Carry?, measured 10 Sep 2026. https://www.viallo.app/blog/what-exif-data-photos-still-carry
Frequently Asked Questions
Do my photos still have a camera serial number in them?
It depends on what took the picture, and not in a small way. In this sample 28.5% of the files from a camera or a scanner carried one, against 1 of the 315 files from a phone. A serial number is written by the device, so the only reliable way to know is to read the file you are about to share.
Is a camera serial number a privacy problem?
It can outlast a location. A coordinate places one photograph in one place, while a serial number identifies one camera, so it ties together every photograph that camera ever took, across sites and across years. In this sample it survives in 28.5% of the files from a camera.
Why do some tools not show the serial number?
Because a serial number is written in more than one place. Reading the standard fields of the whole files found one in 13.3%, while a reader that also opens the maker’s own note found one in 26.7%. A viewer that reports no serial may simply be looking at one of the places and not the others.
How do I share a photo without its metadata?
Remove the metadata on your own device before the file goes anywhere, since that is the one step no platform can undo for you. Viallo keeps the stored file as it arrived, so the display copy is cleaner than the download and the download is what the camera wrote, with a camera make in 57.2% of the files measured here.