Reference guide · cited sources
How Image Compression Actually Works
Most image problems become obvious once you know roughly how compression works. Why a screenshot balloons as a JPG, why a quality slider is not a percentage, why two photos at identical settings differ in size by a factor of five — all of it follows from a small number of mechanisms that are not difficult to understand.
By Prashantkumar Kishanrao Sundge · Published · Reviewed
What this guide establishes
- Lossless compression finds redundancy and encodes it compactly. Decompressing returns the original bytes exactly.
- Lossy compression additionally discards information judged hard to see. That discard is permanent and not reversible.
- A quality slider scales how aggressively information is rounded away. It is an encoder instruction, not a percentage retained.
- Compression exploits predictability, which is why smooth skies compress enormously and gravel barely compresses at all.
- JPEG handles smooth tonal variation well and sharp edges badly. PNG is the reverse. That single fact explains most format advice.
The idea underneath everything
An uncompressed image is a list of colour values, one per pixel. A 1200×800 photograph in three 8-bit channels is 2,880,000 bytes before anything clever happens. Every compression method is an attempt to describe those same values in fewer bytes.
The way to do that is to exploit predictability. If the next pixel is usually similar to the one before it, you do not need to store its full value — you can store the small difference, which takes fewer bits. If a particular colour appears constantly and another appears rarely, you can give the common one a short code and the rare one a long code, and come out ahead on average.
This is why image content matters so much. Compression is a bet that the data is predictable. A photograph of a clear sky is enormously predictable and compresses to a fraction of its raw size. A photograph of gravel, foliage, or static is close to unpredictable, and there is very little for any compressor to exploit. Two images with identical dimensions and identical settings can differ in file size by a factor of five for this reason alone.
Lossless compression
Lossless methods reorganise the data without discarding any of it. Decompressing returns the original bytes exactly, and no amount of repeated compression and decompression degrades anything.
PNG works in two stages. First it filters each row, replacing pixel values with the difference from a prediction based on neighbouring pixels. In a smooth gradient this turns a row of steadily increasing values into a row of near-zero differences, which is far more compressible. Then it applies DEFLATE, a general-purpose method that combines matching repeated sequences against recent data with assigning shorter codes to more frequent symbols.
The consequence is that PNG is extraordinarily good at flat colour and sharp geometric shapes — a screenshot, a diagram, an interface graphic — where the prediction is usually exactly right and the differences are mostly zero. It is poor at photographic noise, where the prediction is usually wrong and the differences are as unpredictable as the original.
Lossless compression has a hard floor. Genuinely random data cannot be compressed at all; there is no redundancy to exploit. This is not a limitation of any particular algorithm but a mathematical fact, and it is why a heavily detailed photograph stays large in any lossless format.
Lossy compression and the frequency idea
Lossy methods go further by discarding information deliberately, choosing what to lose based on what human vision is least likely to notice. JPEG is the canonical example and the mechanism is worth following, because it explains almost all of JPEG's behavior.
JPEG divides the image into small blocks, conventionally 8×8 pixels, and transforms each block from pixel values into frequency components. Low frequencies describe the broad, gradual variation across the block — its overall brightness and gentle shading. High frequencies describe rapid variation: fine texture and sharp edges. The transform is lossless on its own; it is a change of description, not a reduction.
The lossy step is quantisation. Each frequency component is divided by a value from a table and rounded to a whole number. High-frequency components are divided by larger values, so many of them round to zero and disappear entirely. Zeros compress to almost nothing, which is where the size reduction comes from.
JPEG also stores colour information at lower resolution than brightness, because human vision is markedly less sensitive to fine colour detail than to fine brightness detail. This alone removes a substantial fraction of the data before anything else happens.
What the quality slider actually does
The quality number scales that table of divisors. A higher number means smaller divisors, less rounding, fewer components lost, and a larger file. A lower number means larger divisors, more components rounded to zero, and a smaller file.
This is why a quality number is not a percentage of the original retained. It is a parameter controlling how aggressively one step of the process rounds, and the relationship between that parameter and visible quality is neither linear nor consistent between encoders. Two libraries can both offer 'quality 80' and produce measurably different results from the same input.
It is also why the loss is irreversible. Once a value has been divided and rounded, the original cannot be determined — many different inputs round to the same result. No software can undo JPEG compression, however sophisticated, because the information required to do so was not kept.
And it explains generation loss. Re-saving an already-compressed JPEG runs the whole process again on data that has already been through it, rounding what was already rounded. Each save is permanent, and several rounds of ordinary editing produce visible degradation that cannot be recovered.
Why artifacts look the way they do
Each of JPEG's characteristic failures maps directly onto the mechanism. Blocking — the faint grid visible in smooth areas like skies — happens because each 8×8 block is quantised independently, so adjacent blocks can end up with slightly different average values and the boundary becomes visible.
Ringing, the halo of disturbed pixels around sharp high-contrast edges, happens because a sharp edge is made mostly of high-frequency components. Discarding them leaves the edge to be reconstructed from the low frequencies that remain, which overshoot and undershoot around the transition. Text on a plain background is the worst case, which is exactly why screenshots should not be JPEGs.
Colour bleeding, where colour smears across a boundary, comes from storing colour at reduced resolution. A sharp red-on-white edge has its colour information averaged over a larger area than its brightness information, so the red appears to leak.
Recognising the artifact tells you the cause, and the cause tells you the fix. Blocking and ringing mean the quality was too low or the content was wrong for the format — not that the file is corrupt.
What newer formats do differently
WebP, AVIF, and other more recent formats use the same broad strategy — transform, discard what is hard to see, encode compactly — with more sophisticated machinery at each step. They predict blocks from already-decoded neighbouring blocks in more ways, use variable block sizes rather than a fixed grid, and apply filters that smooth block boundaries before they become visible.
The result is meaningfully smaller files at comparable visual quality, and fewer of the characteristic artifacts at aggressive settings. They also fix capability gaps: WebP and AVIF both support transparency in lossy mode, which JPEG cannot do at all.
The cost is compatibility and computation. Newer formats are slower to encode, and support is broad but not universal — particularly in upload forms, desktop applications, and email clients, where accepted-format lists are often years out of date. This is the entire trade-off, and it is why a better format is not automatically the right choice for a given file.
Practical consequences
Choose the format by content, because the mechanisms are content-specific. Photographs are mostly smooth tonal variation, which is what lossy frequency-based compression handles well. Screenshots, diagrams, and text are mostly sharp edges, which is what it handles worst and what lossless prediction handles best.
Compress once, at the end. Keep a lossless master, edit the master, and export a fresh lossy copy whenever you need one. This way compression happens a single time rather than once per edit, and generation loss never accumulates.
Expect content to dominate settings. If a file will not fit under a size limit, the composition is a legitimate variable — a subject against a plain background compresses far better than the same subject against foliage. This is not a trick; it is the direct consequence of compression being a bet on predictability.
And do not trust percentage claims. A tool advertising a fixed reduction is describing its result on some image. Yours will differ, and an already-optimised file cannot be reduced dramatically again without visible damage.
How this guide was written
This is a reference guide. It explains published specifications, platform rules, and format behavior, and every factual claim is attributed to the sources listed below. It does not report an in-house PixelConvert measurement, and no result here should be read as one.
Read how PixelConvert separates measured and reference guides →Limitations
- This is a conceptual summary of how the major formats work, simplified deliberately; it is not a specification-level account, and individual encoders make implementation choices not described here.
- It reports no in-house PixelConvert measurement. The measured figures cited across this site appear in the guides that produced them.
- Format capabilities and browser support change over time; verify current support before relying on a newer format for delivery.
Frequently asked questions
- What is the difference between lossless and lossy compression?
- Lossless reorganises data without discarding any of it, so decompression returns the original exactly. Lossy additionally throws away information judged hard to see, which makes files much smaller and cannot be undone.
- Why does a quality setting of 80 not mean 80% of the original?
- The number scales a table of divisors used to round away frequency information. It is an encoder instruction, not a proportion retained, and the same number in a different library produces a different result.
- Why do two photos at identical settings have very different file sizes?
- Compression exploits predictability. A smooth sky is highly predictable and compresses enormously; gravel or foliage is nearly unpredictable and barely compresses. Content dominates settings.
- Why does text look so bad in a JPG?
- Sharp edges are made mostly of high-frequency components, and those are exactly what JPEG quantisation discards first. The edge is then reconstructed from what remains, producing the halo known as ringing.
- Can JPEG compression be undone?
- No. The loss happens by rounding, and rounding is not reversible — many different original values produce the same rounded result. Tools that appear to repair JPEGs are smoothing artifacts, not recovering data.
- Are newer formats like WebP and AVIF always better?
- They generally produce smaller files at comparable quality and support transparency in lossy mode. They are slower to encode and less universally accepted, particularly by upload forms and older applications, so the right choice depends on the destination.