Skip to content

How diff works and how to read one

By · Developer Text Tools · 5 min read · Published

Developers read diffs all day: in pull requests, in git diff, in deployment previews and in test failures. Most of the time the output is obvious. Occasionally it is baffling, showing a whole block as removed and re-added when you only changed one line, or matching closing braces from different functions. Understanding how diff tools work explains these results and helps you get cleaner diffs. This guide covers the algorithm, the output formats and practical tips.

The problem diff solves

Given two versions of a text, a diff tool tries to find the smallest set of changes that turns the first into the second. It works on units, usually lines, and describes the result as lines kept, lines deleted and lines inserted. There is no "modified line" at the algorithm level: a changed line is a deletion followed by an insertion, which tools then display as a modification.

Longest common subsequence

The classic approach is to find the longest common subsequence (LCS) of the two texts: the longest sequence of lines that appear in both, in the same order, though not necessarily next to each other. Everything in the old text that is not part of the LCS was deleted; everything in the new text that is not part of it was inserted.

Old:  A  B  C  D  E
New:  A  C  D  X  E

LCS:  A  C  D  E
Diff: keep A, delete B, keep C, keep D, insert X, keep E

A straightforward LCS computation compares every line of one text with every line of the other, which takes time proportional to the product of their lengths. That is fine for files of a few thousand lines, but too slow for large inputs.

Myers and other algorithms

In 1986 Eugene Myers published an algorithm that finds a minimal diff in time proportional to the size of the input multiplied by the number of differences. When two files are mostly the same, which is the usual case, it is very fast. It is the default in Git and GNU diff.

A minimal diff is not always the most readable one. When several equally short edit scripts exist, Myers may pick one that matches a stray closing brace or blank line from the wrong place. Git offers alternatives:

  • --patience first matches lines that are unique in both files, such as function signatures, and then diffs between those anchors. It produces more intuitive results for code that was moved or restructured.
  • --histogram is a faster refinement of patience and is often the best choice; you can make it the default with git config --global diff.algorithm histogram.

Reading a unified diff

The unified format is what git diff, diff -u and patch files use:

--- a/config.yaml
+++ b/config.yaml
@@ -3,7 +3,7 @@ server:
   host: 0.0.0.0
   port: 8080
   workers: 4
-  timeout: 30
+  timeout: 60
   log_level: info
   cors: true
   compress: true
  • The --- and +++ lines name the old and new files.
  • Each hunk starts with an @@ header. -3,7 +3,7 means the hunk covers 7 lines starting at line 3 in the old file, and 7 lines starting at line 3 in the new file. Git appends the nearest preceding function or section heading for orientation.
  • Lines starting with a space are unchanged context, three lines by default. Lines starting with - were removed and + were added.

Context lines let a patch be applied even if the file has shifted slightly, and help reviewers understand where a change sits.

Side-by-side and word-level views

Side-by-side views show the old version on the left and the new on the right, aligned so unchanged lines sit next to each other. They are easier to read for changes spread across many lines; unified views are more compact and better for small edits and for patches.

Line-based diffs mark the whole line even if only one character changed. Better tools run a second, finer diff on each pair of changed lines to highlight the exact words or characters that differ. This is what makes a renamed variable inside a long line easy to spot. The Text Diff tool does both: line-level matching with word-level highlights, in side-by-side or unified view.

Why some diffs look noisy

Whitespace changes

Re-indenting a block, converting tabs to spaces or removing trailing whitespace changes every affected line. Use an ignore-whitespace option, git diff -w, to see only content changes, then check the whitespace separately if it matters, as it does in YAML and Python.

Line endings

A file saved with Windows line endings (CRLF) differs on every single line from the same file with Unix endings (LF). If a diff shows the entire file as changed but the lines look identical, line endings are the likely cause. A .gitattributes file with * text=auto normalises them in a repository.

Reordered content

Diff algorithms find what stayed in the same order. If you move a function or sort keys in a JSON file, the diff shows the moved content as deleted in one place and inserted in another. Git's --color-moved option highlights moved blocks differently. For data files, normalise first: format both JSON documents with sorted keys in the JSON Formatter and then compare.

Long lines

Minified files and long paragraphs of prose produce enormous one-line changes. For documentation, writing one sentence per line makes diffs dramatically clearer.

Tips for reviewable diffs

  1. Separate formatting changes from logic changes, in different commits or pull requests.
  2. Run the same formatter everyone else runs, so your diff contains only your change.
  3. Keep line endings and indentation consistent across the team with .editorconfig and .gitattributes.
  4. When a diff looks confusing, try another algorithm or ignore whitespace before assuming the change is wrong.
  5. For text that is not in Git, such as configuration from two servers or two API responses, use a standalone diff tool rather than comparing by eye. Our eyes are very good at missing a single changed character.