What Is a Checksum?

Technician comparing storage devices at a workstation with an abstract checksum verification display

A checksum is a value calculated from a block of data so that a computer can later check whether the data changed. The sender or publisher calculates the value with a specified algorithm and stores or transmits it alongside the file, message, or disk block. A receiving system runs the same calculation on the data it actually received. If the two results match, the data passed that particular check; if they differ, something changed. NIST therefore describes a checksum as a content-dependent value used to detect changes, errors, or manipulation. That compact result is much smaller than the original data, so repeated verification is practical even when the underlying file contains gigabytes.

The basic idea is deterministic: the same input processed by the same algorithm produces the same result. Change even one bit and the result will usually change, giving software a compact way to compare large objects without comparing every byte to a second copy. A checksum can travel in a file manifest, sit inside a network packet, or be stored as filesystem metadata. The algorithm matters, because a small arithmetic checksum, a cyclic redundancy check, and a cryptographic hash are designed for different levels and kinds of error detection. Programs must agree on details such as byte order, input range, and output format, because a correct calculation with different rules will not match.

Simple sums are inexpensive and can catch common accidental mistakes, but some changes can cancel one another and leave the same total. Internet checksums combine fixed-size words and are useful for detecting many transmission errors. Cyclic redundancy checks, or CRCs, are especially good at detecting burst errors in noisy links and storage media. Cryptographic hash functions such as SHA-256 are also used as checksums in everyday language. They produce a fixed-length digest with much stronger resistance to finding two useful inputs that share a result. Designers choose among these tools by considering expected faults, processing cost, result size, and whether a malicious opponent is part of the threat model.

Checksums appear throughout computing. Download sites publish a digest so users can compare a downloaded installer with the publisher’s expected value. Backup and synchronization tools use checksums to notice damaged or changed files. Filesystems and storage devices can attach checks to blocks and verify them when data is read. Network protocols check headers or payloads before accepting a packet. Archives can store a checksum for every member, allowing extraction software to warn that an item no longer matches the value recorded when the archive was created. These checks can also help identify silent corruption that causes no obvious error when a file is copied but becomes important much later.

A matching checksum is evidence of consistency, not a complete security guarantee. Every fixed-length result represents many possible inputs, so collisions are mathematically unavoidable. Weak algorithms may allow an attacker to alter data while preserving the same checksum. Even a strong cryptographic hash cannot prove who supplied a file if an attacker can replace both the file and the published hash. Authenticity requires a trusted comparison value, a digital signature, or a message-authentication code delivered through a channel the attacker cannot silently rewrite. A signed manifest solves part of that trust problem by letting software verify that the list of expected digests came from an authorized publisher.

When checking a file manually, first identify the exact algorithm the publisher used, then calculate that algorithm over the unmodified file and compare every character of the result. A mismatch can mean corruption, an incomplete transfer, a different software version, or deliberate modification; it does not identify which cause occurred. A match means the file agrees with the supplied value under that algorithm. The practical lesson is simple: checksums are excellent alarms for change, but their reliability depends on the algorithm and on whether the expected value itself can be trusted. Automated tools often normalize capitalization and ignore spaces in a displayed digest, but they must never silently switch to a different algorithm.

Explore more "Explainers"

Discover additional explainers across politics, science, business, technology, and other fields. Each explainer breaks down a complex idea into clear, everyday language—helping you better understand how major concepts, systems, and debates shape the world around us.