QR Code Data Encoder/Decoder

Explore how QR codes encode data at the bit level. Enter data and see the encoding mode, character count, bit stream, and error correction codewords.

Detected Mode
Character Count
1

Mode Indicator

The first 4 bits identify the encoding mode used.

=
2

Character Count Indicator

Indicates the number of characters in the data, with bit length depending on mode and version.

= characters

bits for mode, version

3

Data Bit Stream

... + more bits

data bits total

4

Complete Bit Stream Structure

Mode Indicator: 4 bits
Character Count:
Data:
Terminator: 0000
Total: ( bytes)

Character-by-Character Encoding

# Char Code Point Binary Bits

Showing first 50 of characters.

Error Correction Overhead

How to Use This Tool

  1. 1
    Enter your data and select encoding mode

    Type the text or number string you want to analyse and choose the encoding mode (numeric, alphanumeric, byte, or Kanji) to see how each mode represents your input.

  2. 2
    Explore the bit-level representation

    Step through the encoding stages: mode indicator, character count indicator, data codewords, and error-correction codewords, displayed as binary sequences with annotations.

  3. 3
    Examine the final module placement

    View how the data bits are mapped to the QR code grid, including the function patterns, data region, and the effect of the applied masking pattern.

Frequently Asked Questions

How does numeric mode encoding work at the bit level?

In numeric mode, groups of three decimal digits are converted to a 10-bit binary number (0–999 fits in 10 bits). If the digit count is not divisible by three, the remaining one or two digits are encoded in 4 or 7 bits respectively. The mode indicator for numeric mode is 0001 (4 bits), followed by a character count indicator whose length varies by version (10 bits for Version 1–9, 12 for Version 10–26, 14 for Version 27–40). This encoding is approximately 3.32 bits per digit, compared to 8 bits per character in byte mode, making it highly efficient for numeric-only payloads such as tracking numbers, serial numbers, or phone numbers.

What is Reed-Solomon error correction and how is it applied to QR codes?

Reed-Solomon (RS) codes are a class of polynomial error-correcting codes invented by Irving Reed and Gustave Solomon in 1960. In QR codes, RS is applied to groups of 8-bit data codewords, adding redundant error-correction codewords computed from Galois Field GF(256) arithmetic. The number of error-correction codewords per block is determined by the version and error-correction level. Each RS block can detect and correct up to t errors where 2t equals the number of error-correction codewords; beyond this capacity, the codewords can still detect but not correct errors. Multiple RS blocks are interleaved before placement in the symbol so that a physical burst error (such as a scratch) affects multiple blocks, reducing the probability that any single block exceeds its correction capacity.

What are the eight data masking patterns and why are they needed?

The eight masking patterns defined in ISO/IEC 18004 are XOR functions applied to the data modules of the QR code to prevent large areas of uniform dark or light modules, isolated modules, or patterns that resemble the finder patterns from appearing in the symbol. Such visual anomalies cause decoding errors in some scanner optical systems. Each of the eight mask patterns is applied in turn, and the resulting symbol is scored using a four-penalty rule system that assigns penalty points for runs, boxes, and finder-pattern-like sequences. The mask with the lowest total penalty score is selected and its index encoded in the format information strip. Different encoder implementations may choose different masks for the same data, which is why two encoders can produce visually different QR codes from identical input.

What is the quiet zone and how are its bits encoded?

The quiet zone is not encoded in bits — it is a physical border of blank (light) modules around the outside of the QR symbol, not part of the data region. Its purpose is to provide a contrast boundary that allows the scanner's image processor to locate the edges of the symbol against the background. ISO/IEC 18004 requires a minimum quiet zone of four module widths on all sides. The quiet zone carries no data and is not subject to error correction; it is simply part of the physical printing specification. Errors in the quiet zone (such as background pattern encroaching on the margin) are detected at the scanner hardware level before the RS decoder is invoked.

How is kanji mode different from UTF-8 byte mode for Japanese text?

Kanji mode uses 13-bit encoding for double-byte characters in the Shift-JIS character set, specifically the ranges 0x8140–0x9FFC and 0xE040–0xEBBF. The encoding subtracts a base value from each two-byte character and packs the result into 13 bits, achieving roughly 8 bits per character for standard Japanese kanji and kana — more efficient than encoding the same characters as UTF-8 (which requires 3 bytes = 24 bits per character in byte mode). However, Kanji mode only covers Shift-JIS encoded characters; emojis, Simplified Chinese characters, and other Unicode outside the Shift-JIS range must use byte mode with UTF-8. For modern multilingual content, byte mode with UTF-8 is more universal, though less space-efficient for pure Japanese text.