⚡ Zig Guide LiveUnofficialbut fully verified
✓ Zig 0.17.0What's newOn an older Zig?Blog

The File That Was Not There

· Written against Zig 0.17.0-dev.2056+79a9897cd

The runnable examples on this page are checked every night and currently run on Zig 0.17.0. The plain code blocks are not checked, and may no longer compile.

Monday, 9:12 AM

Vishaka had written a small tool for the team.

It was called phrase.

It made passphrases out of random words, like grape-habit-buddy-hammer-diner-river.

The team used it for shared test accounts, so nobody had to invent a password.

On Monday morning, a message came from Tejas.

~/Downloads $ ./phrase
error: FileNotFound

It worked on Vishaka’s laptop.

It had worked on Vishaka’s laptop all of last week.

9:20 AM: what the program was looking for

The words came from a file called words.txt.

It has 256 words, one per line.

The first version read it like this:

const text = try std.Io.Dir.cwd().readFileAlloc(io, "words.txt", init.gpa, .limited(1 << 20));

cwd() is the current directory.

That is the directory you are in when you run the program, not the directory where the program is.

Vishaka always ran phrase from the project folder, and words.txt was in it.

Tejas ran it from ~/Downloads.

There was no words.txt there.

So the program was correct, and it still failed on every machine except Vishaka’s.

9:31 AM: three ways to fix it

Vishaka thought of three fixes.

The first was to send two files and tell people to keep them together.

People would not keep them together.

The second was to look for words.txt next to the program instead of in the current directory.

That is better, but it is still two files.

Somebody would copy one and forget the other.

The third was to put the words inside the program.

Zig can do this with one line:

const words_txt = @embedFile("words.txt");

@embedFile runs when the program is compiled, not when it runs.

The compiler reads words.txt from beside the source file, and stores its bytes in the program.

words_txt is a pointer to those bytes.

Its length is known at compile time: 1,582 bytes.

At run time, there is no file to open, and nothing that can fail to open.

9:48 AM: let the compiler do the parsing

Vishaka could have split words_txt into lines every time the program started.

But the words never change while the program runs.

So Vishaka asked the compiler to split them instead, once, while it builds the program.

snippets/blog/phrase.zigwords
/// The table, built while compiling. A bad line in words.txt is a compile
/// error, so a program with a broken word list cannot be built at all.
const words = parseWords(words_txt);

That line is outside any function.

In Zig, a const outside a function is computed while compiling.

So parseWords runs inside the compiler.

snippets/blog/phrase.zigparseWords
fn parseWords(comptime text: []const u8) [word_count][]const u8 {
    // The duplicate check below compares every pair of words, about 32,000
    // comparisons. The compiler stops a long computation unless it is told to
    // expect one.
    @setEvalBranchQuota(2_000_000);

    var list: [word_count][]const u8 = undefined;
    var n: usize = 0;
    var line_number: usize = 0;
    var lines = std.mem.splitScalar(u8, text, '\n');
    while (lines.next()) |line| {
        line_number += 1;
        if (line.len == 0) continue;

        for (line) |c| {
            if (c < 'a' or c > 'z') @compileError(std.fmt.comptimePrint(
                "words.txt line {d}: \"{s}\" must be one word in lowercase a to z",
                .{ line_number, line },
            ));
        }
        for (list[0..n]) |seen| {
            if (std.mem.eql(u8, seen, line)) @compileError(std.fmt.comptimePrint(
                "words.txt line {d}: \"{s}\" is already in the list",
                .{ line_number, line },
            ));
        }
        if (n == word_count) @compileError(std.fmt.comptimePrint(
            "words.txt has more than {d} words",
            .{word_count},
        ));

        list[n] = line;
        n += 1;
    }
    if (n != word_count) @compileError(std.fmt.comptimePrint(
        "words.txt has {d} words, and it needs exactly {d}",
        .{ n, word_count },
    ));
    return list;
}

The result is a table of 256 words, already split, stored in the program.

When phrase starts, the work is already done.

The checks in parseWords also run inside the compiler.

@compileError stops the build with a message.

std.fmt.comptimePrint builds that message, with the line number in it.

@setEvalBranchQuota is there because the duplicate check makes about 32,000 comparisons.

The compiler stops long computations unless you tell it to expect one.

Wednesday, 2:05 PM: Tejas edits the list

Tejas wanted a home city in the word list.

Tejas replaced line 87, crate, with navi mumbai.

Then Tejas rebuilt the program.

phrase.zig:35:37: error: words.txt line 87: "navi mumbai" must be one word in lowercase a to z
            if (c < 'a' or c > 'z') @compileError(std.fmt.comptimePrint(
                                    ^~~~~~~~~~~~~
phrase.zig:15:25: note: called at comptime here
const words = parseWords(words_txt);
              ~~~~~~~~~~^~~~~~~~~~~

A word with a space in it would put a space inside some passphrases, between two dashes.

The compiler refused to build it.

Next, Tejas tried maple.

phrase.zig:41:46: error: words.txt line 200: "maple" is already in the list

Line 200 was mosaic, and maple was already on line 194.

A duplicate word makes some passphrases more likely than others, which makes them easier to guess.

Then a line was deleted by mistake.

phrase.zig:54:26: error: words.txt has 255 words, and it needs exactly 256

Each mistake stopped on Tejas’s machine, at build time, with the line number.

A program with a broken word list was never built.

Why exactly 256

One byte holds a number from 0 to 255.

So with 256 words, one random byte picks one word, with no bias and no extra arithmetic:

snippets/blog/phrase.zigwritePhrase
/// One word per byte, joined with dashes.
fn writePhrase(out: *Io.Writer, bytes: []const u8) !void {
    for (bytes, 0..) |b, i| {
        if (i > 0) try out.writeByte('-');
        try out.writeAll(words[b]);
    }
    try out.writeByte('\n');
}

words[b] is the whole lookup.

Each word adds 8 bits, so six words make a 48-bit passphrase.

Press Run

const std = @import("std");

const Io = std.Io;

/// The bytes of words.txt, read by the compiler and stored in the program.
const words_txt = @embedFile("words.txt");

/// The table, built while compiling. A bad line in words.txt is a compile
/// error, so a program with a broken word list cannot be built at all.
const words = parseWords(words_txt);

/// 256 words, so one random byte picks one word, and each word adds 8 bits.
const word_count = 256;

fn parseWords(comptime text: []const u8) [word_count][]const u8 {
    // The duplicate check below compares every pair of words, about 32,000
    // comparisons. The compiler stops a long computation unless it is told to
    // expect one.
    @setEvalBranchQuota(2_000_000);

    var list: [word_count][]const u8 = undefined;
    var n: usize = 0;
    var line_number: usize = 0;
    var lines = std.mem.splitScalar(u8, text, '\n');
    while (lines.next()) |line| {
        line_number += 1;
        if (line.len == 0) continue;

        for (line) |c| {
            if (c < 'a' or c > 'z') @compileError(std.fmt.comptimePrint(
                "words.txt line {d}: \"{s}\" must be one word in lowercase a to z",
                .{ line_number, line },
            ));
        }
        for (list[0..n]) |seen| {
            if (std.mem.eql(u8, seen, line)) @compileError(std.fmt.comptimePrint(
                "words.txt line {d}: \"{s}\" is already in the list",
                .{ line_number, line },
            ));
        }
        if (n == word_count) @compileError(std.fmt.comptimePrint(
            "words.txt has more than {d} words",
            .{word_count},
        ));

        list[n] = line;
        n += 1;
    }
    if (n != word_count) @compileError(std.fmt.comptimePrint(
        "words.txt has {d} words, and it needs exactly {d}",
        .{ n, word_count },
    ));
    return list;
}

/// One word per byte, joined with dashes.
fn writePhrase(out: *Io.Writer, bytes: []const u8) !void {
    for (bytes, 0..) |b, i| {
        if (i > 0) try out.writeByte('-');
        try out.writeAll(words[b]);
    }
    try out.writeByte('\n');
}

pub fn main(init: std.process.Init) !void {
    const io = init.io;

    var out_buf: [1024]u8 = undefined;
    var stdout = Io.File.stdout().writerStreaming(io, &out_buf);
    const out = &stdout.interface;
    defer out.flush() catch {};

    try out.print("{d} words, {d} bytes of words.txt, inside the program\n", .{ words.len, words_txt.len });
    try out.print("first word: {s}, last word: {s}\n\n", .{ words[0], words[words.len - 1] });

    // Six words of 8 bits each is a 48-bit passphrase.
    var bytes: [6]u8 = undefined;
    try out.print("{d} words, {d} bits:\n", .{ bytes.len, bytes.len * 8 });

    // A fixed seed, so these three are the same on every run and the page can
    // check them. Never use a seeded generator for a real passphrase: anyone
    // with the seed gets the same words.
    var prng: std.Random.Xoshiro256 = .init(2026);
    for (0..3) |_| {
        prng.random().bytes(&bytes);
        try writePhrase(out, &bytes);
    }
    try out.flush();

    // The real one. `io.random` comes from the operating system's secure
    // random source, so this line is different every time. It goes to stderr,
    // because the check that runs this program compares stdout, and this line
    // can never match.
    io.random(&bytes);
    var err_buf: [256]u8 = undefined;
    var stderr = Io.File.stderr().writerStreaming(io, &err_buf);
    try stderr.interface.writeAll("\nyours, from the system's random source:\n");
    try writePhrase(&stderr.interface, &bytes);
    try stderr.interface.flush();
}

The program ran in your browser.

There was no words.txt to download.

The words came inside the 51 KB of WebAssembly that the page fetched.

The three passphrases under “6 words, 48 bits” are the same every time.

They come from a generator with a fixed seed, so this page can check them every night.

The last one, under “yours”, is different every time you press Run.

It comes from io.random, which reads from the operating system’s secure random source.

Never use a seeded generator for a real passphrase.

Anyone who knows the seed gets the same words.

Friday: one file for everyone

Vishaka built phrase for the three systems the team used.

$ zig build-exe -O ReleaseSmall -fstrip -target x86_64-linux-musl phrase.zig
$ zig build-exe -O ReleaseSmall -fstrip -target x86_64-windows phrase.zig
$ zig build-exe -O ReleaseSmall -fstrip -target aarch64-macos phrase.zig
TargetSize
Linux, x86-64153 KB
Windows, x86-64442 KB
macOS, Apple silicon166 KB

Each one is a single file.

Each one has the word list inside it.

Tejas ran the new one from ~/Downloads.

~/Downloads $ ./phrase
256 words, 1582 bytes of words.txt, inside the program

Check it yourself

You can see the words inside the program with strings, which prints the readable text in a binary:

$ strings -n 3 phrase | grep -c -x -F -f words.txt
256

All 256 words are in the file.

-n 3 matters here, because strings skips anything shorter than four characters by default, and fig has three.

What else to embed

@embedFile suits any small file that the program cannot work without, such as:

  • a SQL schema or a list of migrations,
  • an HTML template,
  • a default config file,
  • the text for --help,
  • a test fixture.

It is the wrong tool for a large file, because every byte ends up in the program.

These chapters explain the parts in more detail:

About the numbers

Vishaka and Tejas are made up.

The error messages, the sizes and the 256-word list are real.

Every one was produced with the Zig version shown at the top of this post.

The 256-word list is short on purpose, so the byte-to-word step is easy to see.

Six words from it give 48 bits.

That is fine for a test account.

For a password that guards something important, use more words, or a longer list: with the 7,776-word list that diceware uses, six words give about 77 bits.

← All posts