The 2 GB Log File at 11 PM
The runnable examples on this page are checked every night and currently run on Zig 0.17.0. The plain code blocks are not checked, and may no longer compile.
10:52 PM
Shivam’s phone buzzed on the kitchen table.
It was the alert nobody wants: Checkout error rate above 5%.
Shivam was on call this week.
That means when production breaks at night, Shivam is the one who fixes it.
Shivam opened the laptop.
A message from Urvashi in support was already waiting.
Customers say payment hangs and then fails. Started maybe 20 minutes ago?
“Maybe” was not good enough.
Shivam needed to know which part of the site was slow, and when it started.
10:58 PM
Shivam logged in to the web server, web-03, and looked at the access log.
$ ls -lh /var/log/shop/access.log
-rw-r----- 1 shop shop 2.0G Oct 3 22:58 /var/log/shop/access.log
Two gigabytes.
Fifty million lines, one line for each request.
Each line looked like this:
2026-10-03T22:41:07Z POST /api/checkout 504 5000
The time, the method, the path, the status code, and how many milliseconds the request took.
Shivam knew what the report had to show:
- how many requests got each status code,
- which endpoint was slow,
- and the time of the first
504, which means the request timed out.
The server had awk.
It did not have Python.
And this week, nobody was allowed to install software on production servers.
The report could be written in awk.
But a per-endpoint average, a maximum, and “the first 504” is a long awk line.
Shivam wanted something to test on the laptop first, and trust on the server.
So Shivam opened a new file: logscan.zig.
11:04 PM: one line at a time
The smallest piece came first: turning one line of text into fields.
/// Returns null for a line we cannot use. An aborted request is logged with
/// `-` where the milliseconds go, and one bad line must not stop the scan.
fn parseLine(text: []const u8) ?Line {
var fields = std.mem.tokenizeScalar(u8, text, ' ');
const time = fields.next() orelse return null;
_ = fields.next() orelse return null; // the method: GET, POST
const path = fields.next() orelse return null;
const status = std.fmt.parseInt(u16, fields.next() orelse return null, 10) catch return null;
const ms = std.fmt.parseInt(u32, fields.next() orelse return null, 10) catch return null;
return .{ .time = time, .path = path, .status = status, .ms = ms };
}The function returns ?Line.
The ? means “a Line, or null”.
Shivam knew the log had bad lines in it.
When a customer closes the tab in the middle of a request, the server writes - instead of the milliseconds.
parseInt fails on -.
So parseLine returns null, and the caller counts the line as skipped.
The caller adds one to skipped and moves on to the next line.
11:10 PM: reading 2 GB with 64 KiB
The server had less than 1 GB of free memory.
So reading the whole file into memory first was not an option.
Shivam gave the reader a buffer of 64 KiB instead:
var buf: [64 * 1024]u8 = undefined;
var reader = file.reader(io, &buf);
The reader fills the buffer from the file.
takeDelimiter('\n') returns the next line as a slice of that buffer.
When the buffer is used up, the reader refills it.
The file can be any size.
The memory used for reading stays at 64 KiB.
fn scan(r: *Report, in: *Io.Reader) !void {
while (true) {
const text = in.takeDelimiter('\n') catch |err| switch (err) {
// A line longer than the buffer. Skip it and keep going.
error.StreamTooLong => {
r.bytes += try in.discardDelimiterInclusive('\n');
r.lines += 1;
r.skipped += 1;
continue;
},
else => |e| return e,
} orelse break;
r.lines += 1;
r.bytes += text.len + 1;
const line = parseLine(text) orelse {
r.skipped += 1;
continue;
};
try r.add(line);
}
}There is one more case here.
A line longer than the whole buffer gives error.StreamTooLong.
The loop skips that line and goes on.
11:17 PM: the bug that almost happened
For each endpoint, Shivam kept a running total in a hash map.
The first version stored line.path as the key.
Then Shivam stopped.
line.path is a slice of the reader’s buffer.
On the next refill, those bytes are overwritten.
Every key in the map would change to some other text, and nothing would report an error.
So the program copies each key the first time it is seen:
fn add(r: *Report, line: Line) !void {
if (line.status < r.statuses.len) r.statuses[line.status] += 1;
if (line.status == 504 and r.first_504_len == 0) {
const n = @min(line.time.len, r.first_504.len);
@memcpy(r.first_504[0..n], line.time[0..n]);
r.first_504_len = n;
}
const slot = try r.endpoints.getOrPut(r.gpa, line.path);
if (!slot.found_existing) {
// `line.path` points into the reader's buffer, and the next read
// overwrites it. The map keeps the key, so it needs its own copy.
slot.key_ptr.* = try r.gpa.dupe(u8, line.path);
slot.value_ptr.* = .{};
}
const e = slot.value_ptr;
e.requests += 1;
e.total_ms += line.ms;
e.max_ms = @max(e.max_ms, line.ms);
}In many languages, reading a line gives you a new string that you own.
In Zig, the line is a view into the reader’s buffer.
That is why reading a line allocates nothing.
It is also why a line you want to keep must be copied.
11:24 PM: test it on the laptop
The real log was not on the laptop.
So Shivam gave the program a second mode.
With no file name, it writes a sample log of 50,000 lines, then scans that.
The sample is one hour of made-up traffic.
From 22:40, the checkout endpoint gets slow and starts timing out, like the real one.
Press Run to see the report.
const std = @import("std");
const Io = std.Io;
pub fn main(init: std.process.Init) !void {
const io = init.io;
const gpa = init.gpa;
var out_buf: [4096]u8 = undefined;
var stdout = Io.File.stdout().writerStreaming(io, &out_buf);
const out = &stdout.interface;
defer out.flush() catch {};
const dir = Io.Dir.cwd();
const args = try init.minimal.args.toSlice(init.arena.allocator());
const path = if (args.len > 1) args[1] else blk: {
try writeSampleLog(io, dir, "access.log", 50_000);
break :blk "access.log";
};
const file = try dir.openFile(io, path, .{});
defer file.close(io);
// The whole program reads through these 64 KiB. A 2 GB file and a 2 KB
// file use the same amount of memory for reading.
var buf: [64 * 1024]u8 = undefined;
var reader = file.reader(io, &buf);
var report: Report = .init(gpa);
defer report.deinit();
try report.scan(&reader.interface);
try report.print(out, buf.len);
}
/// One request, as the server logs it:
/// `2026-10-03T22:41:07Z GET /api/checkout 504 5000`
const Line = struct {
time: []const u8,
path: []const u8,
status: u16,
ms: u32,
};
/// Returns null for a line we cannot use. An aborted request is logged with
/// `-` where the milliseconds go, and one bad line must not stop the scan.
fn parseLine(text: []const u8) ?Line {
var fields = std.mem.tokenizeScalar(u8, text, ' ');
const time = fields.next() orelse return null;
_ = fields.next() orelse return null; // the method: GET, POST
const path = fields.next() orelse return null;
const status = std.fmt.parseInt(u16, fields.next() orelse return null, 10) catch return null;
const ms = std.fmt.parseInt(u32, fields.next() orelse return null, 10) catch return null;
return .{ .time = time, .path = path, .status = status, .ms = ms };
}
const Endpoint = struct {
requests: u64 = 0,
total_ms: u64 = 0,
max_ms: u32 = 0,
fn average(e: Endpoint) u64 {
return e.total_ms / e.requests;
}
};
const Report = struct {
gpa: std.mem.Allocator,
lines: u64 = 0,
bytes: u64 = 0,
skipped: u64 = 0,
/// One counter per possible status code. No allocation, no hashing.
statuses: [600]u64 = @splat(0),
endpoints: std.StringHashMapUnmanaged(Endpoint) = .empty,
first_504: [20]u8 = undefined,
first_504_len: usize = 0,
fn init(gpa: std.mem.Allocator) Report {
return .{ .gpa = gpa };
}
fn deinit(r: *Report) void {
var keys = r.endpoints.keyIterator();
while (keys.next()) |key| r.gpa.free(key.*);
r.endpoints.deinit(r.gpa);
}
fn scan(r: *Report, in: *Io.Reader) !void {
while (true) {
const text = in.takeDelimiter('\n') catch |err| switch (err) {
// A line longer than the buffer. Skip it and keep going.
error.StreamTooLong => {
r.bytes += try in.discardDelimiterInclusive('\n');
r.lines += 1;
r.skipped += 1;
continue;
},
else => |e| return e,
} orelse break;
r.lines += 1;
r.bytes += text.len + 1;
const line = parseLine(text) orelse {
r.skipped += 1;
continue;
};
try r.add(line);
}
}
fn add(r: *Report, line: Line) !void {
if (line.status < r.statuses.len) r.statuses[line.status] += 1;
if (line.status == 504 and r.first_504_len == 0) {
const n = @min(line.time.len, r.first_504.len);
@memcpy(r.first_504[0..n], line.time[0..n]);
r.first_504_len = n;
}
const slot = try r.endpoints.getOrPut(r.gpa, line.path);
if (!slot.found_existing) {
// `line.path` points into the reader's buffer, and the next read
// overwrites it. The map keeps the key, so it needs its own copy.
slot.key_ptr.* = try r.gpa.dupe(u8, line.path);
slot.value_ptr.* = .{};
}
const e = slot.value_ptr;
e.requests += 1;
e.total_ms += line.ms;
e.max_ms = @max(e.max_ms, line.ms);
}
fn print(r: *Report, out: *Io.Writer, buffer_size: usize) !void {
try out.print("read {d} lines, {d} bytes, through a {d}-byte buffer\n", .{ r.lines, r.bytes, buffer_size });
try out.print("skipped {d} lines we could not parse\n\n", .{r.skipped});
try out.print("status requests\n", .{});
for (r.statuses, 0..) |count, status| {
if (count > 0) try out.print(" {d} {d:>8}\n", .{ status, count });
}
const Row = struct { path: []const u8, e: Endpoint };
var rows: std.ArrayList(Row) = .empty;
defer rows.deinit(r.gpa);
var it = r.endpoints.iterator();
while (it.next()) |kv| try rows.append(r.gpa, .{ .path = kv.key_ptr.*, .e = kv.value_ptr.* });
// Slowest first. Ties go by name, so the order never depends on the
// hash map's.
std.mem.sort(Row, rows.items, {}, struct {
fn lessThan(_: void, a: Row, b: Row) bool {
if (a.e.average() != b.e.average()) return a.e.average() > b.e.average();
return std.mem.order(u8, a.path, b.path) == .lt;
}
}.lessThan);
try out.print("\nendpoint requests avg ms max ms\n", .{});
for (rows.items) |row| {
try out.print(" {s:<15}{d:>8}{d:>8}{d:>8}\n", .{ row.path, row.e.requests, row.e.average(), row.e.max_ms });
}
if (r.first_504_len > 0) {
try out.print("\nfirst 504 at {s}\n", .{r.first_504[0..r.first_504_len]});
}
}
};
/// A made-up hour of traffic from 22:00 to 23:00. From 22:40 the checkout
/// endpoint slows down and starts timing out, which is what the report has to
/// find. The seed is fixed, so every run writes the same file.
fn writeSampleLog(io: Io, dir: Io.Dir, name: []const u8, count: u32) !void {
const file = try dir.createFile(io, name, .{});
defer file.close(io);
var buf: [8192]u8 = undefined;
var writer = file.writer(io, &buf);
const w = &writer.interface;
var prng: std.Random.Xoshiro256 = .init(2026);
const rand = prng.random();
const paths = [_][]const u8{ "/", "/api/cart", "/api/checkout", "/api/search", "/login" };
const base_ms = [_]u32{ 12, 40, 90, 60, 25 };
for (0..count) |i| {
const second: u32 = @intCast(i * 3600 / count);
const minute = second / 60;
// u32, not usize: usize is 64 bits here and 32 bits in wasm, and the
// two widths draw different numbers from the same seed.
const p = rand.uintLessThan(u32, paths.len);
var status: u16 = if (rand.uintLessThan(u32, 100) < 2) 404 else 200;
var ms = base_ms[p] + rand.uintLessThan(u32, base_ms[p]);
if (p == 2 and minute >= 40) {
ms *= 9;
if (rand.uintLessThan(u32, 100) < 15) {
status = 504;
ms = 5000;
}
}
try w.print("2026-10-03T22:{d:0>2}:{d:0>2}Z ", .{ minute, second % 60 });
try w.print("{s} {s} {d} ", .{ if (p == 2) "POST" else "GET", paths[p], status });
// About one request in a thousand was aborted, and has no time.
if (rand.uintLessThan(u32, 1000) == 0) {
try w.writeAll("-\n");
} else {
try w.print("{d}\n", .{ms});
}
}
try w.flush();
}The program ran in your browser.
It wrote a 2 MB file, then read it back through the 64 KiB buffer.
Look at the checkout row.
The average is 699 ms.
That number looks bad, but not terrible.
It mixes forty calm minutes with twenty slow ones.
The last line is the useful one: the first 504 was at 22:40:00.
11:31 PM: one file, no install
Now the program had to run on web-03.
The server runs Linux on x86-64.
Shivam’s laptop could build for it directly:
$ zig build-exe -O ReleaseFast -fstrip -target x86_64-linux-musl logscan.zig
$ ls -l logscan
-rwxr-xr-x 1 shivam shivam 273624 Oct 3 23:31 logscan
$ file logscan
logscan: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), statically linked, stripped
-target names the machine to build for: x86-64, Linux.
This program does not use a C library at all.
Zig’s standard library talks to the Linux kernel directly.
So the result is “statically linked”: the file needs nothing else on the server.
No runtime, no shared libraries, no package to install.
It was 274 KB.
Shivam copied it over.
$ scp logscan web-03:/tmp/
11:33 PM: the real file
$ /usr/bin/time -f "%e s, %M KB" /tmp/logscan /var/log/shop/access.log
read 50000000 lines, 2075125000 bytes, through a 65536-byte buffer
skipped 51000 lines we could not parse
status requests
200 48471000
404 934000
504 544000
endpoint requests avg ms max ms
/api/checkout 9996000 699 5000
/api/search 9970000 89 119
/api/cart 9912000 59 79
/login 10115000 37 49
/ 9956000 17 23
first 504 at 2026-10-03T22:40:00Z
6.45 s, 652 KB
Fifty million lines in about seven seconds.
The whole process used 652 KB of memory.
Out of curiosity, Shivam also ran awk on the same file, counting only the status codes.
That took about 13 seconds.
11:36 PM
Shivam posted the report in the incident channel.
Only checkout is slow. Everything else is normal. First timeout at 22:40:00. What shipped around 22:40?
Two minutes later, Urvashi replied.
Payments team changed the timeout config for the card provider at 22:39.
They rolled the change back at 11:41 PM.
At 11:52 PM, the checkout error rate was back to zero.
The next morning
Shivam added logscan.zig to the team’s tools repository.
It is about 220 lines.
Anyone on the team can build it for any server with one command, including the ARM machines:
$ zig build-exe -O ReleaseFast -fstrip -target aarch64-linux-musl logscan.zig
Try it yourself
The full program is above the Run button.
Press Copy and save it as logscan.zig.
Then, with Zig installed:
$ zig run logscan.zig # writes access.log and scans it
$ zig run logscan.zig -- /path/to/your.log
Your own log will have a different format.
Change parseLine to match it.
Everything else stays the same.
These chapters explain each part in more detail:
- Reading line by line, including lines longer than the buffer.
- Hash maps, including who owns the keys.
- Cross-compilation, for building for another machine.
About the numbers
Shivam, Urvashi and web-03 are made up.
The numbers are not.
The 2 GB file is the sample log repeated 1,000 times.
Every time and size above was measured on a laptop with an Intel Core i5-1235U, using the Zig version shown at the top of this post.
Over three runs, logscan took between 6.4 and 8.2 seconds, and its memory stayed at 652 KB every time.
The output above is from the fastest run.
The awk command was awk '{c[$4]++} END{for(k in c) print k, c[k]}', which counts status codes and does nothing else.